Remix.run Logo
madrox 11 hours ago

Mechanical Turk had a good run, but not surprised it's shutting down. I'm sure the platform was getting flooded with people doing task arbitrage and using lots of AI anyway.

I believe the issue is that this can no longer be a horizontal play. MTurk was mostly for unskilled tasks...the kind AI can do well enough that it isn't worth the cost differential to verify it or keep farmed to humans. The "trust but verify" AI output is now the kind that requires domain expertise. This is what most full stack AI companies are bringing to industries.

Curious if this kind of work will come around again one day or was just a moment in time. If it does I'm sure it will be specifically about generating training data.

johnsmith1840 11 hours ago | parent | next [-]

I think yes but on different categories. First one I imagine is robotics control and support.

"This robot is having trouble folding a tshirt help it out for 1$"

Unless they go the waymo route of highly trusted people but I think mass deployed robots are a bit safer than a car for this.

georgefrowny an hour ago | parent | next [-]

Giving out control of industrial machinery that interacts in the human environment without the physical interlocks (i.e. humanoid robots in a house) to random internet people seems like a problem.

I can imagine a carefully orchestrated plot to assassinate someone by having an embedded agent in the task delegation pool command the laundry bot to punch the target's head off their shoulders.

dividedbyzero an hour ago | parent [-]

Giving out control of such things to LLMs is already complete madness, so once the first pleasure bot powered by Grok has dismembered a few thousand users, they'll get sophisticated safety mechanisms.

Though like as not you're still going to be right, after all, Stuxnet happened.

madrox 10 hours ago | parent | prev | next [-]

This is what I meant by full stack AI companies. I don't think you could get humans into the loop fast enough if they didn't have some idea of the type of task involved. I don't want people to be asked to fold a tshirt one moment and do a difficult traffic merge the next.

its-summertime 9 hours ago | parent | next [-]

There is training systems and validation of skills in mturk iirc: for tshirt folding, you'd be given fake setups to be able to get used to controlling the robot, if you can't do it, you won't ever get assignments to do it. For traffic overrides, you'd be tested on having correct knowledge, and once again given supervised tasks to show you can actually be trusted (and there would be safety systems, elevating tasks that can't be performed at your level to people who can, etc)

The worker needs to do a bunch to opt into any given work group, which makes the (lack of) payments extremely unreasonable on top of everything else

inigyou 6 hours ago | parent | prev | next [-]

It doesn't really matter what you want though, only what CEOs want and that's low costs. I can see a combined shirt folding/traffic merging platform taking off.

bonoboTP 2 hours ago | parent [-]

CEOs are loser nobodies. Real influence is with owners, boards, investors.

ntauthority 10 hours ago | parent | prev [-]

warioware shows it can be fun though but i'd indeed not like to see that applied to safety-critical tasks

mike_hearn 3 hours ago | parent | prev | next [-]

There are already robotics models that can fold shirts and similar just fine. Progress in VLA models is good, I don't think this would form the foundation of a business. Humanoid robots are going to be another ChatGPT, it's going to seem to happen almost overnight because people aren't paying attention to the underlying research papers.

simsla 2 hours ago | parent [-]

Any papers you'd recommend?

vidarh 5 hours ago | parent | prev | next [-]

There are already a number of companies providing RLHF and SFT services for AI providers that does a lot of validation/prequalification of people that'd be well placed to take on tasks like that, but the big problem to solve would be latency if you don't have people contracted to carry out a task right now.

mikestorrent 10 hours ago | parent | prev [-]

That's a remarkable idea. It could be heavily gamified, it could train models, and it might actually be mentally stimulating since you'd be facing different situations all the time.

Except, I'm a grown adult and I can't fold a t-shirt properly

madrox 10 hours ago | parent [-]

It could require listing your credentials to get you the proper tasks: doctor, lawyer, or tshirt folder

Yokohiii 8 hours ago | parent | prev | next [-]

It is weird because the last time I've heard about MTurk was about developing countries being rather reliant on it for doing AI grunt work. If am not totally wrong this must mean that the data work has moved to other services.

toyg 3 hours ago | parent [-]

Probably AI models got good enough to bootstrap their own training systems.

SV_BubbleTime 7 hours ago | parent | prev | next [-]

> task arbitrage and using lots of AI anyway.

Oh, I remember UpWork.

shuwix 3 hours ago | parent [-]

Oh ... upwork ... put offer, get 50 replies from people which jump for every penny without even being able to understand the task.

raverbashing 5 hours ago | parent | prev | next [-]

"Fun" fact, they were also used by psychology/etc students when they need to do 'research' with X amount of people

timcobb 11 hours ago | parent | prev | next [-]

How could it come around again?

Legend2440 10 hours ago | parent | next [-]

There are like three dozen companies selling similar services for generating AI training data.

charlieyu1 36 minutes ago | parent [-]

Good data isn’t cheap and keeping it in-house gives you more control.

madrox 11 hours ago | parent | prev [-]

I'm unsure. If it does, it will be work that is too expensive or inaccurate or regulatory for current AI methods. For example, you want a doctor to sign off on some AI output on a diagnosis.

However, I'm not sure a single platform will be how it emerges

yashvg an hour ago | parent | prev [-]

[flagged]

ZitchDog an hour ago | parent | next [-]

How do you make sure they aren’t using AI to pretend to be a domain expert? I’d think an AI would be pretty good at that.

yashvg 41 minutes ago | parent [-]

We do our best to identify AI use and ban those experts - its not perfect just yet. Long term we are thinking of moving towards proctoring experts using their camera and screen capture. Hard to think of another reliable way.

fireant an hour ago | parent | prev [-]

Some doctors are known to be a hip shooters making snap decisions, but is 30s really enough time for any kind of "expertise" from a real human?

yashvg an hour ago | parent [-]

You get matched with an expert in 30s to 3min, which then starts a back and forth with the AI agent and the expert which goes on for 10 to 60mins till the AI is happy.

This works well for many tasks, but for others a more async mechanism where the expert doesn't feel rushed might be better.

testplzignore 30 minutes ago | parent [-]

> Experts are vetted upfront by an AI interviewer

> till the AI is happy

My god this is dystopian.