Remix.run Logo
jmugan 3 hours ago

A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)

IsTom an hour ago | parent | next [-]

Aren't they still bad at understanding how bicycle frame works? Especially the steering part?

gegtik an hour ago | parent | next [-]

Maybe that makes them human..

https://www.booooooom.com/2016/05/09/bicycles-built-based-on...

nozzlegear 30 minutes ago | parent [-]

Humans aren't machines trained on the entire stolen corpus of human knowledge. We expect that a human will do poorly at arbitrary tasks they have no experience doing – especially drawing, which many (most?) humans aren't trained in at all. The same isn't true of the AI, whose proponents and priests have, for over a year or more, spent time, energy and billions of dollars attempting to convince us that the clankers can do anything.

zh3 an hour ago | parent | prev | next [-]

Shhh...you'll alert the models :)

Totally agree though, anyone with a vague understanding of how bikes works ignores the pelican because they know the bike is unrideable in the first place.

edaemon an hour ago | parent | prev [-]

Yes, but I think the idea here is that most models produce very similar pelicans on bicycles, so a different test might be more useful in gauging the differences in models.

maxutility 3 hours ago | parent | prev | next [-]

Agree. The pelican benchmark was interesting a year ago when most models struggled and a good pelican indicated an unusually capable model. Now it’s saturated and uninteresting.

A good new benchmark should have awful performance to start and there should be a lot of headroom for improvement. This benchmark is also intentionally difficult and requires the LLM to develop the animation through spatial reasoning and first principals rather than existing video generation pipelines. Similar to how SVG generation was out of distribution for most models a year ago.

irthomasthomas 31 minutes ago | parent | next [-]

I think it's interesting to see them visibly struggling to improve. Claude pelicans aren't much better today than they where 18 months.

sixtyj 30 minutes ago | parent | prev [-]

Have you seen pelicans in Simon Willison’s tests? It is still not a pelican on bike I would like to publish :)

Imho we don’t need to make benchmarks that draw the whole 3D world. Pelican’s drawing is really nice in its simplicity and complexity at the same time.

It seems to be an obscene waste of compute time to generate useless 3D worlds that are just a bragging - 3D is really heavy discipline to make it right, see Mark Zuckerberg’s ceased attempt with 3D VR…

Multiply it by thousands times as a lot of people have found out threejs lib and prompt “generate 3D world and make no mistake” are new orange/black.

dofm 2 hours ago | parent | prev | next [-]

But it's another benchmark on how good models are at generating intensely average, unwanted things with unthinking design. Just scaled up.

charcircuit an hour ago | parent | prev [-]

Bad? It has a charming style. I would watch the whole book if it was made like this.

trial3 an hour ago | parent | next [-]

yeah, definitely, in the same way that we all regularly go and look back fondly at our chatgpt ghiblified family photos

altmanaltman 42 minutes ago | parent | prev [-]

Very few things are universally hated. One can love something truly that is hated by most. But it doesn't change the fact that it's still hated by most. An objective and a subjective opinion can exist at the same time on this.