Remix.run Logo
fragmede 4 hours ago

We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison.

marcus_cemes 4 hours ago | parent | next [-]

You would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare.

ArvidSu 4 hours ago | parent | prev | next [-]

You only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that

luckydata 4 hours ago | parent | prev [-]

someone already does that https://aistupidlevel.info/

chaimtweiss 3 hours ago | parent [-]

It's actually a extremely cool site, and fascinating to view the results off the AI bots i use.