| ▲ | simlevesque 5 hours ago |
| We can't have hard evidence. It's a SaaS and they own the code and the machine it runs on. So it may be a widespread hallucination. But there's no evidence of that either. |
|
| ▲ | dooglius 2 hours ago | parent | next [-] |
| Run a benchmark with a large number of samples, rerun a few days later. Compare results, use statistics to see if there's a statistically significant difference. |
|
| ▲ | fragmede 4 hours ago | parent | prev [-] |
| We could still have soft evidence though. Make a Todo app on Monday, and make a Todo app on Tuesday, and see what it makes in comparison. |
| |
| ▲ | marcus_cemes 4 hours ago | parent | next [-] | | You would need a significant sample size to make any sort of conclusion from such a probabilistic process. Then there's the issue of how you would actually grade/compare. | |
| ▲ | ArvidSu 4 hours ago | parent | prev | next [-] | | You only need to come up with a catchy "SomethingBench" name, post it on reddit/x and now you're an ai sage. Not to disparage the launch/after comparison though, I'd genuinely enjoy a data point like that | |
| ▲ | luckydata 4 hours ago | parent | prev [-] | | someone already does that https://aistupidlevel.info/ | | |
| ▲ | chaimtweiss 3 hours ago | parent [-] | | It's actually a extremely cool site, and fascinating to view the results off the AI bots i use. |
|
|