| ▲ | tripleee an hour ago | |
He's intentionally forgoing all nuance in order to make this sound dramatic > Somewhere around 20M people (0.2%) see first-hand that large, complex projects that used to take them weeks/months can now be completed by agents with a prompt. No, they can't, at least not any semblance of quality. The cases we're seeing where this does kinda work is in ports and translation where all the rules are already documented in the best specification language possible with a way for the LLM to verify itself: code. We saw this close to a year ago now with Cloudflare and NextJS > The impact scales with ambition, problem size, and horizon. A question with a paragraph answer barely stresses the system. You need a reservoir of big, difficult problems that you really care about These are operating on different capabilities - AI's ability to answer informational queries as a chatbot frankly sucks and can't be trusted without verifying it. I run up against this every day. A problem with a verifiable answer on the other hand it's very good at solving. He knows this (his next paragraph) but he's putting them on the same scale of "stressing the system" to attempt to add proof to his introductory claim And then there's the completely unverifiable scare that there are internal frontier models way beyond anything we've seen "swarms of thousands of agents collaborating over weeks on software mega projects: minting zero days, running cyber attacks and defenses at machine speeds, discovering new science, advancing the frontier of mathematics" I dunno. I haven't been able to set up OpenAIs remote codex connection, their shit is buggy as hell and the web UI keeps crashing and making messages disappear. Is this what their internal superhuman "Things that would have taken top professionals in the industry years of work" looks like? Granted Claude has been really smooth, but still.. | ||