| |
| ▲ | baq 17 hours ago | parent [-] | | Yup that’s quite literally what ‘agency’ is and the whole point of agentic workflows. Personally I’ve had them running for days with good results and as you can see OpenAI had them running for months, and yes indeed the things made some very questionable decisions and assumptions… but they unquestionably did a lot of stuff correctly, for some definitions of ‘technically correct’. | | |
| ▲ | realusername 16 hours ago | parent | next [-] | | Personally I don't believe in agentic workflow. I don't think that's a coincidence that both OpenAI and Anthropic chose math problems to test their long agentic workflows, they are well defined, with a clear finish line and with a 100% clear progress path, most of real life tech projects aren't like that. No matter how clever the model is, most problems have multiple valid, invalid and unclear decisions to make, running it for a long time is just picking the first option on everything, which isn't usually what you want | |
| ▲ | koonsolo 15 hours ago | parent | prev [-] | | Can you describe what came out, after letting it run for days? I always have a hard time understanding what kind of task would be worth it. | | |
| ▲ | baq 13 hours ago | parent [-] | | Fed it a list of bugs (with good descriptions, screenshots, etc., pointers to code) and it fixed them |
|
|
|