Remix.run Logo
tzone 5 hours ago

Which model are you using? There is monumental difference between models. Even between "frontier models". When people tell these stories, it would be great to also add which model you were using.

As an example Opus 5.0 is in completely different class compared to Cursor Grok 4.5 even if the benchmarks don't show such massive difference. Not even talking about regular stuff like Sonnet or Composer or stuff like that.

jrockway 5 hours ago | parent | next [-]

Yeah. If you need something to dig deep, you need to try Fable (optionally in /goal mode).

For performance testing, I wrote isolated testbeds that try to impair the system in realistic ways (latency/jitter/bandwidth limit on logical WAN hops when load testing), and Fable is happy to send a bunch of agents at it and iterate until it gets the results it's looking for.

I think that if you are used to Sonnet medium or something, this will surprise you, but models like Fable and Sol on high/xhigh will really dig deep until they meet your goal. (I mostly use this for bug hunting and not perf, but ... I think it can do perf if you set it up right.)

vonneumannstan 5 hours ago | parent | prev [-]

It's almost always a version of "I used the free version of ChatGPT 10 months ago and it couldn't code well."