Remix.run Logo
▲ dotancohen 6 hours ago

  > Ask for something difficult from GPT-6 Sol and Opus 5.5 and watch what each one does.
That's far too vague. I found Opus to be terrific at coding, but human text just seems so robotic with it. OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural. What is "something difficult" in your workflow?
▲notatoad 2 hours ago | parent | next [-]

My side by side evaluation this week was to build a tool for mounting my app’s UI components in a headless chrome and feeding mock data into them, for the purpose of taking screenshots for help docs. Not super complicated, but a real task I needed done.

I gave the task to codex first, sol 6 xhigh. it took a couple back and forth prompts to define the project and then it worked for a bit and to took a couple more prompts before I decided it was good enough - not perfect, but close. It re-implemented some wrapper components in a simplified way that lost some of the UI, but it would work.

Opus 5.5 high took the same prompt with no back and forth, it just went off and one-shotted a tool that takes pixel-perfect screenshots of exactly what my app looks like.

▲peterbell_nyc 6 hours ago | parent | prev | next [-]

You HAVE to have a set of personal evals for each class of task you want to use models against at scale so you can test plausible candidates and compare output on your work against your evals.

There is way too much subtlety in what does and doesn't work for a given problem, context/prompt, tool set and eval. I can tell you Fable is generally better than Haiku, but comparing similar tiers really does depend on your exact context.

▲Starlevel004 5 hours ago | parent | prev [-]

> OpenAI models used to be the prototype for robotic text, but lately I've been finding them much more natural.

This was the biggest thing I noticed in the 6 models; their conversational prose is dramatically less grating.