Remix.run Logo
A $500 RL fine-tune of a 9B open model beat frontier models on catalog review(fermisense.com)
39 points by ilreb 2 hours ago | 7 comments
himata4113 an hour ago | parent | next [-]

What I really started to notice is that SOTA models are really good at putting themselves out of the job.

We can see this already with GPT how luna can do 90% of what sol is used for. The only reason why china still bothers 'distilling' models is accurate training data generation, something that oai and anthropic had to spend years collecting while trying to dodge legal challenges.

The more intelligent models get, the more people will offramp to cheaper solutions that get the job done. There's no real benefit to using a sota model when the accuracy is already 99% and I think that is the biggest danger to US labs.

com2kid 33 minutes ago | parent [-]

The upper end is all about coding. If I have terra on extra high write code, Sol will find a plethora of bugs and rip the code apart.

Anything else? Sure use a cheaper model.

himata4113 15 minutes ago | parent [-]

You can have sol write code and terra will find a plethora of bugs and rip the code apart. In reality this is just the nature of advisory prompting and why advisor from omp.sh is such a great feature. They get caught as they're being written.

JSR_FDED an hour ago | parent | prev | next [-]

I like the 2x2 grid that describes when to fine-tune a model, when to use a frontier model, etc.

From the article it’s not clear how the scorer grades every episode - was it a frontier model that assigned the grade? How does that continue to work as the model that is being fine-tuned becomes better at the task than the frontier model?

nzeid an hour ago | parent | prev | next [-]

I didn't read the Ramp article but this reads like a post hoc fallacy. Companies with 2x revenue have money to spend on AI. Companies with 1.15x revenue don't.

sudo_cowsay 38 minutes ago | parent | prev [-]

What benchmark is it? Is it super niche?

stldev a few seconds ago | parent [-]

They built their own benchmark and then trained directly against its scoring function.. seems to be the rage, but nothing convincing from the article alone.