| ▲ | Art9681 a day ago | ||||||||||||||||
The absolute best way to prove this works is by releasing a model that was fine-tuned with this method and then showing benchmarks depicting the improvement delta between the base model and the fine tuned one. The work is not done. Then release it to the masses and wait a few days for the actual real world anecdotes. Until then, this is noise. | |||||||||||||||||
| ▲ | Reubend a day ago | parent | next [-] | ||||||||||||||||
Yeah, this is just slop. No benchmarks, no concrete case studies, just some vibecoded "platform" to finetune models on your own traces. Which is an idea that has some value, but also some weaknesses. And this implementation of it isn't forthcoming with that concept. You have to really dig in to understand what they're even talking about. | |||||||||||||||||
| |||||||||||||||||
| ▲ | SilenN a day ago | parent | prev | next [-] | ||||||||||||||||
Valid criticism. Happy to answer any qs. We're still working on solidfying results. | |||||||||||||||||
| |||||||||||||||||
| ▲ | irishcoffee a day ago | parent | prev | next [-] | ||||||||||||||||
Benchmarks are the ultimate consolidation of halnons razor. | |||||||||||||||||
| ▲ | teravor a day ago | parent | prev [-] | ||||||||||||||||
[flagged] | |||||||||||||||||
| |||||||||||||||||