| ▲ | 44% on ARC-AGI-1 in 67 cents(mvakde.github.io) | ||||||||||||||||||||||||||||||||||||||||
| 40 points by porridgeraisin an hour ago | 13 comments | |||||||||||||||||||||||||||||||||||||||||
| ▲ | pwmglenn a minute ago | parent | next [-] | ||||||||||||||||||||||||||||||||||||||||
Really impressive and creative research. I wonder if the leading labs do anything similar with their models? It doenst look like the open source labs do? | |||||||||||||||||||||||||||||||||||||||||
| ▲ | larodi 5 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
"I don’t understand why others didn’t figure this out" - how about we allot the possibility that so many of presumed ML experts don't have any clue what they be doing, and are eventually API bitches, nothing more. | |||||||||||||||||||||||||||||||||||||||||
| ▲ | xeonax 43 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Even cooler is his about me mention of saving his own life https://mvakde.github.io/ > Saved myself in a medical emergency (doctors didn't know what rhabdomyolysis was) | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | embedding-shape an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||||||||||||||
Is the author only running their model against one benchmark? I don't think anyone finds that difficult to achieve, the difficulty comes when you want to make the model not benchmaxxed to a specific benchmark, and generalize so it can solve problems not part of the training data, but seems this model is specifically for not this? How useful is that? If you just wanted to pass these specific tasks in this specific benchmark, and wanted to do so cheaply, I'm sure a non-LLM-based approach would yield better results for even cheaper, since what the author's model does, seem to basically be "solve ARC puzzles", not a general LLM or "coding" LLM. | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||
| ▲ | eis 33 minutes ago | parent | prev [-] | ||||||||||||||||||||||||||||||||||||||||
> Increases in LLM scores are now mainly driven by post training (evidence in next section) and are probably a function of amount of synthetic data. They are learning to solve ARC tasks, not learn general abstract reasoning Agreed and that's for any benchmark. Private tests are better but you still have to trust the provider to not log and use them for training. That's why I like when a new set of tests like a new ARC-AGI version is published, that's where you can see which of the models abstracted to more general capabilities instead of being focused on the previous tasks. Most models completely fail new ARC-AGI tests. The "67 cents" part though is misleading imho. You can't extrapolate from there and think that investing say $100 will get you a lot better results. You hit a ceiling very fast and investing into more compute will give you diminishing results. So yes, you can train a custom model to do somewhat decently on a specific set of tasks but then what? | |||||||||||||||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||||||||||||||