| ▲ | radicalriddler 8 hours ago | |
Huh, according to some of those charts, it's both dumber, and more expensive to run against their benchmarking tasks than Fable??? Seems crazy to me. | ||
| ▲ | sejje 6 hours ago | parent [-] | |
Perhaps the model is able to evaluate that it's not done, and to keep pressing on in the face of mounting failures, until it eventually arrives at a solution. Where Fable can skip that. | ||