| ▲ | Farmadupe an hour ago | |
Yeah, I'm with you on this, I think this is just what fable/opus-5 slop looks like now... * "A frozen verify.py accepts the claim" (what does it mean to freeze a python script?) * "which trains the recipe eight times on fixed seeds it can't touch" (what does it mean to not be able to touch a seed) * "One other detail is that we gave an estimation of the speedrun noise in program.md that was slightly too large. 62 out of ~100 runs measured it themselves instead of trusting our number" (What does it mean for a "run" to "distrust" a noise measurement) * "One important disclaimer is that our benchmark has a lot of variance" (Actually this one makes sense, but congratulations for burying the lede) * "Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind" (This implies the the graph would show every model finding a plateau in whatever metric the experiment is measuring, but I don't see every model scoring the same in teh graph) | ||