| ▲ | minimaltom an hour ago | |
Worth distinguishing knowledge/task benchmarks from IF / agentic. It doesn't seem out of the question that you can have a small model thats generally good at instruction following and long-horizon agentic, as usually in those cases any requisite knowledge is in the context. Most of the benchmark improvements afaict are in agentic and instruction following benchmarks. | ||