| ▲ | JCharante an hour ago | |
I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead. | ||
| ▲ | barake an hour ago | parent | next [-] | |
Anecdotally, it feels like Opus, Fable, and Sol "get distracted" when you use them for writing code. Great at reasoning and coordination but they will go off on a tangent and refactor half the code base. I only use them for reasoning (of course) and coordinating subagents. | ||
| ▲ | andrenotgiant an hour ago | parent | prev [-] | |
Any data or public links you can share? That surprises me | ||