| ▲ | saretup 3 hours ago | |
To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it. | ||
| ▲ | jjcm 3 hours ago | parent [-] | |
Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down. | ||