| ▲ | dannyw 40 minutes ago | |
Small models are still great for lots of “simple intelligence” use cases, like annotating or summarising files and media; or even just basic chat when given web search tools. My local NAS is private and I’m not going to send it off to APIs for captioning or metadata; but even Qwen3VL 8B does an excellent job at this, despite being quite old. They are also really excellent for fine tuning. Unsloth and Tinker (from Mira’s TML) are great places to start. If your use case is narrower than “coding agent for everything”, you can probably match frontier performances on that narrow domain with ~30b and exceed it with ~100b+. | ||