Remix.run Logo
bad_haircut72 an hour ago

not an AI researcher - this is probably true for these "everything" LLMs but I think specialized models are gonna be the next big thing

ACCount37 an hour ago | parent [-]

"Specialized models" are a bit of a doozy.

The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very little performance, and for too much effort.

Generality and intelligence seem to be entangled very heavily in LLMs.

CamperBob2 21 minutes ago | parent [-]

And yet, there's VibeThinker 3B to bring this long-held premise into question (if not to blast it to pieces.) It is practically illiterate by the standards of larger models, yet performs like models 100x its size on mathematical and logical reasoning tasks.