| ▲ | mapontosevenths 2 hours ago | ||||||||||||||||
Capability in LLM's is distributed throughout the manifold in subspaces. Even worse, the subspaces exist in superposition. That is to say, there is no single 'python' part of the model. The python bit is spread throughout the entire model and overlaps with other pieces that have similar, but unrelated, capabilities. For example the python subpspace might be partially in superposition with cupcake recipes, Esperanto, and calculus. We need calculus in a coding agent but not the other two. However, separating them cleanly is almost impossible, and even identifying them is tough. Internally the manifolds are highly inefficient and nothing like you would imagine something humans built would be designed. It's more like something that evolved in nature. | |||||||||||||||||
| ▲ | Manfrednotfunny 2 hours ago | parent | next [-] | ||||||||||||||||
My current image from a MoE is that the base/core might be the more generic thing and that things like python are part of one expert though. | |||||||||||||||||
| |||||||||||||||||
| ▲ | dist-epoch 2 hours ago | parent | prev [-] | ||||||||||||||||
There was some paper about routing at training bio-knowledge into a particular region of the model, which you then can cutoff when serving. But you probably lose some efficiency since maybe you sized that region too small/too big. | |||||||||||||||||
| |||||||||||||||||