| ▲ | mdp2021 a day ago | |
Sure, the "6b subset" can be more knowledgeable on its area than a whole 27b generalist (and more efficient), but where is the simulated Intelligence encoded? A 6b subset as or more intelligent than a 27b raises the question of how metacognition skills are stored. | ||
| ▲ | mapontosevenths a day ago | parent [-] | |
MOEs are built by training a second "router" model to identify which parts matter inside the dense model. Think of MOE as ignoring noise, rather than a more efficient encoding of the data we teach it which ones can be ignored. Turning down the noise actually sharpens the results sometimes. Modern LLM's are wildly inefficient. | ||