| ▲ | volotat 3 hours ago | |||||||||||||||||||||||||
There is no special algorithm, the finding is that slowing down the LR or the trunk, while keeping the LR of the experts is enough to eliminate most of the forgetting in the network. You can see in that experiment where chess data was the only thing the model read for 524K characters, yet it kept almost the same performance (i.e. held-out loss) on all other domains. If you keep LR the same across the whole network the loss in other domains degrades dramatically - this is a clear sign of catastrophic forgetting in action. What I can say for sure is that any traditional network that does pose a sign of catastrophic forgetting would not be able to learn any patterns from a single stream of data. There are no benchmarks published as the model is heavily undertrained, but it is learning. And you can see this clearly in the loss and samples even though they are still barely coherent. I am not an academic and am not trying to publish a paper about a “major breakthrough” or something like this. I am just a small person who found a cool thing that clearly works and wants to share it with the world. That’s it. | ||||||||||||||||||||||||||
| ▲ | ilaksh 3 hours ago | parent [-] | |||||||||||||||||||||||||
You can't claim it "works" if it hasn't produced any coherent responses and is still early in your first training attempt. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||