| ▲ | vkaku 3 hours ago | |
This is great. All algorithmic efficiencies are amazing! One thing I'd remind all scientists and the wonderful people here is this wonderful meme/line from Jurassic Park: "Your scientists were so preoccupied with whether they could they didn't stop to think if they should." What is the actual amount of data that needs to be pre-trained and what is not? Nobody has come up with great answers to this question, and I'm already seeing amazing 0.5b-2b parameter models working very well with n-Gram corpuses of data. So, how many parameters do you really need for a given workload? | ||