| ▲ | fph 2 hours ago | |||||||||||||
But do you really need a model that has the complete Duran Duran discography memorized and preloaded in RAM at all time? | ||||||||||||||
| ▲ | amelius 2 hours ago | parent | next [-] | |||||||||||||
That's a different question. Probably not. But: 1. Training a large model with lots of information, then stripping the "useless" information from that model to obtain a small model => nobody has shown this. 2. Training a small model, letting it use a database tool so it scores the same as a large model without database => nobody has shown this. | ||||||||||||||
| ||||||||||||||
| ▲ | rotis 33 minutes ago | parent | prev | next [-] | |||||||||||||
Sure. I don't give a damn about them. But Phil Collins man. I love reviewing his body of work before I get worked up. Obvously we cannot cut him, because I need him. So how do you decide what to leave out? | ||||||||||||||
| ▲ | embedding-shape 2 hours ago | parent | prev | next [-] | |||||||||||||
I mean maybe yes? The hypothesis from the early GPT days was (and in a small way still remains): "If we just chuck more data into the training, does it get better at X, even if the data was seemingly unrelated to X?", and the workings of LLMs seem to kind be pointing in that direction, although with some ceiling. But seemingly models good at programming for example, would get worse at programming if you removed everything not-programming. Train a model solely on syntax, and it'll be worse than a general purpose LLM on syntax, in general at least. | ||||||||||||||
| ||||||||||||||
| ▲ | medwards666 an hour ago | parent | prev [-] | |||||||||||||
But what if I _really like_ Duran Duran??? | ||||||||||||||