| ▲ | Garlef an hour ago | |
> if it was smart enough and i think this is exactly the crux; the really big models need really big datasets and current gen LLMs get a lot of training data beyond "all books + all of the internet" the objection is then that producing this additional data would already confound it with pre "virtual cutoff date" knowledge (since the training data probably implies mathematical and SWE concepts that were developed post "virtual cutoff date") | ||