| ▲ | worldsavior a day ago | |
That's a 2.4T model, how would they reduce this to 35B and still give some accuracy? That's a completely different arch. | ||
| ▲ | cyanydeez a day ago | parent [-] | |
there's been a lot of research about reducing models by taking out layers; there's also using it to train smaller models by optimizing parameters. I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling. | ||