| ▲ | NitpickLawyer 4 hours ago | |||||||
It's interesting that this is the third lab to find problems with larger models. Earlier last year oAI was rumoured to have failed their large pretrain. Now google has problems with their pro series, and ds just announced the same. There are some rumours on chinese forums talking about problems with the pretraining phase, so this is not mid/post training related. I wonder if this comes from using the bad architecture scaled up (and it hits some limits) or if this is a data problem (undertrained? bad data? bad pre-processing using smaller models?)... | ||||||||
| ▲ | wolttam 3 hours ago | parent | next [-] | |||||||
Just my intuition about it but it does seem like a data issue. V4 flash and V4 pro feel very similar, which would make sense if they were pre-trained on largely the same corpus. All that would suggest to me is that V4 Flash is capable of absorbing the data they’re throwing at it, and we’re still nowhere near the data limits of their larger 1.6T model | ||||||||
| ▲ | pixelesque 3 hours ago | parent | prev [-] | |||||||
Where are you seeing them having an issue with the larger (Pro) model? The announcement specifically says 4.1 Pro will be released in the future. | ||||||||
| ||||||||