| ▲ | boredatoms 6 hours ago | |||||||||||||||||||||||||
For sure they quickly move to q8, the output quality difference to bf16 is small compared to the speed/capacity gain | ||||||||||||||||||||||||||
| ▲ | NineStarPoint 5 hours ago | parent | next [-] | |||||||||||||||||||||||||
Yeah q8 made so littler difference back when I was testing such things I'd be surprised if people could quickly notice that as a change. It's got to be either further quantized or some other type of optimization that kicks in when people notice the drop. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | torginus 4 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
Some people here have remarked previously that while reduced precision doesn't show up in quick prompts, it does severely impact these models' ability to perform long running tasks - to the point that running these big models with severe quantization might be counterproductive as smaller but less quantized ones perform better. | ||||||||||||||||||||||||||
| ▲ | nonethewiser 4 hours ago | parent | prev [-] | |||||||||||||||||||||||||
Could this explain Opus? | ||||||||||||||||||||||||||