|
| ▲ | NineStarPoint 5 hours ago | parent | next [-] |
| Yeah q8 made so littler difference back when I was testing such things I'd be surprised if people could quickly notice that as a change. It's got to be either further quantized or some other type of optimization that kicks in when people notice the drop. |
| |
| ▲ | selectodude 5 hours ago | parent [-] | | NVFP4 would buy them a huge increase in capacity but I think it would be noticeable. | | |
| ▲ | Caracas288 4 hours ago | parent [-] | | Why doesn't someone just try to measure this next time!? | | |
| ▲ | embedding-shape an hour ago | parent [-] | | Can't really measure without being sure you aren't being messed around with, when it's a remote platform. Stupidly easy to detect when people run such benchmarks/tests against you as well. |
|
|
|
|
| ▲ | torginus 4 hours ago | parent | prev | next [-] |
| Some people here have remarked previously that while reduced precision doesn't show up in quick prompts, it does severely impact these models' ability to perform long running tasks - to the point that running these big models with severe quantization might be counterproductive as smaller but less quantized ones perform better. |
|
| ▲ | nonethewiser 4 hours ago | parent | prev [-] |
| Could this explain Opus? |