| ▲ | Mercury 2.5 LLM hits 770 tokens per second(artificialanalysis.ai) | ||||||||||||||||||||||
| 11 points by Retro_Dev an hour ago | 6 comments | |||||||||||||||||||||||
| ▲ | bearjaws 13 minutes ago | parent | next [-] | ||||||||||||||||||||||
If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking. I've used it on a few for fun projects and its decent but the speed is crazy to watch. | |||||||||||||||||||||||
| ▲ | walrus01 31 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||
Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve. | |||||||||||||||||||||||
| ▲ | rvz an hour ago | parent | prev [-] | ||||||||||||||||||||||
The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies. | |||||||||||||||||||||||
| |||||||||||||||||||||||