Remix.run Logo
Mercury 2.5 LLM hits 770 tokens per second(artificialanalysis.ai)
11 points by Retro_Dev an hour ago | 6 comments
bearjaws 13 minutes ago | parent | next [-]

If you care about speed Cerebras gpt-oss-120b is 1400tk/s and "just as smart" in ranking.

I've used it on a few for fun projects and its decent but the speed is crazy to watch.

walrus01 31 minutes ago | parent | prev | next [-]

Pricing at $0.25 and $0.75 already puts its cost well above reasonably reputable inference providers for deepseek v4 flash or qwen 3.8-flash-next or similar class of open weight LLMs that fit in under 170GB of RAM, so I don't see the point. I think this is probably also stupider than laguna s 2.1 which can also be very cheap to serve.

rvz an hour ago | parent | prev [-]

The speed means absolutely nothing when it is finishing almost dead last when compared to the frontier AI companies.

glouwbug 5 minutes ago | parent | next [-]

Some of us want fast food

copperx 30 minutes ago | parent | prev [-]

Ah, the old "good, fast, or cheap; pick two" proves true once again.

downrightmike 11 minutes ago | parent [-]

Give it a few months.