Remix.run Logo
▲ soerxpso 7 hours ago

I don't see why I would be interested in this model, considering the price difference. They advertise that it's the same as Kimi K3 in half the tokens. But the pricing is double the pricing of K3. So why do I care if it uses fewer tokens, if I'm paying double per token?

▲bigmadshoe 6 hours ago | parent | next [-]

Because you care about how many tokens are used per task. What you said is like only caring about the price of gas and not gas mileage of your car.

▲verdverm 6 hours ago | parent [-]

it's not exactly the same, the model stills "weighs" the same

here, it does less work, it's more like driving half as far but still paying the same total cost

this being said, K3 and E1 models are priced the same at $3.00 / $0.30 / $15.00

https://fireworks.ai/models/fireworks/kimi-k3

https://fireworks.ai/models/fireworks/ember-1

▲bigmadshoe 6 hours ago | parent [-]

From reading the blog post, it is essentially exactly the same as the car example. It delivers the same performance on tasks, but using 40% fewer tokens. This is the same as a car getting you to the same destination but wasting less energy on excess heat, wind resistance, or whatever else affects fuel economy (I am not an expert, obviously). I am paying for an LLM to complete tasks for me, not for the intermediate tokens.

▲verdverm 6 hours ago | parent [-]

I'm with you, I wrote prior comment under the assumption that they were priced differently (from GP claim as such), but they are priced the same (on Fireworks)

Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs

▲seizethecheese 5 hours ago | parent | prev | next [-]

Half the tokens presumably means tasks get done twice as fast.

▲ersiees 6 hours ago | parent | prev | next [-]

It’s same price for lower latency.

▲verdverm 6 hours ago | parent | prev [-]

the pricing is the same, where are you seeing double?

▲tyingq 4 hours ago | parent [-]

K3 is cheaper from other providers. $1/$9, though I can't speak to how good those providers are.

https://openrouter.ai/moonshotai/kimi-k3

▲polski-g an hour ago | parent | next [-]

K3 TOS says they must sell no lower than what Moonshot charges. If you have a provider selling for less than $15/mtok, they are violating TOS from Moonshot.

▲verdverm 13 minutes ago | parent [-]

looks like the underlying vendor is charging correct prices https://inference.net/models/kimi-k3/

however the listing on open router has a `/fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?

▲verdverm 3 hours ago | parent | prev [-]

as the saying goes, you get what you pay for

we require ZDR and Fireworks provides that on contract, so for us they are the same price

▲indigodaddy 2 hours ago | parent [-]

Check out Neuralwatt. They are ZDR and great energy based K3 pricing. They also have a K3-fast which is basically no reasoning (in addition to regular K3).

▲verdverm 2 hours ago | parent [-]

that is nothing like how we consume Ai

industry standard is price-per-1M tokens, don't do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)

GPUs are rented in $/h, like every other piece of hardware in cloud