| ▲ | peri-cl 6 hours ago | |||||||
> "Also maybe my setup (OMP) doesn't do the cache correctly but that's a huge cost driver... so atm it's quite pricy" I don't believe Cerebras has a cached input pricing? They don't list one on the model page: https://inference-docs.cerebras.ai/models/qwen-3.8-27b edit: See the sibling discussion, https://news.ycombinator.com/item?id=49554520#49555094 ("Input tokens, whether served from the cache or processed fresh, are billed at the standard input token rate") | ||||||||
| ▲ | hexa00 6 hours ago | parent | next [-] | |||||||
lol yeah just saw that, yeah that makes it unusable I think at least for me. I wonder if they will do that with sol ultrafast! | ||||||||
| ▲ | olivermuty 6 hours ago | parent | prev [-] | |||||||
They have cache, but it costs the same indeed, no idea what the point of the cache is | ||||||||
| ||||||||