| ▲ | nostrebored 6 hours ago |
| 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool. Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information. ```
Billing access restricted
Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions.
``` We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives: ```
{"message":"Model does not exist or you do not have access to it.","type":"not_found_error","param":"model","code":"model_not_found"}
``` When the error is really about billing. I always want to like Cerebras, but I get the vibe that as a tokens in tokens out consumer you are not valued at all. |
|
| ▲ | Aurornis 4 hours ago | parent | next [-] |
| > 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. I don't understand. How does that make it unusable? Is the limit shared by an entire team at once? 150,000 tokens per minute is a lot. You could start hitting that with a lot of concurrent requests in your session, but even throttled to 150k TPM it's still going to be faster than anything else you find. I think the 128K context limit is the real ceiling. These models aren't amazing at long context, but once you account for a short input prompt, the input files, and headroom for a compaction summary, there isn't a lot left for the problem. |
| |
| ▲ | wild_egg 3 hours ago | parent | next [-] | | It's a limit on input tokens. So that's 3 50k requests per minute. At Cerebras speeds, that's about 5 seconds of usage per minute. I was very excited last year for their coding plan but seeing a burst of requests pulse and then sitting there watching the cooldown reset is really not a great time. Even though each individual request was fast, the sessions were only maybe 10% faster on wall clock time since there was so much waiting time. | | |
| ▲ | amelius 3 hours ago | parent [-] | | Can't you do something with multiple accounts? | | |
| ▲ | sandworm101 3 hours ago | parent [-] | | Or just buy a 5060. This will run on most any 16gb card. Slower for sure but far cheaper than another subscription. | | |
| ▲ | embedding-shape an hour ago | parent [-] | | Or buy a raspberry pi with a SSD, about the same difference, if you're giving up on the 1500 tokens/s anyways. |
|
|
| |
| ▲ | gerdesj 3 hours ago | parent | prev | next [-] | | 128k context is not a limit of the model, that's a limit of implementation: "Context Length: 262,144 natively and extensible up to 1,000,000 tokens." https://huggingface.co/Qwen/Qwen3.8-27B | |
| ▲ | datadrivenangel 3 hours ago | parent | prev | next [-] | | 150k tokens per minute at 1.5k tokens per second means you can have like 3 users concurrently and that's not a lot. | |
| ▲ | conception 4 hours ago | parent | prev | next [-] | | 150k by account. At 1.5k a second you hit it very quickly. | | |
| ▲ | devy 3 hours ago | parent [-] | | Exactly, it burns the tokens 3000x faster, which means the budget ($$$$$$) runs out so faster it will stop super quick, not able to perform long-duration work. At 27B parameter size, the intelligence is not able to accomplish work within a short amount time. Consequently, it become not usable. | | |
| ▲ | a012 33 minutes ago | parent | next [-] | | Unusable is too stretch IMO, you can still use it in tiny tasks that’ll respond almost instantly | |
| ▲ | gerdesj 3 hours ago | parent | prev [-] | | I (we) run Qwen3.8-27B-FP8 on a DGX Spark box - that's roughly £4000 of hardware. I did benchmark it in various ways and it runs quite well but it is a quantised jobbie and 1.5k t/s is also rather faster than anything I can possibly hope to achieve. To run that model at those sorts of speeds is going to need some serious investment and you are going to have to pay for it. | | |
|
| |
| ▲ | 2 hours ago | parent | prev [-] | | [deleted] |
|
|
| ▲ | olivermuty 6 hours ago | parent | prev | next [-] |
| Cerebras the tech is awesome, cerebras the company is a trainwreck |
| |
| ▲ | dd8601fn 3 hours ago | parent [-] | | Is this the chatjimmy asic approach with a bigger model? | | |
| ▲ | ericd an hour ago | parent [-] | | No, the asic could only ever run one model/set of weights, no updates possible, ever. These are general purpose processors that can have their models updated. But the chips are enormous, with a substantial amount of on-die memory alongside the execution units, for a relatively insane amount of memory bandwidth. |
|
|
|
| ▲ | ricardobeat 5 hours ago | parent | prev | next [-] |
| What kind of coding tasks would you expect to hit that limit? In my setup, on a very large codebase, it takes each agent 3-4 minutes at minimum to go past 100k tokens. (note it's 150k uncached tokens, the total limit is 450k/min) |
| |
| ▲ | nostrebored 4 hours ago | parent [-] | | in my last tests with cerebras for coding tasks, most large tasks or anything greenfield would hit token limits. note that smaller models and the gpt-oss-120b style models they used to run are very prone to overthinking, so individual turns may be 3-10k tokens of just thinking + input + output. i don't think it's quite apples-to-apples to compare to a frontier model or even a k3. the odds of success (file compiles? read the right context?) are lower and thinking is longer. |
|
|
| ▲ | collin 5 hours ago | parent | prev | next [-] |
| This was my experience a year ago on some other model they could run super fast. Routine coding tasks would hit the per-minute token limits. Just the math there... 150k TPM... and 15k TPS means... you can run for 10 seconds every minute? The basic math boggles the mind. |
| |
| ▲ | baegi 5 hours ago | parent [-] | | Not sure how the rate limiting works, but it's 1.5k TPS, not 15k, so you could run it for 100s/min, which seems good enough to me | | |
| ▲ | nostrebored 5 hours ago | parent | next [-] | | iirc input (uncached) goes towards the limit as well | | | |
| ▲ | fc417fc802 4 hours ago | parent | prev [-] | | It seems you forgot to account for the fact that cerebras uses a baker's minute which is 144 seconds instead of 60. (Seriously though what's the supposed issue here?) | | |
| ▲ | RussianCow 4 hours ago | parent [-] | | The issue is that all input (including context) counts towards that limit. So 10 requests with 50k of context will blow through the limit, even if little to no output was generated, which is incredibly easy to do with agentic workloads. |
|
|
|
|
| ▲ | 0xbadcafebee 5 hours ago | parent | prev | next [-] |
| Yeah, their public service isn't a serious/competitive offering. They don't have the capacity to serve all the customers who might want to use them at that speed. The public service exists so they get some users on OpenRouter, and that shows them as #1 on speed, which proves their tech is very fast, which gets them billions in hardware sales/licensing. If you have big enough pockets they can probably dedicate capacity to you. But for reliably fast small models you might want to rent some GPUs. |
|
| ▲ | 5 hours ago | parent | prev [-] |
| [deleted] |