| ▲ | wild_egg 4 hours ago | ||||||||||||||||||||||||||||
It's a limit on input tokens. So that's 3 50k requests per minute. At Cerebras speeds, that's about 5 seconds of usage per minute. I was very excited last year for their coding plan but seeing a burst of requests pulse and then sitting there watching the cooldown reset is really not a great time. Even though each individual request was fast, the sessions were only maybe 10% faster on wall clock time since there was so much waiting time. | |||||||||||||||||||||||||||||
| ▲ | kristjansson 6 minutes ago | parent | next [-] | ||||||||||||||||||||||||||||
They made the coding plan a bit better toward the end, but it was pretty tough to use throughout. Seems like an Amdahl’s law of inference economics? there’s so much compute relative to SRAM on the chip and shoreline bandwidth onto the chip that caching buys ~nothing? The contended resource is SRAM and a given token of context needs just as much as another. | |||||||||||||||||||||||||||||
| ▲ | amelius 4 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||
Can't you do something with multiple accounts? | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||