Remix.run Logo
wild_egg 4 hours ago

It's a limit on input tokens. So that's 3 50k requests per minute. At Cerebras speeds, that's about 5 seconds of usage per minute.

I was very excited last year for their coding plan but seeing a burst of requests pulse and then sitting there watching the cooldown reset is really not a great time.

Even though each individual request was fast, the sessions were only maybe 10% faster on wall clock time since there was so much waiting time.

kristjansson 6 minutes ago | parent | next [-]

They made the coding plan a bit better toward the end, but it was pretty tough to use throughout.

Seems like an Amdahl’s law of inference economics? there’s so much compute relative to SRAM on the chip and shoreline bandwidth onto the chip that caching buys ~nothing? The contended resource is SRAM and a given token of context needs just as much as another.

amelius 4 hours ago | parent | prev [-]

Can't you do something with multiple accounts?

jychang 43 minutes ago | parent | next [-]

You would lose caching (if they cache)

sandworm101 3 hours ago | parent | prev [-]

Or just buy a 5060. This will run on most any 16gb card. Slower for sure but far cheaper than another subscription.

ma2kx 31 minutes ago | parent | next [-]

Thats not the point if you choose Cerebras as provider.

embedding-shape 2 hours ago | parent | prev [-]

Or buy a raspberry pi with a SSD, about the same difference, if you're giving up on the 1500 tokens/s anyways.