Remix.run Logo
▲ hansvm 3 hours ago

The best models have O(1k tokens per second). A billion tokens, sequentially, takes 1-2 weeks. 500B takes 500x that ... not fast. The problem might not have required 500B tokens, but if it needed anything within a couple orders of magnitude then something like the given approach (or anything else yielding equivalent results in exchange for parallelism) was mandatory.

▲speedstyle 3 hours ago | parent [-]

Yeah, I don't think the problem needed billions of tokens

▲hansvm 2 hours ago | parent [-]

That's plausibly true. I've definitely seen absurd token costs abused and wasted. I've also seen simple problems require absurd token counts regardless of prompt quality. Which factors made this problem require 1000x fewer tokens than they used?