| ▲ | ricardobeat an hour ago | |
Everyone is doing this to emulate Jev, but... I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast. | ||
| ▲ | prometheus1992 an hour ago | parent [-] | |
was the answer correct? i have tested jev for my use cases and its horrendously wrong, but then the follow up from jev's team is "oh, you need to boil the question down further". it's a spiral of how much do you wanna dumb down the ask so that it answers it correctly. i'll pass for now. also, 30k input tokens is a lot. | ||