| ▲ | kflansburg 13 hours ago | |
Some users on X [1] were discussing the use of DiffusionGemma for Jev-like use cases. Similar to Jev, diffusion models generate output fully in parallel. We achieve similar estimation of relative confidence by looking at token log probabilities. Finally, by disabling reasoning we achieve low output tokens and similar latencies. This demo lets you compare the output to Jev for some canned prompts, or see the DiffusionGemma output for your own prompts. [1] Original X post with vLLM PR https://x.com/mmastrac/status/2100626193943052784 | ||