| ▲ | Reverse-engineered Jev-like model(github.com) | ||||||||||||||||
| 81 points by rochansinha 7 hours ago | 12 comments | |||||||||||||||||
| ▲ | steeve 5 hours ago | parent | next [-] | ||||||||||||||||
https://x.com/harshagundal/status/2100044305536889015?s=20 > They were building in stealth for 2 years, I was building in stealth for 2 hours… > Happy to open source Qwen-2.5-1B-RLCD, 5x faster on-device inference for JSON workloads that need to be type-safe. | |||||||||||||||||
| |||||||||||||||||
| ▲ | mmastrac 2 hours ago | parent | prev | next [-] | ||||||||||||||||
Any diffusion model is potentially a Jev in disguise: https://github.com/vllm-project/vllm/pull/57250 Runs ~0.2s per decision on my DGX Spark.
All incorrect answers are marked with low-P.It (DiffusionGemma with the Jev mode) can also solve an ASCII maze. | |||||||||||||||||
| ▲ | vrc an hour ago | parent | prev | next [-] | ||||||||||||||||
Out of curiosity and semi unrelated — why do so many of these projects with customized encoder-decoder setups use earlier Qwen versions like 2.5 and 3 and not the smallest 3.5? Purely the few 100m params, or something else in the latter’s arch or pretraining? | |||||||||||||||||
| |||||||||||||||||
| ▲ | rochansinha 7 hours ago | parent | prev | next [-] | ||||||||||||||||
Can play Doom too - https://x.com/vinnylarouge/status/2100281651930513460 | |||||||||||||||||
| |||||||||||||||||
| ▲ | tomrod 3 hours ago | parent | prev | next [-] | ||||||||||||||||
I like it! I suspect Jev may have more going on under the hood, but I like the idea of efficient universal transformers | |||||||||||||||||
| ▲ | _superposition_ 5 hours ago | parent | prev [-] | ||||||||||||||||
That was super quick. | |||||||||||||||||
| |||||||||||||||||