| ▲ | j_maffe 3 hours ago | ||||||||||||||||
Anyone has a link to a report of its capabilities? I can't find a reliable source. | |||||||||||||||||
| ▲ | vblanco 3 hours ago | parent | next [-] | ||||||||||||||||
completely vibes based, but ive been using it to port Mindustry game from Java to C# with agents, and its been working for 50 hours (its 15-20 tks so super slow inference). Its done a fantastic work and its almost finished now. Better results than deepseek flash and gpt luna by a mile on this kind of long term work. Less good than gpt sol or opus. We dont know the param count but my guess is 200-300 range. | |||||||||||||||||
| |||||||||||||||||
| ▲ | daveyoung 3 hours ago | parent | prev | next [-] | ||||||||||||||||
likely a distilled glm 5.3 that will punch within 20% of that at 2-3x less size. you'll find that capability is typically very jagged on models that are distilled | |||||||||||||||||
| ▲ | kristofferR an hour ago | parent | prev [-] | ||||||||||||||||
63% at DeepSWE. https://x.com/davis7/status/2091285712566140986 Wenghi is behind DeepSWE, one of the best benchmarks. | |||||||||||||||||