| ▲ | petesergeant 4 hours ago | |||||||||||||||||||||||||||||||||||||
Jev has taught me the same lesson three times over now. When it first came out, I thought "this weekend, I'll do a little open-source Jev based on single-token prediction and the token logit output", but of course when it came to it, there were at least 5 that had already been done between me thinking that and getting around to it. So I wrote up[0] what other people had done, but wasn't happy with how weak the benchmarks were, but in the time between writing the first word and the last few, two excellent sets of benchmarks had been written, so I was able to incorporate those. I published the article, and one of the authors of one of the implementations commented that I'd beaten him to doing the write-up he'd wanted to. This morning I thought "huh, you could have some fun giving Jev a single letter or token at a time, turning it into a chatbot", but as the time of looking two people had already done this (and taken the gag further than I would have), and ... this is isn't either of the ones I'd found. I bet if you scratch the surface there already at leat 5. Time from idea to output has dropped off a fucking cliff. | ||||||||||||||||||||||||||||||||||||||
| ▲ | radarsat1 2 hours ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||
> I'll do a little open-source Jev based on single-token prediction and the token logit output How is this idea generally working out in comparison with Jev? I'm curious, from what I read so far it seems like Jev is still beating this kind of thing. But it's curious because it's not entirely clear why, from an architecture point of view for all we know that's exactly what they're doing. So it must come down to the quality of those logits, ie., model size and training details. It seems to me that what most of these single-token-prediction projects are missing is that Jev seems to be claiming they predict well-calibrated probabilities. This is an incredibly valuable thing that LLMs simply can't deliver unless they are trained specially for it. | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | jorl17 3 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
I had the same sort of thing going on ahah. But I was convinced some hoje must have already done it and I decided I’d research when I got home (I’m out today). Didn’t expect it to reach hacker news so soon, though!! | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||
| ▲ | moffkalast an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||
Everything old is new again, huh. I remember people doing this back in the early llama days, restricting grammars to yes and no tokens or 0 and 1 and then classifying questions. Usually it was rather ass in terms of performance cause no model is tuned to reply that way and it was WAY out of distribution, and yet then it got turned into the main way to run multiple choice benchmarks, and then everyone benchmaxxed it. Doesn't the normal MMLU/Pro also just do the same thing, restrict the output to one token, top-k=1, and it has to be one of the choice letters? I think the real difference Jev makes is the fast parallel decode, it just seems rather bizzare how that works. | ||||||||||||||||||||||||||||||||||||||
| ▲ | applfanboysbgon 3 hours ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||
Time from pointless idea to bad output, anyways. We aren't seeing good software, and now neat hobby project ideas are getting harvested pointlessly when the only purpose of those ideas was the fun and learning of doing. | ||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||