Remix.run Logo
moffkalast an hour ago

Everything old is new again, huh. I remember people doing this back in the early llama days, restricting grammars to yes and no tokens or 0 and 1 and then classifying questions. Usually it was rather ass in terms of performance cause no model is tuned to reply that way and it was WAY out of distribution, and yet then it got turned into the main way to run multiple choice benchmarks, and then everyone benchmaxxed it. Doesn't the normal MMLU/Pro also just do the same thing, restrict the output to one token, top-k=1, and it has to be one of the choice letters?

I think the real difference Jev makes is the fast parallel decode, it just seems rather bizzare how that works.