| ▲ | Beating GPT-5.6 Sol on retrieval with 100x cheaper open models(neon.com) | ||||||||||||||||||||||||||||
| 75 points by moonikakiss 2 hours ago | 13 comments | |||||||||||||||||||||||||||||
| ▲ | mrinterweb an hour ago | parent | next [-] | ||||||||||||||||||||||||||||
There is so much opportunity for purpose built models like this. Ideally a harness should spin up a subagent to offload to targeted models for specific tasks like this. I know this is not a novel idea. Claude code does some of this by handing off the "explore" agent work to haiku. I just love seeing that specialized LLMs are being developed. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | breadislove 12 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
On what do you guys test the model. Its very dubious that there is no common retrieval benchmark such as browsecomp plus or similar tested. And what metric do you report? | |||||||||||||||||||||||||||||
| ▲ | aliljet an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
There is a more serious question in here that's not being answered. How effective is the retrieval in finding buried needles in larger and larger haystacks. And there's a correlary question, how effective could you be in finding paired needles in that haystack where you need to hold a needle to unlock finding another needle. | |||||||||||||||||||||||||||||
| ▲ | JCharante 40 minutes ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
I have done my own testing and found that smaller models can beat their larger siblings on fact retrieval from documents. I haven’t investigated it in depth with a large enough dataset but my guess is that larger models overthink it while smaller ones just do it. I would like if they compared this with 5.6 Luna instead. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | ramon156 an hour ago | parent | prev | next [-] | ||||||||||||||||||||||||||||
Bit unrelated, I realized that z.ai gives you access to deepseek 4 flash. It's incredible how well it performs when given a detailed spec. I'm not sure I've seen a model one-shot like that, and I was already impressed by gemma 4's speed and efficiency. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | richwater 42 minutes ago | parent | prev [-] | ||||||||||||||||||||||||||||
One thing that plagues [insert current FAANG] is the large amount of corpus knowledge that is outdated/misleading or just plain wrong. I'm curious how this addresses that if it's deriving the reward function from the corpus itself. | |||||||||||||||||||||||||||||