| ▲ | Needle: The benchmark your search engine can't memorize(keenable.ai) | |||||||||||||
| 32 points by matt4711 2 days ago | 12 comments | ||||||||||||||
| ▲ | MarkusQ 2 days ago | parent | next [-] | |||||||||||||
I wonder if search engines linked to human-use-case engines (e.g. google/bing) start at a disadvantage because they have been historically incentivized to break themselves to support their business models? It seems reasonable to suppose that "good at selling ads" ≠ "good at finding results". | ||||||||||||||
| ▲ | terno 2 days ago | parent | prev | next [-] | |||||||||||||
do you somehow control how non-trivial the queries are? The LLM generates them, right? what if every engine returns garbage, or on the other hand, handles them too well? building a benchmark like this in a genuinely fair way seems extremely hard to me. I’m very curious about the details, of course within what you can share. | ||||||||||||||
| ||||||||||||||
| ▲ | matt4711 2 days ago | parent | prev | next [-] | |||||||||||||
One of the authors here. We have been seeing lots of benchmaxxing and leakage in standard web search benchmarks such as BrowseComp. We developed this live benchmark with daily/hourly sampled fresh queries matching real agentic search traffic to estimate actual search performance of different AI search providers. | ||||||||||||||
| ||||||||||||||
| ▲ | daft_pink 2 days ago | parent | prev | next [-] | |||||||||||||
so is this an independent search engine benchmark or a blog post from a search engine provider showing their search engine at the top of the benchmark? i'm a little bit confused as at first when I was reading it i thought it was a search engine benchmark, but it seems that keenable is at the top which i assume is related to the web domain owner? i've never heard of keeenable | ||||||||||||||
| ▲ | Alexwortega 2 days ago | parent | prev | next [-] | |||||||||||||
Do you think it will be possible to train on this bench? | ||||||||||||||
| ||||||||||||||
| ▲ | mpalmer 2 days ago | parent | prev | next [-] | |||||||||||||
The blog post appears to get confused and devotes its entire second half to pitching Keenable itself. If the idea is to build credibility for the new benchmark, this maybe was not the best choice.
Besides the clear AI smell, this nonsensical claim also plainly contradicts the methodology's key evaluation claim that the quality of an engine's results should be measured against how much it overlaps with the reranked aggregate of the other engines. The benchmark thus seemingly values an engine's ability to "answer unanswerable questions" at zero.
Yeah? Care to cite anything for that? | ||||||||||||||
| ||||||||||||||
| ▲ | andrewww95 a day ago | parent | prev [-] | |||||||||||||
[flagged] | ||||||||||||||