Remix.run Logo
saejox a day ago

This is a project i wanted to implement for a long time. It regularly benchmarks cloud hosted models with private benchmarks. Not just openai & anthropic, popular openrouter models too.

Tests their intelligence, not their diligence.

Sadly i cant think of a way to monetize the service. Also if it ever gets famous enough labs would try to game the system, it would be cat&mouse game that i am not willing to waste time on without any monetary gain.

arcanemachiner 21 hours ago | parent | next [-]

The only revenue model I for this is ads (like AI Stupid Level[0]). Or as a loss leader to get eyeballs to your service (like Margin Lab[1]).

EDIT: I forgot (and am shocked) that HN still doesn't seem to support Markdown-style links.

[0] https://aistupidlevel.info/

[1] https://marginlab.ai/trackers/claude-code/

mox1 21 hours ago | parent [-]

I mean I think if this is done well, lots of companies would pay for access to that data. Think like Enterprise subscriptions.

Its similar to other data services I see around my F500 company.

adrianco 21 hours ago | parent | prev [-]

I built GitHub.com/adrianco/retort to do this. It’s runs lots of experiments and you can contribute results if you have some spare tokens. You can add your own tests, and it runs Claude, Codex, Gemini, Hermes for local models.