Remix.run Logo
frangonf 2 days ago

For evaluating subs you have to put the money and try by yourself to get your own measurement, since the subs quotas and new models are always changing and overall info is unreliable.

20$ of Codex/Google/Chinese gives you some fair amount of usage to test them, and Opencode Go for 10$ lets you try a good amount of models with good quota. I don't use Openrouter because it gets more expensive than the subs, but for testing, swapping and being completely independent, a proxy service is the best solution.

About benchmarks, I usually agree with DeepSWE. Looking at usage rankings of models in Openrouter is also a good signal.

For changing models locally I just use pi (used also opencode in the past) or if the sub does not allow login in external harnesses I just use whatever cli they have, last year there were differences but today they are all good enough for my needs and have basically the same features.