Remix.run Logo
▲ Zigurd 4 hours ago

If you're actually applying LLMs, all of the things around the LLM that adapt it to coding, for example, that enable it to use existing validation tools for code, and enable it to diagnose and fix tool chain issues that aren't directly coding problems, are what makes the difference between a model that that scores a little higher on a coding benchmark and a model that's useful in a particular code base on a particular platform.

Are there any use cases that have enabled one customer of a frontier LLM to outperform a competitor using a different frontier LLM? Or is this why we are seeing confected points of comparison like solving challenge problems in mathematics?