Remix.run Logo
johnmlussier 2 hours ago

I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.

verdverm 29 minutes ago | parent | next [-]

Interestingly, MiMo was trained across multiple harnesses and it improved capabilities

https://mimo.mi.com/docs/en-US/news/latest/v2-6

stogot an hour ago | parent | prev | next [-]

I’m the opposite. I want one open source harness to rule them all

Cost efficiency is a plus

sanderjd 27 minutes ago | parent [-]

Right! I had the exact opposite view when reading this comment. The good timeline is where one of (or perhaps a small number of) the open source harnesses becomes so dominant that the models compete to have the model that is the best trained to work with that harness. (I'm hoping this would be Pi, because it's my favorite, but mostly I just want it to be some model-agnostic open source harness that wins.)

avaer 2 hours ago | parent | prev [-]

This matters less as models get better and everyone settles on the same overall harness architectures. The model matters more than the harness anyway.

The bigger issue is that the use cases and harnesses for models is infinite, which is hard to compress into benchmark numbers that actually apply to you.

Everyone is benchmaxxing, desperate to sell, and almost nobody except the labs is doing actual science on the results, so harnesses tend to be chosen on voodoo and hunches, like which company made it. There isn't necessarily a good alternative though, bearing the cost of being a harness researcher is probably not many people's goal.

altcognito an hour ago | parent [-]

Agreed, especially since the more frontier models are able to accomplish in a vacuum, the more people will trust them. That being said, tool use is still really important for pulling in the right information.