Remix.run Logo
epolanski 6 hours ago

I was writing a benchmark for my own harness, and DS4 flash answers as well as Fable 5 on any query.

The specific agent is focused on getting precise and on point answers about a codebase.

The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.

The benchmark included more than 50 questions or different difficulty.

But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.

Just to say that the quality of the harness is as important as agents intelligence.