Remix.run Logo
travishcronin 3 hours ago

Interesting to see harness benchmarks for coding agents. The same problem exists for conversational and data agents but I dont see anyone benchmarking them yet. Seems like manual-spot checks are the norm