Remix.run Logo
ramon156 4 hours ago

But the evidence is not there...

pennomi 4 hours ago | parent [-]

Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.

simianwords 3 hours ago | parent [-]

They gesture at not using benchmarks for some reason...

meric_ 3 hours ago | parent [-]

https://typesafe.ai/blog/antibenchmaxxing

But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do