| ▲ | ac29 3 hours ago | |||||||
Not sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last place | ||||||||
| ▲ | svachalek 3 hours ago | parent | next [-] | |||||||
Fable got heavily beaten down by its refusals, which is not too surprising; although a couple of problems got refused for reasons I can't even imagine and the page doesn't quote the refusal. Some of the other failures like the colicky baby one are also probably soft refusals, it's not clear what the grading criteria are but I'm guessing it got docked for not going anywhere near a possible diagnosis. | ||||||||
| ||||||||
| ▲ | poincareball 2 hours ago | parent | prev [-] | |||||||
I've been not just unimpressed by Fable, but actively find it to generate negative value. It hallucinates more, and in more destructive ways, than other models I've worked with and generates truly atrocious jargon and bizarre inhuman explanations that end up cluttering things. The code it writes is terrible too. Overly complex with a lot of technical debt. | ||||||||
| ||||||||