Remix.run Logo
hellohello2 an hour ago

Respectfully, the linked website reads exactly like the Claude artifacts I read all day, i.e., it is low effort. I do not mind AI generated writing at all, but I do mind bad writing.

Further, you ignored the actual contents of my comments, to latch onto a superficial aspect. Please tell me: why do these results contradict existing attempts at benchmarking LLMs, which were designed with considerably more effort? Because the website certainly doesn't explain why in a way that is human-readable, which is why I asked.

EDIT: as explained by another commenter below, its because Fable refused to perform some of the tasks.