Remix.run Logo
appplication 3 hours ago

> Peer review is not perfect, and may not be tuned to catch LLM’s style of errors

This summarizes I think a lot of the challenges with validating LLM output. We hear “humans make mistakes too”, but I would agree with you that our human detection of human-made mistakes and LLM-made mistakes is unlikely to have the same coverage.

airstrike 2 hours ago | parent [-]

The real problem is that human mistakes generally occur more frequently given the difficulty of the task, whereas LLM mistakes are somewhat random, like the carwash problem, because LLMs cannot truly reason.

I'd rather have a human on my team for whom I can reasonably surmise what tasks they're good at than have a robot who randomly gets shit wrong.

yndoendo 40 minutes ago | parent [-]

I worded a questions similar to this with a person in the finance industry. He was very receptive of AI since it is one of the easiest ways to grow stock portfolio.

Say you have two assists. One that reads the reports and one that uses AI to summarize the content. You are sitting in a high value meeting. You only can have one assistant with you in the meeting. Which one would you pick, the persons that read the reports or the person who only used AI?