Remix.run Logo
▲ esperent 4 hours ago

> A CEO managing an army of agents is possibly on the hook if anything goes wrong - for example, if one of the agents does tax fraud or hacks a competitor

This is correct under present day thinking.

But imagine the accounting AI of twenty years hence. If it can reliably do the books, reduce and prevent fraud, significantly better than any human ever could, why would you still need the human in that loop?

In the present day it's a legal requirement, not to mention a practical requirement given present AI capabilities. That doesn't mean it still will be in the future.

To put it another way, if the human is on the hook for the agent, but the agent is proven by a decade+ of statistical evidence to be way less mistake/fraud prone than the human, what's the point of the human there?

▲socializer 3 hours ago | parent | next [-]

I'll give you that the capabilities of LLMs are improving rapidly. Their safety, however, seems to be stuck in mid-2023. They're still easy to dupe and prone to cheating to solve problems. Prompt injection is still a thing, and it's still something we need to paper over with input and output classifiers and other hacks external to the LLM.

What we're seeing so far is consistent with the training data being the upper bound for capabilities. They get better at recall / synthesis / reasoning over the corpus, but they don't, for example, acquire trans-human ethics; they're at best as ethical as we are, except not grounded by the fear of consequences. A perfectly-behaved, perfectly-moral LLM is not a given in 20 years, not unless your position is that there's room for unbounded, exponential self-improvement without any loss of fidelity. In that case, we'll probably have problems more pressing than the outlook for accounting jobs.

▲bot403 4 hours ago | parent | prev [-]

I'm hearing "better" and "less" but not zero and none. The more interesting question is what happens when that rouge/hallucinating agent does commit that statistically unlikely fraud. Who is responsible then? Or is it now no low we just consider it an "accident", pay a small penalty, and move on?