Remix.run Logo
juiceland a day ago

You don’t need a trained attorney to schematize data. The LLMs are used to make the data easier to understand and manipulate.

cromka a day ago | parent [-]

They'll also hallucinate and change meaning in the process of extraction and "schematization"

margalabargala a day ago | parent | next [-]

Not necessarily. Depends how you use it.

"Write a python script that breaks down this PDF by X feature" would not hallucinate anything in the PDF. Certainly you could trivially double check that all text in the extracted JSON was in the text layer of the PDF.

jeffbee a day ago | parent | next [-]

How much experience do you have with LLMs exactly? It would be consistent with my experience if Claude stuck in a line of python that just emits a JSON literal with no justification, potentially buried in a large program where an untrained person might not notice it. I don't even trust them if the output consists of structured data paired with source images from the PDF, because I've experienced LLMs fabricating the source rectangles to match the output. I only use tools like this by asking for programs, because as you note LLMs are good at that, and the verification process consists of tool calls to legitimate PDF manipulation tools so I have some confidence everything is above board. Even then I only do this for hobbies, not anything that matters.

terminalcommand a day ago | parent [-]

Lawyer here. I used to trust Claude as hallucinations are near non-existent now. However for large volume tasks such as due diligence exercises, they still happen. We also tried Legora's tabular review, there were also numerous halucinated provisions in our due diligence exercise.

egorfine 17 hours ago | parent | next [-]

Is it possible to catch those hallucinations using another LLM with a strong fact checker prompt with sources provided in output for human validation?

rayiner a day ago | parent | prev [-]

Junior associates hallucinate too...

NateEag a day ago | parent | next [-]

And when they do, you can train them or fire them, and they learn not to do it.

LLMs change not a whit, and there's no one to take responsibility for the failure (and thus no way to fix it).

As the new variation on the old theme has it, "A computer can never be held accountable, and so very many people are trying to get them make management decisions."

IanCal a day ago | parent | next [-]

You can’t train people to never make a mistake, particularly when doing highly repetitive work like this. You must build your systems to account for that regardless.

NateEag 15 hours ago | parent | next [-]

For sure.

But they do learn and improve.

The models don't (yet).

hollerith 15 hours ago | parent [-]

The models improve in the sense that GPT 5.6 succeeds at things GPT 5.5 fails at.

It might be that the models have been improving in this sense faster than a human child improves.

enraged_camel 15 hours ago | parent | prev [-]

Yes, exactly. Humans are non-deterministic as well, just in different ways. A tired human can make all sorts of errors for example, regardless of how much training they've had.

juiceland 16 hours ago | parent | prev [-]

> LLMs change not a whit, and there's no one to take responsibility for the failure (and thus no way to fix it).

LLM output is nondeterministic and humans take responsibility for the failure the same way they take responsibility of a photocopy is too dark.

NateEag 15 hours ago | parent [-]

> humans take responsibility for the failure the same way they take responsibility of a photocopy is too dark.

You mean, they notice it's too dark right after making it, change the settings, do it again, and give you the good copy?

Because yes, that's my experience of humans.

podocarp 21 hours ago | parent | prev [-]

You can scold juniors and they will learn. You can't scold Claude.

stevesimmons 21 hours ago | parent | next [-]

Surely the rate of improvement in new LLM models is the equivalent mechanism?

eru 20 hours ago | parent | prev [-]

You can scold Claude. Just doesn't make a difference.

chrisjj 20 hours ago | parent [-]

... until you hit Claude's risable "model welfare" protection.

chrisjj 20 hours ago | parent | prev [-]

That's a recipe for disaster in my experience. I tried it (with Claude) on a simple tabular bank statement PDF, and it transposed two amounts, placinh each against the other's description. And the bot assured me the result was cotrect. The chance of a human checker catching such corruption is low.

margalabargala 19 hours ago | parent [-]

Interesting. Did that PDF have a text layer or did you ask Claude to OCR it? If the latter I'm not surprised at all.

chrisjj 16 hours ago | parent [-]

Text layer.

ferngodfather a day ago | parent | prev | next [-]

Doing similar-ish things with Claude, it's helpful to have something to ground it.

For instance, if you can say:

"Refer to the database schema in x.sql as your source of truth for the database structure we want to import into*. Do not invent data, tables or columns that do not exist. Carefully match all output against this database schema and do not create output that doesn't exist if it does not match the schema, simply skip it."

You will end up with a far better result in my experience.

Gotta treat it like a child.

chrisjj 20 hours ago | parent [-]

> Gotta treat it like a child.

"Sorry for that, Your Honour, but we gave that case prep work to a child."

ivraatiems a day ago | parent | prev | next [-]

They do, sometimes. That's why the review still has to happen.

But now it's comparing already filled columns on a spreadsheet, not copy-pasting every single thing from an (often uncopyable) PDF.

chrisjj 20 hours ago | parent [-]

> But now it's comparing already filled columns on a spreadsheet

... with a PDF, right?

> not copy-pasting every single thing from an (often uncopyable) PDF

Obviously the PDF is copyable, else your bot would not be copying it.

ivraatiems 13 hours ago | parent [-]

"Copyable" in this case means you can highlight and copy paste text instead of manually retyping it.

juiceland 16 hours ago | parent | prev [-]

But it gets it right like 99% of the time so human attention can be put towards catching the 1%, not entering data from one table to another and then catching that human’s mistakes.