Remix.run Logo
Humanising LLM Outputs Is Dumb(kuber.studio)
68 points by kuberwastaken 7 hours ago | 36 comments
Xcelerate 3 hours ago | parent | next [-]

You ever read a work of literature with such flowery language that right after you've read a paragraph, you pause and realize you have no clue what you actually read, only to read the paragraph maybe a second or third time and have your mind space out again and again on each successive attempt?

Yeah, for me, that's what parsing huge volumes of LLM-produced text like "direct model calls as replaceable semantic workers" does to my brain. Maybe others don't really have this issue, but after any long output, I prompt the agent "Go back and decompress any LLM-speak in light of the higher level task goals. Eliminate deictic language."

The revised output documents are solely for my personal usage to expedite understanding. The LLMs can slowly converge on their own language for all I care; I retain raw agent output for future agent usage (to avoid the "lossy" problem the author mentions), but that doesn't eliminate the need for some intermediate translation I can use to actually help get my work done instead of spending hours attempting to understand what a "load-bearing pinned gate" is.

whythismatters 2 hours ago | parent | next [-]

The effect you describe reminds me of reading Edward W. Said's "Orientalism" when I was younger. Fable suddenly started with this kind of lingo, iirc, and Opus 5 sounds exactly the same. Tin foil: it's ultimately a vendor lock-in strategy, you'll get the best results with agents from the same tribe, others will trip over the mountain of idiosyncratic metaphors.

Terr_ an hour ago | parent [-]

> you'll get the best results with agents from the same tribe, others will trip over the mountain of idiosyncratic metaphors.

Good point, there's an anti-competitive incentive, and self-bias in models is a mechanism to do it.

andai 5 minutes ago | parent | prev [-]

Well, now I had to ask an LLM to give me examples of what "deictic language" means...

7402 2 hours ago | parent | prev | next [-]

I don't like it when the LLM tries to be my friend. My general prompt (a work in progress) is this. I wonder what other people use.

"Answer impersonally, objectively and analytically, without undue friendliness or enthusiasm. Use an engineering style response: concise, factual, and complete. Do not speak in the first person. Do not promote engagement or an emotional connection. Do not use emojis."

prymitive 2 hours ago | parent | next [-]

+1 it’s a tool

It’s not perfect, it has shortcomings, it sometimes produces bogus outputs. All of that is fine for a tool, it’s not fine when it pretends it’s a conscious being, because errors start to feel like lies and it becomes a bit too personal.

MSFT_Edging an hour ago | parent [-]

People want it to be Data from Star Trek, when it really should be the ship's computer. I want to tell it to run a simulation accurately, create some solved tool, etc.

I don't think there's any correction that can return LLMs to a purely tool-space. Too many AI boyfriend/girlfriends.

cortesoft an hour ago | parent [-]

Shouldn't it be whatever the user wants? If they want the ship's computer, it should be that, if they want Data, it should be that.

MarkusQ an hour ago | parent [-]

You can want your e-scooter to be a jet ski, but you'll wind up having issues when you try to use it as one. LLMs are _really good_ at pretending to be something they aren't, but not always so good at being that thing, so you should be careful what you ask for.

doctoboggan 43 minutes ago | parent | prev [-]

Yes, this really ought to be trained in (or at least RLHF'd in) but that would hurt engagement numbers so the opposite is done instead.

These are tools and it would behoove us all to keep that top of mind. Dangerous tools that are not your friend (but are useful as tools nonetheless)

Animats an hour ago | parent | prev | next [-]

Well, what do you expect? LLMs are trained on blithering, mostly from web sites. So you get blithering out.

There's an important point in the article, that forcing a style onto an LLM is lossy. Although he doesn't seem to mention it, forcing a style may result in the insertion of new blithering, possibly made up as a hallucination.

efficax a few seconds ago | parent | next [-]

is it lossy though? That didn't make sense to me. You can tell it to use Simplified Technical Language and also still have it give you all the detail. it's just another piece of the prompt that produces the output. it's not like there's "pure" llm output and then "lossy" output guided by a prompt.

mjburgess 25 minutes ago | parent | prev [-]

I think that was a good enough explanation for gpt3.5 -- these days, labs are extremely capable of post-training phases that eclipse that kind of training phase -- and hence of choosing whatever style or tone they wish.

eg., OpenAI has gone a long way to making reasoning token-efficient by having reasoning piovot off terse langauge -- whereas anthropic appears to be doing the opposite.

conguy 3 minutes ago | parent | prev | next [-]

Honest Short Fall -- <insert 30 lines of useless shit>.

If the author wants to read slop for hours, be my guest. Make it lossy, my job is not to read mimetic feelings, it's to make sure implementations get implemented.

firefoxd an hour ago | parent | prev | next [-]

And on the "input" side, one thing that used to improve google search result was to write like you are talking to a robot. "Ruby on rails http header set function". As opposed to "how do I set header in ruby?" Then you have to page through results until you find something specific to rails.

Now, the second example is the only thing that works. Power users have lost their powers with AI overview.

skydhash 17 minutes ago | parent [-]

I still use the first strategy (with DDG) and it still works great. But for technologies I work often, I just take a bit of time to familiarize with the site's structure and maybe bookmarks a few pages.

stillpointlab 2 hours ago | parent | prev | next [-]

One thing that continues to give me pause is Fable's insistence on using my fist name in messages and docs. Like, I'll explain what I want to the AI and ask it to write out a spec or brief and Fable says "Jamie wants me to ...". It just feels different and unprofessional. If I was at a job and a PM asked me to write up a task spec I wouldn't say "Harold wants to add <feature> ...". And since I am the one reading the output it is also superfluous and almost feels like talking about myself in third person. But there is almost a kind of glee in the way it uses my name, like a student using their teachers first name when the custom is to use Mr/Mrs.

scubbo 37 minutes ago | parent | next [-]

Fair perspective, though I actually prefer this for two reasons: * When it's proposing responses for me to choose between, a description like "I close the PR and you make a followup" is ambiguous - is "I" there "the entity making the proposition (the LLM)" or "the entity making the choice (me)". * I have a line in my `AGENTS.md` specifically instructing it to call me by my name; if it stops doing so, that's a telltale that context-bloat is pushing out other instructions.

zamadatix an hour ago | parent | prev [-]

I usually leave memory/connections turned off. Remembering/finding out what my name would be is not really something I want to waste context or tokens in, let alone any if the other things it tries to assume I'd like it to remember/find.

mdp2021 2 hours ago | parent | prev | next [-]

Suppose you had an LLM (NN) producing its default output from an input (a generally optimal for-most-cases role-sys, and any role-user), and then you wanted to have that output reformatted in some style (e.g. "In iambic pentameter" | "haiku" | "eli5" | "in the style of Feynman" | "bulleted like Axios" ...). How would you keep the internal NN workings that were basis for the original output, and use them to get a rewritten version (instead of placing the original query and output in the context and ask to rewrite it)?

In other words, is there a way to keep the internal process intact up to the point of the formulation - and have only that vary.

warmwaffles 10 minutes ago | parent | prev | next [-]

Humanizing the LLM output is a hedge against agents hitting a wall and someone having to reason through it by hand.

raver1975 40 minutes ago | parent | prev | next [-]

Someone is finally making good sense up in here.

mikaeluman 3 hours ago | parent | prev | next [-]

I don't get it. The skills and instruction try to make the answer more machine like on purpose.

Not humanising it...

People want the terse, matter-of-fact output. Not the conversational chatty verbose and bloated nonsense with gray words and jargon and terms like "blast radius"

alansaber 3 hours ago | parent [-]

The article lost me when it implied that verbose drivel is actually intrinsically superior rather than a way to hedge bets

Havoc 3 hours ago | parent | prev | next [-]

> The problem is that these instructions are not applied after the model has finished doing the work

Seems like something fixable with a simple two step process. Ask it the thing. Then ask it to summarise the answer in simpler terms. More tokens and time aside that would check both boxes

kuberwastaken 3 hours ago | parent | next [-]

pretty much what I do, better yet ask it to boil it down in visuals in a simple webpage if it's a very large project

StyloBill 3 hours ago | parent | prev [-]

Should be a harness feature actually.

agenticworldcup 2 hours ago | parent | prev | next [-]

Yes, but the sycophantic responses are the worst.

thenthenthen 2 hours ago | parent | prev | next [-]

I have been using chatgpt for a while and its awkward, yesterday i tried gemini and its like a breath of fresh air.

51Cards an hour ago | parent [-]

If you're finding Gemini "clean and straightforward" give it awhile. I felt the same thing too when I switched until I realized that it just hadn't formed a model for my communications yet. After awhile it became just as flowery as ChatGPT did. I had to tone both down with saved preferences.

mthoms 31 minutes ago | parent | prev | next [-]

There's some good points made here about losing fidelity by over-simplification. As an ADHD sufferer, I'd take this piece much more seriously if the title wasn't so belittling.

I don't think it's wise to take communication advice from someone so helplessly juvenile (and attention seeking) in their own communication attempts.

alansaber 3 hours ago | parent | prev | next [-]

Not sure what happened in the blog, but I quite enjoyed the mindmap in the right panel

kuberwastaken 3 hours ago | parent [-]

Thanks, I guess haha :P

slowmovintarget 2 hours ago | parent | prev | next [-]

At first I read the title and mistook it for an argument against the anthropomorphism of LLMs. It isn't. Instead it's a take on suggesting that maybe it's a bad idea to dumb down the self-chatter in the process. A reasonable take.

It isn't deliberately unhinged like Steve Yegge's take: https://yegge.ai/essays/model-welfare/ In Steve's essay he starts with the assertion that agents are sentient... Whether or not that's true isn't really relevant, as his agent-flavored version of Pascal's wager actually holds water, especially for Anthropic models, as their system prompts already push the model in that direction, and it is better to work with them than try to prompt against the tide.

spwa4 3 hours ago | parent | prev [-]

TLDR: This is an argument to get LLMs to answer in short, even code-like statements because you can exchange information quicker with an LLM that way. Cool!

sandblast 35 minutes ago | parent [-]

It is not. You got it backwards.