Remix.run Logo
dakolli 4 hours ago

LLM prose has degraded with each model update, the models are now being RL'd into wall of texts that only makes sense to other agents. I think its going to become more obvious as we move forward that is llm generated because labs seem to only be focused on tool-calling/agentic-coding environments, as that's the only thing that drives revenue.

There may come players who focus on models that are good at writing for technical writing/docs , copyrighting ect but I think people will lean towards not using them and will rather have the "human touch" for the things that directly impact brand perception.

Keep in mind, every single AI company that is selling the idea that you don't need to hire designers and web design is "solved" have $100k retainer designers crafting their landing pages.

spijdar 4 hours ago | parent [-]

Yeah, this tracks with my experiments. About every 6 months or so for the past year-and-a-half I've tried using the "frontier" LLMs to as-near-as-possible autonomously write novel length stories, because I find it fascinating.

When I had Gemini 2.5 write a novel, it wasn't really objectively "good" by any stretch of the imagination, but while the prose was very purple and full of cliches and, well, bad writing I guess, it still felt ... subjectively good, at least for what it was.

Last week I did a run with GPT-5.6, and wow. On the one hand, it managed to produce 110,000 words that were "shockingly" coherent. The model was able to maintain state and plot lines and background details extremely well, much better than older models.

But I just don't like the prose. I haven't really liked _any_ prose that GPT-5.6 produces. It's significantly better at "instruction following" and keeping track of things, but, wow.

> “The sequence is consistent with their voluntary choices.” Mara enlarged the uncertainty field rather than the result. “It does not prove what happened to anyone we can’t observe. It does not prove contact did this. And it does not turn the Shard into treatment.”

GPT-5.6 in particular becomes so fixated on certain ideas like "consent" and epistemology, that by the end of the narrative, the prose and dialogue are all just "agent speech", despite the prompt/harness specifying that it's a _novel_ with narrative prose and such.

Interestingly, the model itself produces an accurate critique of its own output:

> The draft has become a *consent-centered medical, legal, and logistical procedural*. The important drift is therefore not that many events were omitted. It is that the retained events now prove a different thesis.

Which begs the question of if it would do better with a couple rounds of output -> critique -> revision. But I think I've had enough LLM prose for a bit...

jedberg 2 hours ago | parent [-]

I'm curious if you've tried Claude. Subjectively, I've always preferred Claude's writing over ChatGPT and Gemini.