Remix.run Logo
▲ throwawayy6767 an hour ago

That's just cocktail party level of thinking. LLMs are getting better at math, code and logic (and marginally better at science and general knowledge) because these are domains that can be objectively verified and thus there is a potentially infinite supply of 'facts' to generate and train on. These are very powerful but ultimately very abstract domains. For everything else the messy real world and its physical bottlenecks gets in the way and there's little reason to expect progress to accelerate. It still takes months to get mice to reproduce and run experiments on, no matter how knowledgeable about biology the models have become.

If anything, in some domains frontier models have become worse - claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.

▲atleastoptimal 37 minutes ago | parent [-]

Is good writing verifiable? I don't think it is, but LLM's have been hillclimbing writing quality. That being said this is with the help of RLHF.

However there are many other domains which have verifiable rewards in the process of learning them, despite their overall impact not being verifiable. For example, the life sciences, an LLM could be given access to data about an organism, and then make predictions about how a drug or gene therapy will affect that organism. In economics, LLM's could create models of behavior, evaluate predictions over time and see how well those predictions match reality.

>claudisms and chatgpt idioms are making them notoriously bad at prose without a considerable amount of prompting and tweaking.

I think this is because a lot of people genuinely like the claudisms, even though a small minority of technical people don't.

▲wizee 15 minutes ago | parent | next [-]

The quality of natural language output from today’s LLMs seems no better and often worse than LLMs from 18 months ago, at least in my experience. It’s much harder to build a good RL loop for something subjective like good writing, compared to something with an objective and verifiable right or wrong answer, such as software or mathematics.

▲howunfortunate 3 minutes ago | parent [-]

If you gave me natural language from a 125 IQ human and a 155 IQ human I'm not sure I could judge which human is smarter from the writing quality.

And that's a massive intelligence gap!

It may just be that writing is mostly saturated, and the remaining perceived gap is not quality but some kind of "cultural fit".

▲throwawayy6767 16 minutes ago | parent | prev [-]

It's not at all clear that LLMs have gotten significantly (much less hillclimbing) better at writing since GPT4-ish. They did get better at oulipo-like challenges (writing under silly, verifiable constraints like Fable's viral 'facetiously' poem) but that's far from being all there is to good writing.

Anecdotally, creative writing communities do report drastic differences in quality between models (it is my understanding that Gemma 4 is especially praised) so there is something to it being verifiable but it's a far cry from being something that can be distilled into objective math/code-like evals. One could even say it comes down to vibes.

>For example, the life sciences, an LLM could be given access to data about an organism, and then make predictions about how a drug or gene therapy will affect that organism.

Yes, biotech companies are doing this right now but ultimately they are just predictions and you still need cold hard biological data to ground them and iterate on. That's slow and expensive, especially for the juicy fields (human biology, food and crops, clinical trials) and you can't just plug billions of VC capital into the pipeline and hope for RSI. Physical (gotta procure all those labs and their equipment), human (gotta hire and pay specialists to run experiments), biological (gotta wait for organisms to reproduce, drugs to take effect, crops to grow), regulatory (gotta convince agencies that your fancy new drugs are legit) bottlenecks get in the way.

And biology is one of the easier fields where a path to RSI (if not accelerating) is conceivable. Good luck iterating on macroeconomics data where there are no replicates and 'experiments' take literal years if not decades.

>I think this is because a lot of people genuinely like the claudisms, even though a small minority of technical people don't.

My uncharitable take is that the abtruse jargon makes people feel smart for understanding it and gives a sense of belonging (as a closed circle of initiates who understand LLM cant), just like rationalists love to repackage old or unsavoury ideas under new nerdy smart sounding names.