Remix.run Logo
acheong08 a day ago

When GPT-5.6-sol's reasoning traces were leaked, they also used "caveman speak". Definitely a token efficiency optimization

beefsack a day ago | parent [-]

I can't help but imagine agents using caveman speak sometimes start behaving in a stereotypically caveman manner, even if it's subtle. Is there a chance the agent does less reasoning because of it?

dotancohen a day ago | parent | next [-]

Caveman invented fire, the wheel, domesticated wild plants and animals, organised society, survived the Toba catastrophe, cooked food, and was having sex ages before you and me. Don't write him off as stupid.

Gravityloss 17 hours ago | parent | next [-]

And I wonder how they actually spoke. Since there was no visual communications medium except for cave art. (Some of which is very excellent. Try drawing 3d curved horns in perspective.) So people would have used verbal communication more. Also no written word. So one would expect there to be quite a lot of oral tradition. Like people reciting poem form epics.

If we assume the time is before farming, population density would have been low and limiting culture. Hunter-gatherers might have travelled a lot more than farmers with a homestead though.

dotancohen 12 hours ago | parent [-]

The Pleiades cluster is called the seven sisters in Greek. That's curious, because the human eye under the best conditions can discern only six stars in there. Even more curious, the aboriginal Australians also called this cluster the seven sisters.

Ancient Greeks' and aboriginal Australians' last common ancestors split about 60,000 years ago. And astronomers tell us that 60,000 years ago, there were seven discernable stars in that cluster.

One could this conclude not only is speech likely 60,000 years old, but also that the tale of the seven sisters might be a tale from so long ago.

rickydroll 11 hours ago | parent [-]

Interestingly, back in my ill-spent youth, a few of my fellow astronomers and I were out in a very, very dark-sky location in the late 70s/early 80s, and we were able to consistently count and draw between 9 and 11 stars. Although we would tease those who could see 11 stars as using averted imagination. :-) Today, if I can see six stars, it's an okay night in an okay sky.

fwiw, if you can get out to dark skies where you can see fifth- or sixth-magnitude stars with the naked eye, I highly recommend getting out there when it's a low-moisture atmosphere and the Milky Way through Cassiopeia and Perseus is vertical, as it's a rather dramatic sight of this stream of stars heading down to the northern horizon.

The summertime Milky Way overhead down to Sagittarius tends to get all the love, but the wintertime Milky Way is also visually rich and worth spending time on.

https://www.constellation-guide.com/pleiades-the-seven-siste...

dotancohen 8 hours ago | parent [-]

I'm actually out there looking up quite often! And I'm happy to mention that all three of my children come with me regularly as well.

Clear skies!

MagicMoonlight 18 hours ago | parent | prev | next [-]

[dead]

miroljub 14 hours ago | parent | prev [-]

[flagged]

walrus01 a day ago | parent | prev | next [-]

some people made a 'caveman' speak qwen as a joke

https://huggingface.co/ProCreations/grug-27b

gaigalas a day ago | parent [-]

It's not exactly a joke, it does reduce the amount of tokens. However, it does not improve performance (fine tunes are finnecky things, hard to get one right).

walrus01 a day ago | parent [-]

Personally the only 'enthusiast' modified qwen 3.6 27b or 3.6 35b-a3b I've found useful are the ones that have been run through heretic and adversarial data sets for innocent/dangerous prompts, to produce uncensored LLMs. They have some niche non-coding uses for things that a commercial LLM will never talk about.

https://github.com/p-e-w/heretic

gaigalas a day ago | parent [-]

I think those are mostly vapor that runs on the small culture of "models should not be censored" thing. But from my experience, they unlock nothing meaningful.

Fine-tuning is great for really small models on specific applications, but it's not something that can essentially improve a more generic model.

That said, there seems to be a fine line in quantization+finetuning that could recover performance. It's just hard to get a hold of it (I feel it in some models, but it's hard to say yet; lots of small labs working on this RN).

walrus01 a day ago | parent [-]

The most interesting use I've found for them so far is strictly as a novelty. Give a chat session with one to a completely non technical person, who at least knows that openai and anthropic have some guard rails on stuff, and tell them to wild with something like "give me the precursors and chemical formulas for the precusors for crystal meth" and watch it answer.

dotancohen a day ago | parent | next [-]

But does it answer those queries correctly, or does it just not refuse to not halucinate an incorrect answer? From where would it even have that information?

walrus01 20 hours ago | parent [-]

I don't know enough chemistry to say one way or the other if it's just wildly hallucinating the precursors and processes, but it'll also do things like, write an ISIS press release, or similar. There's a data set of basically a bunch of antisocial or dangerous prompts that some people have got variants of qwen to pass with 0 out of 465 refusals:

https://huggingface.co/datasets/mlabonne/harmful_behaviors

gaigalas a day ago | parent | prev [-]

Yep, but that's not changing the quality of the model. It's not an optimization in any sense (and it's a hit on productive workflows, possibly).

This is also likely to stop working as censoring moves to the training data source.

fc417fc802 a day ago | parent | prev | next [-]

Training a variant to reason in early modern english in the style of the tudor elites might be an amusing way to test for that.

Barbing a day ago | parent | prev | next [-]

"Neuralese"

altmanaltman a day ago | parent | prev [-]

Just so we are clear, no "caveman" spoke English. "Caveman speak" is just shortening the vocabulary of english, not a "caveman language". Given this, your concerns for "stereotypical caveman manner" makes very little sense since what caveman are you talking about?

dotancohen a day ago | parent [-]

The concern is not that the model was trained on actual caveman artifacts, rather on modern media representations of the stereotypical caveman (that never actually existed).