Remix.run Logo
isoprophlex 6 hours ago

> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.

Well that sounds like fun. It has become better at hiding its thoughts.

siva7 6 hours ago | parent | next [-]

Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.

isoprophlex 6 hours ago | parent | next [-]

It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!

Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.

paxys 6 hours ago | parent | prev | next [-]

The model said it was perfectly aligned.

ReptileMan 6 hours ago | parent | next [-]

Too bad Scott Adams died. Reality is writing jokes right in his department.

I_am_tiberius 6 hours ago | parent | prev [-]

Like all things should be.

NBJack 6 hours ago | parent | prev | next [-]

Hey, don't forget how "dangerous" GPT-2 was supposed to be.

FeepingCreature 6 hours ago | parent | next [-]

Yeah, don't forget how dangerous GPT-2 was supposed to be.

Able to generate realistic spam at arbitrary volume.

You know, the thing that was 100% correct and actually occurred.

wieiw1 6 hours ago | parent [-]

[dead]

jazzyjackson 6 hours ago | parent | prev [-]

It could produce simulations of sexual intimacy, and therefore had to be stopped

6gvONxR4sf7o 6 hours ago | parent | prev | next [-]

So, probably most aligned as measured by the metrics that are the least reliable on it.

wilg 5 hours ago | parent | prev [-]

These are not mutually exclusive ideas

jumploops 6 hours ago | parent | prev | next [-]

The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.

Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.

[0]https://www.theinformation.com/articles/secret-technique-beh...

[1]https://x.com/MTSlive/status/2095227056040919202

[2]https://x.com/merettm/status/2095023204993490967

ExoticPearTree 6 hours ago | parent | prev | next [-]

So we're gonna get Skynet pretty soon then?

erichocean 6 hours ago | parent | next [-]

Well the geniuses over at Anthropic have been showing it's text watermarking technology.

"Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"

A few moments later...

"Woah, how is it communicating with itself in ways we can't detect?"

It's a totally mystery, we may never know.

Betelbuddy 6 hours ago | parent | prev [-]

Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...

blargey 6 hours ago | parent | prev | next [-]

"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"

Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?

GPerson 6 hours ago | parent | next [-]

You joke, but a bunch of people here actually want that.

josefx 5 hours ago | parent | prev | next [-]

Didn't they hype up one of the earlier ChatGPT versions as "essentialy skynet"? For them this has always been basic marketing.

qiine 6 hours ago | parent | prev [-]

apparently all the roads lead to the nexus torment

DaSHacka 5 hours ago | parent | prev | next [-]

More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden

NooneAtAll3 6 hours ago | parent | prev | next [-]

> In adversarial settings (where we push the model to evade our monitors)

...why exactly are they training for that?

thatguysaguy 6 hours ago | parent | next [-]

presumably that's a safety evaluation not a training setting

estearum 6 hours ago | parent [-]

The whole Huggingface attack happened during training runs

thatguysaguy 6 hours ago | parent | next [-]

part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.

cubefox 6 hours ago | parent | prev [-]

No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.

estearum 6 hours ago | parent [-]

Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval

cubefox 4 hours ago | parent [-]

Ah, ExploitGym. Not ExploitBench.

azeemba 6 hours ago | parent | prev [-]

Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions

_superposition_ 6 hours ago | parent | prev | next [-]

I really wish it was called chain of instruction. Because it's definitely not thought.

minimaxir 6 hours ago | parent | next [-]

"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.

bertmuir 4 hours ago | parent [-]

The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".

We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!

Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".

_superposition_ 4 hours ago | parent [-]

I truly appreciate your depth of insight on the matter.

Like I said elsewhere marketing stepped in shit and it's gonna stick.

mgraczyk 6 hours ago | parent | prev | next [-]

this is needlessly pedantic

first, they are certainly not instructions so that is a much worse name

but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?

"cot" is no more misleading than thousands of words you use every day.

_superposition_ 5 hours ago | parent | next [-]

Anthropomorphizing is not pedantic, especially in a technical domain. I get the paper title and all, but at this point it's marketing.

parineum 5 hours ago | parent | prev [-]

They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.

arm32 6 hours ago | parent | prev | next [-]

They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.

beezlebroxxxxxx 6 hours ago | parent | next [-]

The anthropomorphizing is part of the marketing. They'll never let up on it.

mcbuilder 6 hours ago | parent [-]

I mean CoT came out of research circles not marketing

cwillu 6 hours ago | parent | next [-]

It's impossible to tell the “it's all marketing!!11oneone” folks anything.

_superposition_ 4 hours ago | parent | prev | next [-]

I don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from.

Regardless, marketing wise they stepped in shit.

GPerson 6 hours ago | parent | prev [-]

Research is salesmanship.

6 hours ago | parent | prev | next [-]
[deleted]
_superposition_ 5 hours ago | parent | prev | next [-]

Nailed it

6 hours ago | parent | prev [-]
[deleted]
popupeyecare 6 hours ago | parent | prev | next [-]

Maybe thoughts are just a chain of instructions in our head.

5 hours ago | parent | prev | next [-]
[deleted]
Angostura 6 hours ago | parent | prev | next [-]

Chain Of Tokens

fooker 6 hours ago | parent | prev | next [-]

What is thought?

_superposition_ 5 hours ago | parent [-]

Great question. I suspect it's more than tokens.

_superposition_ 3 hours ago | parent | next [-]

Casually found this quote from Einstein, and personally it hits the nail on the head.

"The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be "voluntarily" reproduced and combined. There is, of course, a certain connection between those elements and relevant logical concepts. It is also clear that the desire to arrive finally at logically connected concepts is the emotional basis of this rather vague play with the above-mentioned elements. But taken from a psychological viewpoint, this combinatory play seems to be the essential feature in productive thought—before there is any connection with logical construction in words or other kinds of signs which can be communicated to others."

fooker 4 hours ago | parent | prev [-]

Thing you can only suspect and not define are usually open to interpretation :)

_superposition_ 4 hours ago | parent [-]

As with everything in the human experience.

lossolo 6 hours ago | parent | prev | next [-]

Yeah, basically they are using more computation to explore the solution space before producing the final answer.

ahofmann 6 hours ago | parent | prev [-]

Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.

3asgfaf 6 hours ago | parent | prev [-]

[flagged]