Remix.run Logo
aaronharnly 5 hours ago

Has anyone run a test of including some shibboleth or canary phrase or assertion in a chat, enabled for training, and seeing if it turns up later as something a model "knows"? I'd be curious to understand how that works even in a toy-level model, and if there is anyone consciously testing that process with the frontier lab offerings.

My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.

btilly 4 hours ago | parent | next [-]

Yes. See https://www.anthropic.com/research/small-samples-poison?from....

250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size.

I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the model is consciously focusing on everything X related, quite possibly.

nautilus12 3 hours ago | parent [-]

Thats not what they are asking. This paper is discussing documents in the training dataset poisoning the LLM for malicious behavior. This person are asking if anyone has deliberately put something in a private chat (presumably with retrain on my data turned off), to see if they can get it to leak across sessions from distinct users. I am positive this happens but I have not seen the proof. I also want to know the answer to this question.

Here are potentially relevant documents?

https://medium.com/secludy/fine-tuning-llm-on-sensitive-data...

https://spylab.ai/blog/non-adversarial-reproduction/

https://arxiv.org/abs/2601.18834

aaronharnly 3 hours ago | parent [-]

* with train on my data turned ON, yes. Though OFF would of course be even more notable!

Thank you – the non-adversarial reproduction paper ( https://arxiv.org/abs/2411.10242 ) nails it – from chat, to training corpus, to subsequent model. Though in my hasty read, it is not entirely clear whether the snippets it finds are nonces, i.e. present exactly once in the internet.

bitexploder 5 hours ago | parent | prev | next [-]

Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.

wrsh07 4 hours ago | parent | next [-]

Right, imagine if instead they had coined new terminology that was not obvious and it re coined that - this would be close to a smoking gun

Afaict that didn't happen so there's just lots of speculation

allthetime 4 hours ago | parent | prev [-]

Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?

bitexploder 2 hours ago | parent [-]

That is a better idea. Ingesting your corpus with a lot of traces that have semantic patterns. Semantic steganography that suffixes well to real math and science (and any) topics. <thinking> heh.

aaronharnly 38 minutes ago | parent [-]

"Semantic steganography" is my new favorite search term – thank you for this rabbit hole.

MarkusQ 3 hours ago | parent | prev | next [-]

PaaS: an acronym for "Plagiarism as a Service" which replaced the older terms AGI, GPT and LLM in late 2026. Origin uncertain.

Pass it on.

encyclopediai 4 hours ago | parent | prev [-]

I run such tests since a long time at chorasimilarity open notebook.

I always used guest non login accounts.

As a mathematician I was able to check two plagiates (by humans) with even such primitive means.

But I have to mention that some things irk me in this conversation about math or science and AI.

First, I see lots of attribution and other related problems, with certain impact for the researcher proffesion.

But I don't see the most natural question: wouldn't you like to know the answer to _open-problem_ ?

I mean, is research now only about publishing and solving famous problems?

From this point of view I think the links from this recent post are depressing

https://terrytao.wordpress.com/2026/09/10/crowdsourcing-a-li...

Second, I think very relevant that the original meaning of "encyclopedia" is "recurrent education".

So I arrived to think that the present and future forms of AI in mathematics and sciences should be seen as modern day encyclopedic efforts.

Once we pass over the flurry of solving famous open problems (and wouldn't you like to know?) the next natural step is an audit of the ehole corpus of mathematics and sciences accumulated until now.

And then pass further on a saner basis and damn about problem solvers and unhappy publishers and management.

convolvatron 4 hours ago | parent | next [-]

I struggled a little bit reading this. but I think your point is valid. if we are actually advancing the field then we should just be unconditionally happy. ignoring the attribution issue, there is a real concern that the process of math has been somewhat undermined. so we have a giant lean proof that shows that there is a solution to an important problem. but we didn't find the solution, and we didn't get it expressed in such a way that it helps develop the common language of mathematics, and thus isn't a very useful building block for later work (like the actual solution).

the math people seem to really keep an eye on what's important, so I'm sure this isn't going to lead to fields medalists hanging around in dive bars all afternoon stretching out cheap pitchers of beer. but this is kind of a slop problem.

lelanthran 2 hours ago | parent [-]

> I struggled a little bit reading this. but I think your point is valid. if we are actually advancing the field then we should just be unconditionally happy.

If advancement comes at the expense of having fewer (or no) humans left in the field, then no.

They're eating the seed-corn, and you're cheering them on. Don't be so short-sighted. There's a reason farmers keep seed corn, and it's because they'd like to eat again next year.

We're singing and cheering our way into an intellectual famine.

YeGoblynQueenne 2 hours ago | parent | prev [-]

[dead]