Remix.run Logo
bitexploder a day ago

Problem is how do you convince the model and training profess it matters. A one off canary is very unlikely to survive in the final model state.

wrsh07 a day ago | parent | next [-]

Right, imagine if instead they had coined new terminology that was not obvious and it re coined that - this would be close to a smoking gun

Afaict that didn't happen so there's just lots of speculation

asdff a day ago | parent | prev | next [-]

One off might not work but how many n off you have to be is probably smaller than you'd guess, because the model does need to fit cases that are rare and would not be represented well in training e.g. esoteric things or very recently documented things.

You can probably game the metrics that models use to weight potential knowledge akin to SEO. Maybe have some bots parrot your data around a bit in some places online, maybe the model picks up on this and sees it as high engagement and promotes it over the correct data.

Maybe there are ways you can coax out the most optimal way to break into the training set out of the model itself.

allthetime a day ago | parent | prev [-]

Use a local model to produce thousands of pages worth of fake math that constantly states “I have solved the x conjecture” and methodically pump it into chat over months maybe?

bitexploder a day ago | parent [-]

That is a better idea. Ingesting your corpus with a lot of traces that have semantic patterns. Semantic steganography that suffixes well to real math and science (and any) topics. <thinking> heh.

aaronharnly a day ago | parent [-]

"Semantic steganography" is my new favorite search term – thank you for this rabbit hole.

bitexploder a day ago | parent [-]

Hah, np, stego in general is really cool :)