Remix.run Logo
tokioyoyo a day ago

Models aren’t trained as much from random internet text as they used to in pre-2024 era. Like they are, but specialized datasets get more attention.

MikeNotThePope a day ago | parent [-]

Your model/harness will indiscriminately do web searches to get answers. I believe that’s where the real risk is.

brunoarueira 16 hours ago | parent [-]

I don't know for sure, but will it require something like with high "SEO" rankings to fake the reputation of the sources? So if the LLM search for a specific spreaded knowledge splitted across many bad sites, it can poison the model maybe.

Terr_ 3 hours ago | parent [-]

You mean, how difficult will it be for an attacker to get their text seen by the LLM that does arbitrary searches?

It could be as simple as a malicious prospectus for AcmeCo, and then try to get AcmeCo on the radar so that the LLM-tools find your document and incorporate it. The malicious bits don't even need to be AcmeCo-related, they could be to pump (or dump) practically anything.