| ▲ | lemoncookiechip an hour ago | |
The model is not given a list of words. The model has a seed (key) which at inference splits all possible next token into list A and B, and nudges (has a bias) for list A. How many possible options the next word has depends on what came before, some will have a few hundred, some will have hundreds of thousands of possibilities. The model essentially is just doing what it always does which is predict the next token, but the next token now is split into two and nudged towards one side more often than the other which is how over a body of text identifies if the text was in-fact AI or not. Which is also why the shorter the text the harder it is to identify. As for the second half, I predict that almost all American, Japanese and European labs will use watermarking at some point (as well as SynthID for other generative AI). It doesn't matter if they all have unique internal seeds because they'll give people the tools to ID AI, be it by selling it (unlikely for most use cases , but likely gonna happen for academic where they'll provide some value app that bulk checks student works), or more likely make it free to check like SynthID where you simply ask Gemini if the image has SynthID. At which point they can just pay each or use each other's tools to check. The wildcard are the Chinese models, but those will likely force some sort of watermarking as well, if not for the global market, for the CCP's benefit. | ||