Remix.run Logo
bryanrasmussen 11 hours ago

On getting domains that are Human slop.

You anonymized domains of pre-AI sources, are there any domains that had an excessive number of "telling you the same thing three times" or other AI tells among them?

Of the percentage that was misidentified, do they come from any sources in particular?

If you get a lot of content from these sources and run against the model do they perform worse? If they do how do they perform with word choice detectors? I would expect that structural slop is related to word choice slop among humans.

Anyway these are things I would be interested in as being the point where AI slop rubs up against the human slop which it learned from.