| ▲ | bryanrasmussen 12 hours ago | |
somebody named JochenMadler says below that #1 was essentially the study (for some reason their comment is dead, not sure why) I agree it is somewhat close to the study, but we are not sure, because it is not known how much is really boring dull marketing copy. In choosing blog posts from pre-AI times I suppose you might have difficulty finding the worst examples, and might accidentally get higher quality work. >Using the Wayback Machine, we collected 2,250 blog posts from 268 B2B company websites that were written before ChatGPT existed Not sure what metric was used to determine these 2250 blog posts? But there are certainly a lot of ways they can select higher quality posts by accident. on edit: evidently the Jochen from the study, maybe they thought your comment was AI written. | ||
| ▲ | bryanrasmussen 11 hours ago | parent [-] | |
On getting domains that are Human slop. You anonymized domains of pre-AI sources, are there any domains that had an excessive number of "telling you the same thing three times" or other AI tells among them? Of the percentage that was misidentified, do they come from any sources in particular? If you get a lot of content from these sources and run against the model do they perform worse? If they do how do they perform with word choice detectors? I would expect that structural slop is related to word choice slop among humans. Anyway these are things I would be interested in as being the point where AI slop rubs up against the human slop which it learned from. | ||