| ▲ | paxys 9 hours ago |
| All the information Gemini surfaced was created with human effort and published on the internet with the expectation that humans would visit the website and the creator would get some reward - advertising dollars, bragging rights, popularity, subscribers or whatever else. If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content? |
|
| ▲ | dwayne_dibley 3 minutes ago | parent | next [-] |
| Easy to fix a well documented router now. Difficult to fix a non documented router in five years time because no one has been contributing to the web about its bug fixes. |
|
| ▲ | neuralkoi 8 hours ago | parent | prev | next [-] |
| I've seen websites put up some draconian measures to try and get a grip on the scraping. So much for the sub-second loading experience when you have Cloudflare, Google, Anubis, and all these other captcha services trying to see if you're a human. It's made the web browsing experience so much worse. Some of the proposals to address this include charging bots for access to web resources, but they will also have repercussions for regular users. I don't see how you solve this cleanly. |
| |
| ▲ | robinsonb5 2 hours ago | parent | next [-] | | Sure - it sucks, unfortunately the alternative is the sites going away entirely. When the load from scraper bots is constantly knocking the site offline the choices are literally to allow it to remain inaccessible for much of the time, put up a layer of defenses with all the user-annoyance compromises that entails, or just give up and unpublish the site. | | |
| ▲ | Borg3 2 hours ago | parent [-] | | The alternative is simple.. Go dark. VPN tech is known from like 30 years. Pretty much everyone can use it (VPN providers). But instead using it to browse net, build VPN overlay networks of interest for people. Gaming networks, R&D networks, Retro Networks. People will peer to PoP and use resources. Bad actor? BAN it from network. You have control. This could be done in Internet, but big corpos and big money won the battle. Just wake F*ing up... | | |
| ▲ | kukkeliskuu 22 minutes ago | parent [-] | | Continuing on your suggestion. There could be open source tooling to create custom private "closednets", with - trust ring mechanism to allow invitations, flagging, banning, and banning those that invite people who were banned - the rules of the closednet - search engine with opt-in scraping - portal (remember the 80s?) with all the registered nodes, perhaps by service category such as public git repo hosts, web sites etc. etc. The first closednet could be Hacker News. |
|
| |
| ▲ | Grimburger 2 hours ago | parent | prev | next [-] | | Cloudflare specifically has a block for LLM and AI training bots now. Not sure of the effectiveness but it's there. | | |
| ▲ | matherial 2 hours ago | parent | next [-] | | Minimal. I'm behind Cloudflare and 90% of the traffic is still scrapers. I don't think they're serious about the long tail. I think the main thing Cloudflare is trying to do is block direct traffic from frontier labs and then start charging them for access. They might end up shooting themselves in the foot, as this simply empowers sketchy residential-proxy outfits to undercut Cloudflare and sell the data to labs for less. | |
| ▲ | bakugo an hour ago | parent | prev [-] | | It still only blocks "well-behaved" bots that have proper User-Agents and respect robots.txt, so it's largely pointless. The problematic bots are all disguising themselves as Chrome and sending requests from millions of residential proxy IPs, and the only real solution to those is some sort of captcha or PoW page on first visit. |
| |
| ▲ | JKCalhoun 8 hours ago | parent | prev [-] | | Me, I'm just scraping the parts of the internet I like, toying with local LLMs… ready really to just shove off. |
|
|
| ▲ | littlecranky67 2 hours ago | parent | prev | next [-] |
| > "humans would visit the website and the creator would get some reward" That expectation is a problem, has always been a problem, and Tim Berners Lee never mentioned anything about a reward structure when coming up with the WWW. |
| |
| ▲ | pferde 2 hours ago | parent | next [-] | | Your thinking too narrowly about the reward. Sometimes, it's just about the getting the knowledge out there that's motivating the creator, not anything tangible for themselves. | | |
| ▲ | boredhedgehog an hour ago | parent | next [-] | | > it's just about the getting the knowledge out there that's motivating the creator In that case the creator should welcome AIs with open arms; a human reader will forget eventually, but the AI will preserve the knowledge forever. | | | |
| ▲ | esposm03 2 hours ago | parent | prev [-] | | > Your thinking Should be "You're thinking". |
| |
| ▲ | gspr 2 hours ago | parent | prev [-] | | Did the operators of HN mention anything about a reward structure for posting comments here? I'm sure you can see how that's still attractive to some. |
|
|
| ▲ | jefftk 9 hours ago | parent | prev | next [-] |
| > If the only visitors to websites are now LLM training bots then what incentive is there to publish anything new? I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me. |
| |
| ▲ | exmadscientist 8 hours ago | parent | next [-] | | The danger that's concerning people (rightly or wrongly) isn't that LLMs are going to be an intermediary to your website. It's that they'll be the only thing reading it. No one will ever read your post or know what you wrote. The only consumers will be LLMs, they'll train on a version that strips out you as the author (probably more due to expedience than any sort of malice; it's not like you're famous, are you?), and your idea might get embedded into a set of model weights somewhere. No human will see a byte of it. Are you actually saying you'd be OK with that? | | |
| ▲ | JKCalhoun 8 hours ago | parent | next [-] | | Not the OP, but I suspect no one goes to my website anyway. (I write nonetheless.) | | | |
| ▲ | jefftk 7 hours ago | parent | prev [-] | | Yes, that would be fine. I write primarily communicate ideas, not for credit or fame. Empirically, however, LLMs don't strip out the author: the big models know a lot about what I've written even with search disabled. Ex: https://claude.ai/share/8cbcdf88-a360-421a-8c06-ae7b7992e866 | | |
| ▲ | frabcus 16 minutes ago | parent [-] | | Sure they know about you if you ask, but generally they won't credit you if they cite an idea from their latent space that came from you. |
|
| |
| ▲ | shuwix 4 hours ago | parent | prev | next [-] | | Your level of thinking is defined by ICD-10. Yeah, authors don't want to be recognized as authors, they don't want any reward for their work, they don't want to amass pool of loyal readers, interact with them, etc. All they want is for halucinating AI to take excerpts of their work and compile it with random sh!t. GENIUS | | | |
| ▲ | mittensc 2 hours ago | parent | prev | next [-] | | I mainly shared my projects for learning, discussion and bragging rights. LLMs just use everything, generate similar code with no attribution and keep users from visiting, so no bragging rights or attention. Worse, there are some PRs that seem fully generated ... So i mostly stopped sharing and started pulling my old repos offline. At this pace, i don't want to compete with a clone of myself in the future that will do my work for much cheaper. | |
| ▲ | laughing_man 7 hours ago | parent | prev | next [-] | | There are different types of writing. If we depend on people writing because it's enjoyable at some level, we're going to lose writing that's important but also a bit tedious. | | |
| ▲ | jefftk 7 hours ago | parent [-] | | Of course! I do think we'd lose a lot of great writing if it went amateur-only. But my parent seemed to be saying the incentive would entirely disappear, so I wanted to give my perspective. |
| |
| ▲ | gspr 2 hours ago | parent | prev [-] | | > I write because I have ideas I want to share, and whether that happens with LLMs as an intermediary isn't important to me. Sure. But you can see that for some people (myself included), writing for peers is part of the joy? And that if instead a megacorp places an opaque computer program between the author and the readers, that joy might be ruined? |
|
|
| ▲ | swingandamiss 8 hours ago | parent | prev | next [-] |
| Gemini can just consume the device documents. There's an incentive for device makers to publish this content. |
| |
| ▲ | dspillett 8 hours ago | parent [-] | | There had always been some incentive for manufactures to publish device documentation, and yet it has often been quite lacking either in quality or overall existence. I doubt LLM/agents being the readers will change that at all. What I expect AI scraping and using without credit will impact is people publishing their own unofficial help and guidance, and the affect there is likely to be negative. It won't stop all of them, but enough to be noticeable. Another possible negative is the manufactures documentation being AI generated without sufficient review, so possibly more erroneous than before, or intentionally not producing full documentation at all and expecting AI to fill the gap (MS seems to be heading this way: pushing "ask copilot" all over Azure instead of links direct to good reference material). All this would add up to a situation that is somewhere between "a little worse than pre-AI" and "an absolute shit show". |
|
|
| ▲ | no-name-here 6 hours ago | parent | prev | next [-] |
| At least LLMs almost always transform the original - it usually isn’t as straightforward as “Here’s the original but without the ads that pay for it”. But we already have the latter case that exists - ad blockers. Ad blockers literally serve up the word-for-word original content minus the ads. |
|
| ▲ | penneyd 8 hours ago | parent | prev | next [-] |
| Well in the example above the manufacturer still has incentive to provide the manual's and guides that describe how to use their products, and if that is subsequently served by an LLM that's totally fine. The only sites that LLM's would have a negative effect on are those that are only hosting content for the ad views. |
| |
| ▲ | Eddy_Viscosity2 8 hours ago | parent [-] | | Manuals don't always well explain how to use their products with everybody else's products because there are too many to do that. But there are lots of people trying things out and might figure out the fine details on how to make various things work. They then publish these how-to pieces (which exist no where else) to the internet, or at least they used to when there were incentives to do so. |
|
|
| ▲ | latexr 2 hours ago | parent | prev | next [-] |
| Exactly. What’s problematic about comments like your parent is the absence of mid-to-long term thinking. It’s like bragging about a new highly addictive psychedelic drug that a dealer gave you a taste of for free. The effects are awesome today, you feel so fun and free! Never mind that it’s destroying your body and that the dealer will eventually charge you or demand you pay in other ways, that’s a problem for another day. Weeee! |
|
| ▲ | karmakaze 6 hours ago | parent | prev | next [-] |
| Funny, I published information in the hopes that humans would benefit from it. If it happens to be through collective intelligence of LLMs I'm ok with that--even more so if through open models. |
| |
| ▲ | wartywhoa23 2 hours ago | parent [-] | | > Funny, I published information in the hopes that humans would benefit from it. Sure, humans would benefit. It took them searching, reading themselves, maybe even understanding something in the process, to complete a 360° revolution of their squirrel cages in time T. Now they can omit searching, skip reading to the regurgitated answer, throw away understanding, and complete a full revolution in T/N, where N is a heuristic value directly proportional to the amount of skin in the AI hype. But the catch is that the squirrel cage must run non-stop still. |
|
|
| ▲ | satvikpendem 8 hours ago | parent | prev | next [-] |
| People write and create regardless of profit motive, it has been that way for thousands of years. |
| |
| ▲ | watwut 3 hours ago | parent | next [-] | | People are way less likely to write if there is no one to read it. And blog were also monkey see monley do - people seen other peoples blogs and got inspired. When people wont see others blogs, they wont start writing own. When there will bw no ome to actually read it, they will go to do something else. | |
| ▲ | _DeadFred_ 7 hours ago | parent | prev [-] | | Actually no. The original copyright laws were created in part because of the realities of needed profit motive to have high value writing done, time consuming compilation work done. It was even titled "An Act for the Encouragement of Learning". The thousands of years writing you are talking about was often funded by patrons, who kept the output in their private libraries to show off (and maybe lend out) for prestige. It was a horrible limitation of knowledge and ideas. Much worse than the profit motive, copyright based system that came after that spawned a new age of knowledge in which everyone had cheap access, and those that didn't had access to the (no longer just private) libraries. I'm sure the billionaire class would love a return to patronage based libraries, NDAs on authors of books, and the elitism they would feel with a return to private libraries locking away all kinds of knowledge that would happen if patronage become the only way authors could make money (such as with AI just regurgitating their works, or if the stupid 'do away with copyright' people got their way). | | |
| ▲ | graemep an hour ago | parent | next [-] | | Actually it grew out of censorship and monopolies: https://en.wikipedia.org/wiki/Statute_of_Anne#Background Cheap access came from the invention of cheap printing . The laws were passed to restrict it. | |
| ▲ | nairboon 3 hours ago | parent | prev [-] | | It's not copyright that caused cheap access. The printing machine allowed for cheaper publications, that's what spawned the new age of knowledge. The raw materials and the duplication of knowledge was the bottleneck. With digital systems this cost is minuscule, but still there. |
|
|
|
| ▲ | _dark_matter_ 9 hours ago | parent | prev | next [-] |
| I wonder the same thing. I only imagine that what comes next is worse: AI companies using vast resources to develop new training data, in house, locked down. They are already doing this with developers and code at Meta. Information will become locked away behind AI paywalls and chatbots. |
|
| ▲ | 9 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | ekianjo 2 hours ago | parent | prev | next [-] |
| > incentive is there to publish anything new? For how long can we continue to rely on pre-2024 non-AI generated content? obviously new content still has value because it remains the source layer for LLM agents. it just wont be ads giving you revenues thats all. |
|
| ▲ | Havoc 8 hours ago | parent | prev | next [-] |
| That is next earnings quarters problem is the approach being taken |
|
| ▲ | dudefeliciano 2 hours ago | parent | prev [-] |
| Well companies are starting to put hidden ads in text content if the user agent is an AI crawler |