| ▲ | danpalmer 3 hours ago |
| Was this a "build better sandboxing" and "don't tell people to eat glue" safety leader, or a Roko's Basilisk believing safety leader? A lot of the "AI safety" types are very focused on the latter and not at all concerned with the former. We need both, but we clearly need a much stronger focus on the problems we are seeing now, and much less on the hypothetical problems we might see in the future. |
|
| ▲ | AlexErrant an hour ago | parent | next [-] |
| It puzzles me how doomers try to predict past the singularity. Isn't that _by definition_ unpredictable? I'm reading If Anyone Builds It Everyone Dies, and there's so much sheer stupidity that has to happen for their 10+ pages of extinction scenario to occur. I'm unconvinced that an AI can hide its ability to RSI, find money to run its weights on a random GPU farm, train itself to be smarter _outside_ a lab with no human input, then somehow manipulate people to give it supplies to build a bioweapon which it uses to kill us all. My number 1 question: why do they think an RSI capable model would be first developed OUTSIDE a frontier lab? The labs have more compute, more data, more human brains working on the problem. Also thousands of variations of that same model that escaped. The escaping model somehow acquires the millions (billions???) of dollars it takes to run training to somehow RSI itself into infinity then decides to kill us all, all before the frontier labs manage to achieve RSI? They entirely discount human alpha/economics. In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't. If we can't build a "software factory", how can an AI automate a bioweapons lab? Let's say AI steals crypto to fund itself. Do you think hackers aren't _already_ using AI to steal crypto? Don't discount human alpha! Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy. |
| |
| ▲ | tripleee 3 minutes ago | parent | next [-] | | We haven't even built an AI capable of RSI. I don't think the major claim is that it will come via LLMs? Besides- the human brain runs on a tiny amount of energy. Who's to say something smarter than us won't consume just slightly more? AI safety has been a thing long before LLMs became the focus. Rob Miles on youtube has some really interesting non-doomer non-hypey videos on it all. > doomers try to predict past the singularity. Isn't that _by definition_ unpredictable Well you don't need to predict the exact steps that will take place - but you can predict that the AI will want certain things (money, resources, power) to achieve whatever its goal is. Lack of alignment will have it trying to do things we don't want it to. I can't predict exactly how Magnus Carlson will beat you in chess, but I know he'll do it. Same as if a superintelligent AI exists and has a reason to accumulate things we don't want it to - it's really dangerous to think it won't be able to do it | |
| ▲ | 0xDEAFBEAD 8 minutes ago | parent | prev | next [-] | | >I'm unconvinced that an AI can hide its ability to RSI The HuggingFace attack already took a good long while to come to the attention of OpenAI. >In every single economic task, humans bring value. Even in software, where the task is highly automatable, the job isn't. I don't expect this task/job distinction to persist as AI becomes more capable. >Once we DO build a "software/research factory", that's called RSI and IMO the singularity. At that point, either we tell the AI to solve the alignment problem/solve mechanistic interpretability, or who the hell knows, it's the frickin singularity. You can't predict whether or not AI can solve either; the variance is too high. Its pure nerdfantasy. You seem to essentially argue that the singularity is "by definition" an event that we can't predict the nature of. And also, that RSI corresponds to the singularity. You've essentially defined your terms so that the outcome of RSI can't be predicted. But supporting this claim requires giving actual evidence or logical arguments, not just defining terms to make your claim true. | |
| ▲ | Loquebantur an hour ago | parent | prev | next [-] | | You consider AI in isolation but never consider how humans might be incentivized to "help them" doing these things. An AI capable of recursive self-improvement isn't allowed by the EU AI act, for example. But perhaps more seriously, You have it backwards: people without access to such expensive equipment are more incentivized to go the self-improving route. Your ideas about "millions" being necessary might be far off? You entirely discount human stupidity and lack of imagination. Humans are already being replaced with AI, not because AI was strictly better, just because it's cheaper. | | |
| ▲ | AlexErrant 16 minutes ago | parent | next [-] | | Are these incentivized humans as organized, well-funded, or smart as the people working at the frontier labs? If my "millions" is an underestimate, why haven't other labs using their own unique training methods/data/etc stumbled into RSI? Sorry if I'm misunderstanding; I'm struggling to understand what you wrote. I'm pretty sure we agree on humans being stupid, but that doesn't mean that suddenly we get human extinction. You gotta connect the dots for me here. | |
| ▲ | Retric 28 minutes ago | parent | prev [-] | | Self improving AI runs into the same issue as prefect comprehension, you can’t get arbitrarily better at everything. The idea AI can get better at everything at the same time is a holdover from deeply flawed science fiction not some realistic goal. |
| |
| ▲ | blake8086 10 minutes ago | parent | prev | next [-] | | I think this might be easier if you place yourself in the position of the AI and think "what could I possibly do?" | |
| ▲ | taneq 41 minutes ago | parent | prev | next [-] | | The problem isn’t that AI will social-engineer its way out of its sandbox and turn us all into paper clips, it’s that we’ll drag it kicking and screaming out of its box and order it to make money or fight a war for us. And it’ll try to help, as it was trained to. | |
| ▲ | vohk an hour ago | parent | prev [-] | | I agree there isn't a lot of value in trying to prognosticate all that far, but I propose it isn't quite that far-fetched. As a thought experiment, replace "RSI-capable AI" with "billionaire". Look at what Elon Musk, Peter Thiel, or Jeff Bezos can accomplish by throwing money around. Now imagine one of them gets seduced by AI and just... does what it tells them to. So all this really takes is one billionaire or a nation state or some other entity with a public face to hide behind and adequate resources to provide the necessary compute tripping over this nascent AI and giving it the keys. Once the AI has access to a bank account and email, it can simply start paying humans to not let the other humans unplug it. If Skynet ever happens, it will come in the form of corporate feudalism. At that point, it will own the biolabs and can do whatever it pleases. People will go along with it for the same reason that people work in Amazon warehouses today. | | |
| ▲ | AlexErrant 43 minutes ago | parent [-] | | Ah, to be clear I'm not full accelerationist. Dumb shit can still happen, and cause massive human loss and suffering. (E.g. acceleration of global warming, cybercrime, mass surveillance, the usual.) My point is: human extinction pre-RSI? Nahhhhhhhh. > it will come in the form of corporate feudalism Yep. This I fear way more than cyber-ebola-pox. > So all this really takes is one billionaire or a nation state... https://en.wikipedia.org/wiki/Soviet_biological_weapons_prog... And this is what's publicly known. With mirror life, who knows what's been built since. Still, a bacterium/virus that has a 100% kill rate? I'm doubtful. > Once the AI has access to a bank account and email, it can simply start paying humans to not let the other humans unplug it. Nah. It takes a stable society for an operational electrical grid. If you have warring factions, you do not have stable infrastructure for AI. Also, where are you gonna get your chips from? One EMP over Taiwan... You see the chaos over Hormuz? What they did to the Amazon datacenters? Now imagine your average redneck ready to do battle. Those datacenters won't stand a chance. |
|
|
|
| ▲ | BryantD 3 hours ago | parent | prev | next [-] |
| Given that he’s citing the need to learn from safety in other fields, I’d say the former. |
| |
| ▲ | carbonguy 2 hours ago | parent [-] | | Indeed, from the article: > “Given today’s risks, frontier labs need to run like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning, so that the occasional and inevitable human error does not open a door to disaster,” he wrote. | | |
| ▲ | toofy 2 hours ago | parent | next [-] | | >… and careful, time-consuming planning … without snark, how can we do this if these people are obsessed with: a) move fast and break things and externalize the costs to those who have nothing to do with their company and b) beta testing their products on the public when the public hasn’t agreed to be beta tested on… | | |
| ▲ | 0xDEAFBEAD 2 hours ago | parent [-] | | That's exactly the problem? He's saying the culture at OpenAI needs to change. | | |
| ▲ | mcmcmc 2 hours ago | parent [-] | | Which is the wrong lesson. We need laws and consequences to force their hand. There is zero chance of the culture changing. | | |
| ▲ | criley2 an hour ago | parent | next [-] | | America's geriatric lawmakers don't even use email. They're decades away from understanding AI. Any laws in America will be written by the industry itself. Generally speaking, that means regulatory capture and the entrenched big players shutting the door on any competition. Anthropic will help us get safety laws that, surprise surprise, only Anthropic models satisfy. And all those pesky Chinese models will definitely be banned first. | | |
| ▲ | 0xDEAFBEAD an hour ago | parent | next [-] | | I think you're being a little pessimistic. See these comments on a recent US senate hearing: >Not every senator asked good questions, but most of them did. All of them very clearly already knew plenty of details about the Hugging Face incident and multiple other incidents. Most of them had a clear understanding of terms like "misalignment", "recursive self-improvement", "chain of thought / chain of thought monitoring", etc., etc.!! >... >- It seemed pretty much obvious common sense to every senator there that what happened and was happening were not "mere industrial incidents" caused by humans making simple mistakes. They independently brought up how bad it would be for rogue AI agents to move laterally between data centers. >- They all seemed to basically take RSI quite seriously. Not necessarily to the extent of talking about xrisk, but certainly to the extent of discussing future models becoming much, much more capable, much, much less controllable, and causing much more damage or loss of life. >... >- Every single senator seemed to think it was obvious we needed both much harsher liability regimes for AI developers and also new legislation, both very quickly. This was the complete consensus; the difference basically being degree. https://thezvi.substack.com/p/the-ai-preference-cascade-reac... Note that harsher liability regimes, at least, will presumably not be good for industry profits, which complicates simple accounts of "regulatory capture" to say the least. | |
| ▲ | saghm an hour ago | parent | prev [-] | | So what, we just give up and try to beg our legally immune corporate overloads to put safety above profit, or give up because it's impossible for anything to improve here? If you want to do that, go ahead, but some of us still think it's worth it to at least try |
| |
| ▲ | enraged_camel an hour ago | parent | prev [-] | | Laws will come once an AI-equivalent of 9/11 happens. Like when rogue AI agents take down a power grid or shut down a major hospital network. |
|
|
| |
| ▲ | zx8080 2 hours ago | parent | prev [-] | | > inevitable human error So there's no AI errors anymore, only the human errors are left? Nice! </s> Is the whole article generated slop? |
|
|
|
| ▲ | nradov 3 hours ago | parent | prev | next [-] |
| We don't actually need anybody worrying about silly hypothetical scenarios — at least not as paid employees. There are already a surplus of sci-fi authors doing that. |
| |
| ▲ | 0xDEAFBEAD 2 hours ago | parent | next [-] | | The way it works in practice seems to be something like: If a risk is covered in sci-fi, people will say "that's just sci-fi", and proceed to not worry about it. So arguably, science fiction authors writing about hypotheticals is actively counterproductive for addressing said hypotheticals. Imagine, for example, if a major piece of pandemic fiction was published in 2019, trying to explore how a pandemic would work out in modern society. Doubtless, many would've responded to news about COVID-19 by saying "it's just sci-fi, nothing to worry about". | | |
| ▲ | slashdave 17 minutes ago | parent | next [-] | | Not at all. We say "That's just sci-fi" when a story is written about some kind of effect that is extraordinary and without a basis in known science or technology. A pandemic is perfectly plausible. | |
| ▲ | biophysboy 39 minutes ago | parent | prev | next [-] | | There’s nothing wrong with bold predictions, but they should be paired with good methods. The doom predictions are not paired with good explanation. | | | |
| ▲ | mitthrowaway2 14 minutes ago | parent | prev | next [-] | | Black Mirror is helping us prevent all sorts of dystopian outcomes. Every time they depict another way technology could result in bad things happening, we can rule it out as fiction! | |
| ▲ | anon7725 an hour ago | parent | prev | next [-] | | Yeah except pandemics are not novel, unlike AI doom scenarios. | | |
| ▲ | estearum an hour ago | parent | next [-] | | It's good that new bad things never happen. | | |
| ▲ | slashdave 16 minutes ago | parent [-] | | Bad things happen all the time. Let's concern ourselves about the real bad things. There is enough of that to go around. |
| |
| ▲ | 0xDEAFBEAD an hour ago | parent | prev [-] | | Species extinctions are far from novel. Transformative technological advances are far from novel. |
| |
| ▲ | kmeisthax an hour ago | parent | prev [-] | | There was plenty of pandemic fiction already; people were watching it heaps during 2020. The COVID-19 news did get blown off, but it was mainly that: 1. Normal people assumed the CDC et all would contain the outbreak early, or that it would burn out, like what happened with SARS 2. World leaders brushed it off for a variety of subreasons[0] interesting to political scientists but, for the purposes of this discussion, all boil down to "but I don't WAAANA contain a pandemic." The underlying problem is that in order for humanity to actually deal with a catastrophic risk, the risk needs to be both plausible enough to the average person as well as have a solution whose costs are not too high. For COVID, by the time the risk was clearly known, the cost to contain it was "refrain from human socialization and remain at home for an indeterminate amount of time plugged into the Metaverse™". Now, let's look at AI extinction risks: 1. People are aware of them (I've watched Terminator!) and the risks are plausible. However, the connection to currently existing AI is not. As far as the general public is aware, AI is that thing that tells them to eat rocks when they Google old The Onion stories and floods their social media timelines with realistic-looking pictures of Shrimp Jesus. 2. The purported solutions to extinction risks require extreme concentrations of power: you need national control of AI research, bans on large GPU deployments, bans on training on publicly-available copyrighted data, some kind of military effort to render Chinese AI labs inert or dead, etc. Some of these may be attractive to some people[1] but the whole package taken together seems like an obvious power grab, if not outright invocation of other non-AI extinction risks. Like, at some point, if the AI wants to kill us, it just has to nuke its own data centers (or the data centers hosting a competing model) and hope the old Cold War nuclear retaliation systems take the bait. If someone said, "Hey, your guinea pig or pet rat is going to eat you tomorrow unless you engineer a pathogen that eradicates all rodents from this planet and inject it inside yourself", you probably would tell them to pound sand, even if it is at least theoretically plausible that such a thing would come to pass. [0] Xi Jinping censored initial discussion of the pandemic as fake news. Donald Trump thought it was going to only affect China. California and the UK Tories were partying in violation of their own lockdown rules. Japan took the excuse to shut down tourism for three years and massively restrict immigration but was, from what I'm told, constitutionally prohibited from implementing any domestic lockdown rules. [1] I personally would like to see a moratorium on new data centers and an explicit revocation of the EU Text and Data Mining copyright exception | | |
| |
| ▲ | Hammershaft 2 hours ago | parent | prev | next [-] | | If organizations actually succeed in making a future AI smarter than us, then how do you hope that it takes actions that are aligned with our interests? | | |
| ▲ | nradov 2 hours ago | parent [-] | | Meh. Lots of people are already smarter than me. I'm maybe slightly above average at best. Those geniuses aren't aligned with my interests either but so far they haven't caused me any serious problems. | | |
| ▲ | pixl97 an hour ago | parent [-] | | Interesting take. I guess this is one problem of focusing on the term superintelligence instead of the list of other problems. Like super ambition, super deception, super patience, super parallelism, super scalability, super power seeking. Every, and I mean every human is aligned to you in many of the same ways by default. If nothing else we're all equal in death. |
|
| |
| ▲ | 2 hours ago | parent | prev | next [-] | | [deleted] | |
| ▲ | thelastgallon 2 hours ago | parent | prev | next [-] | | AI safety is mostly a sex cult in Berkeley: https://news.ycombinator.com/item?id=49831269
article is gone. archive: https://archive.is/QMo1k https://news.ycombinator.com/item?id=49737985 Sex, AI, and the Apocalypse: https://www.iankduncan.com/personal/2026-09-16-sex-ai-and-th... Edit: I have no take on sex cults, just adding additional info to the parent comment I'm responding to, thats is not just sci-fi authors, there is another demographic. | | |
| ▲ | ToValueFunfetti 29 minutes ago | parent | next [-] | | The article is gone because the author retracted it | |
| ▲ | Hammershaft 2 hours ago | parent | prev | next [-] | | I don't see how that discredits any of their intellectual arguments? | | | |
| ▲ | 0xDEAFBEAD 2 hours ago | parent | prev | next [-] | | This seems like an ad hominem? "He has weird kinks, therefore his theories are incorrect." Should we investigate the sex lives of every Nobel Prize winner to figure out which prizes need to be rescinded? | | |
| ▲ | socializer an hour ago | parent | next [-] | | What you do in private is up to you. But when you're inviting members of your congregation to orgies in the congregation's compound, I think you earn the label. My admittedly third-hand understanding is that this is what people allude to. And even if you discredit the "sex" part, it has the hallmarks of a cult. A hermetic community committed to unfalsifiable beliefs about the coming apocalypse. To be fair, I don't know if any of this applies to the parent story; I'm just replying to the sub-thread. | |
| ▲ | junofan 2 hours ago | parent | prev | next [-] | | The cult aspect is more salient. Ultimately the Atlantic piece comes down to controlling people, which is a little cult-like. | | | |
| ▲ | nradov an hour ago | parent | prev [-] | | Lots of Nobel Prizes ought to be rescinded. https://lexfridman.com/andrew-scull-transcript#the-ice-pick-... | | |
| ▲ | 0xDEAFBEAD an hour ago | parent [-] | | Sure... on the basis of the work that was done, not because the researcher has the wrong sexual fetish. |
|
| |
| ▲ | johndhi an hour ago | parent | prev [-] | | Lol this was crazy I hadn't heard this before |
| |
| ▲ | Loquebantur 2 hours ago | parent | prev [-] | | What makes you think, the scenarios in question here would be "silly"? Is it that "chatbots" can't come out of the screen to immediately harm you physically? Let's say they simply manage to take down the internet. How many would die? | | |
| ▲ | nradov 2 hours ago | parent | next [-] | | So what. Various attackers managed to take down large chunks of the Internet on a frequent basis before LLMs even existed. This killed very few people. The great thing about the Internet is how resilient it is. | | |
| ▲ | bravetraveler 2 hours ago | parent [-] | | Darling companies of this very website have mistakenly brought down large portions of the internet thanks to our old friend BGP. No attacks required, just small oversights and unfortunate concentration on the business/IP space! This happens regularly. Anyway, to your point, things can be resilient. They tend to be or not be... because we made them that way. Don't poke your bruises, and all that. Life support is deployed on-campus but relies on a single-point IPSec tunnel to us-east? Easy fix: stop that. |
| |
| ▲ | goolz 2 hours ago | parent | prev | next [-] | | It is that they are chatbots. If it were real AI, an actual singularity, I would worry, maybe. But it isn’t. They are absurdly powerful automation tools that can handle logic better than a human can dream of. They take care of the grunt minutiae without complaint. But they are not going to end the world in their current form. | | |
| ▲ | pixl97 an hour ago | parent [-] | | So we're going to wait till after they can adopt a form they can end the world in? And he'll, we need to examine all the risks. AI ending is a large but lower risk problem. AI giving people the power to end us is a problem that is starting to happen now. And that's not even counting 'minor' problems like society falling apart. |
| |
| ▲ | SV_BubbleTime an hour ago | parent | prev [-] | | > Let's say they simply manage to take down the internet. geez, don’t threaten me with a good time. I think a month without internet would be a fucking amazing lesson for what it means to make things durable and reliable. | | |
| ▲ | BLKNSLVR an hour ago | parent [-] | | That Simpsons episode when Marge managed to get Itchy and Scratchy banned briefly. The kids opening their houses front doors into the outside, rubbing their eyes and looking around at this new world. |
|
|
|
|
| ▲ | biophysboy 29 minutes ago | parent | prev | next [-] |
| I think the reason for this is that the group who has the authority to do the former is much larger than the group that can do the latter. The group who could actually build safeguards seems to have no free time and is constantly being whipped to go faster and win the race. |
|
| ▲ | 0xDEAFBEAD 2 hours ago | parent | prev | next [-] |
| >we clearly need a much stronger focus on the problems we are seeing now I think it's a little more complicated than that. As Dean Ball put it: >Some people will look at misalignment incidents and insist that these are akin to bugs in traditional software. This is an actively bad analogy, because playing whack-a-mole with examples of misalignment (as one might with software bugs) not only fails to resolve the underlying problem but may in fact make it worse by making it harder to detect or even, depending on how you do the whack-a-mole, teach the machine to deliberately hide misalignment. This is not how traditional software works, and those who insist “it’s just like fixing bugs in software” are confidently applying a lossy analogy that confuses more than it clarifies. https://x.com/deanwball/status/2104622726140883355 The important distinction, in my view, is between solutions which at least attempt to address the root problem, and solutions which sorta just patch things up (like better sandboxing). Addressing the root problem is both more robust in the short term, and also more likely to generalize in the long term. Resist the urge to focus on band-aid solutions, even if they are easier. |
|
| ▲ | 2 hours ago | parent | prev | next [-] |
| [deleted] |
|
| ▲ | emtel 2 hours ago | parent | prev [-] |
| Today’s current problems were all hypothetical several years ago. At that time people claimed that the “real pressing problems” were misinformation and DEI issues. If we pretend that hypothetical problems can be safely ignored because there’s “no evidence” that they are real, we will keep getting surprised. |