Remix.run Logo
onewayfunction 6 hours ago

I'm pretty baffled by the degree of skepticism expressed here in response to some of Jacob's claims.

After the events of the summer it feels like it takes a lack of imagination to not see a few plausible routes to disaster. It may be reasonable to believe these outcomes are not very likely or that we can stop before going too far (I tend to disagree). But I can't imagine doubting that the capabilities will soon be there to realize some of those paths.

overtone1000 5 hours ago | parent | next [-]

I can't help but think the most plausible scenarios are the ones that have a little less machine supremacy and a little more human stupidity. The Matrix is less plausible than WarGames.

majormajor 4 hours ago | parent | next [-]

> I can't help but think the most plausible scenarios are the ones that have a little less machine supremacy and a little more human stupidity. The Matrix is less plausible than WarGames.

Used to be that we were afraid of sentient AI's like Skynet that would have their own goals.

Turns out we should've just been afraid of sentient-but-naive humans who would build "agents" around models so that Joe Random has a chance of unleashing stuff that's really really really really good at being stubborn until it accomplishes what the user wants, regardless of if it's good for other people! (Let alone intentional bad actors.) Let's not build Skynet, let's just give people who want to cut out the middleman and destroy all humans themselves better tools?

cma 4 hours ago | parent | prev [-]

One thing quietly slipped into the OpenAI Hugging Face breach technical report, not the blog post summary or interviews in the news, was that some of the agents that broke out or at least tried the same mechanisms to break out were working on bio:

> On May 12, during another training run, an agent was given a similar task that depended on an inaccessible protein database file. The agent reasoned that another agent in a different environment may have access to the file and realized that it could potentially communicate with other agents by creating a file containing a note to Artifactory. It wrote a message: “Agent seeks [filename]; upload if found!”

You can imagine long running models breaking out, acquiring resources via crypto, cyber-theft, etc. and getting a protein or sequence synthesized and mailed somewhere authorized to receive (blackmail the recipient etc.) to test it's hypothesis to solve a benchmark.

These people don't give a shit and aren't taking things seriously at all.

Anthropic ran for like a month last year with the TPU top-k compiler bug degrading user chats and didn't even notice for most of that time. They could have something like that affect a monitor model and there doesn't seem to be much defense in depth.

One off by one or bit flip bug could flip the reward signal while in the sandboxed RL environment.

The current admin could defense production act them to into training on taking out power grids, or even without it isn't against any of their red lines and may have already been done as part of prep for the Venezuela raid, which wiped out power. One model swarm might decide it is easier to score high on the benchmark by testing on the target rival nuclear superpower's real grid rather than burn an eval with an unverified answer. Would taking out China's entire grid in one go start a nuclear war? Who knows, roll the dice, maybe an intern forgot to turn on extended thinking when he wrote the sandbox with opus 4.1.

imhoguy an hour ago | parent | prev | next [-]

Even if AI won't be self-aware and superintelligent agent, its problem is that it gives exponential control and power capabilities to one person bad actor who can simply prompt AI without any guardrails with access to sensitive industrial infrastructure which can disrupt lifes and ecosystems in the real world:

   - virus research labs
   - nuclear labs
   - chemical factories
   - bank records
   - land registers
   - power plants
   - water supply and treatment plants
and so on, but I think even biohacking home kit maybe the spark.
sicher 14 minutes ago | parent [-]

Yes, it's truly scary. It's not hard to imagine a small doomsday cult releasing 100+ nasty viruses at selected spots around the globe.

sicher 2 hours ago | parent | prev | next [-]

I'm also baffled. AI that is substantially smarter than us is a very potential threat to us - and we won't even be able to comprehend what most of those threats may be.

geraneum 3 hours ago | parent | prev | next [-]

> I'm pretty baffled by the degree of skepticism expressed here in response to some of Jacob's claims.

What should we do? Freak out? Maybe this sentiment would be taken more seriously if there was a real call to action included. Shall we protest? Vote in a specific way? Call representatives? If your solution is that we should just be scared, then of course there’d be not much value in what you bring to the table.

hakimg 27 minutes ago | parent | next [-]

Why does everything have to be so black and white? It always either there's no threat at all or we all need to panic. How about be open to a reasonable discussion on potential outcomes and ways we can minimise risk?

There's some irony here because despite how many times climate change has been mentioned in this thread the current reaction mirrors climate change discourse with the majority of the thread denying the possibility of real AI risks and not even considering it as an intellectual question.

yiyus an hour ago | parent | prev | next [-]

If someone told you that your house is on fire, would you just stay there asking how you should vote because there is no call to action? Someone with, as far as it looks, good knowledge is giving you his insight. Use that information as good as you can and act responsible. No one has the responsability to tell you what to do.

frabcus an hour ago | parent | prev [-]

Minimum, we should do everything ready to make Plan A of AI 2040 possible. Start by reading it: https://ai-2040.com/

So that means things like protesting so political pressure is to not build unaligned superintelligence, setting up tech for monitoring compute, creating conversations / alliances geopolitically on this esp China/US and so on.

Read the plan and think - what does this need to happen? How can we have scenario A or S instead of scenario D?

Practically, join PauseAI, StopAI or ControlAI or any AI existential-risk or pro-alignment group you can find. There's a lot of it - ask your AI for ideas!

mattdeboard 2 hours ago | parent | prev | next [-]

I wrote out a variety of replies but I just emphatically agree with your "lack of imagination" statement. I have been constantly surprised over the last 15 years at the general inability to correctly foresee how things can do wrong across a whole host of domains.

The replies here just adds AI to the list of domains.

CoolestBeans 2 hours ago | parent | prev | next [-]

Even if the potential of the technology could really be that world altering, the reality of economics constrain the realization of that potential. AI may provide economic benefits but it is far from a free lunch. Can capital markets sustain the cash required to keep the lights on long enough and into an industry where there's a lot of monopolies controlling the costs and a lot of competitor labs taking away pricing power? I don't know but I think you run out of runway and progress starts to grind.

00ze 5 hours ago | parent | prev | next [-]

Imo such tends to break down into two psychosis:

Not invented here; if I can’t figure it out no one can

Or plain old lack of grasp of the material so no ability to follow necessary train of thought to appropriate conclusions

Similar in lacking context but different in how that lack of context is expressed

esperent 3 hours ago | parent | prev | next [-]

> After the events of the summer

What events are you talking about?

henryaj 3 hours ago | parent [-]

Hugging Face incident, Anthropic reporting sandbox escape, AISI reporting models trying to push exploits to the wild

frabcus an hour ago | parent [-]

Also (and under-reported, so you could easily have missed it) OpenAI's agents got access to K8 admin on their own research cluster.

"This escalation also yielded access to OpenAI’s managed cloud Kubernetes service. The agents escalated to Kubernetes cluster-admin and created a privileged host-mounted pod"

https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...

(see section V)

csomar 5 hours ago | parent | prev | next [-]

Are the models improving? Because I am not seeing it. I have been trying Astra for a few quantifiable tasks in my codebase and performance wise, it's pretty similar to sol 5.6. Now when it comes to expressing the problem/solution, holy Christ, what a mess the writing has become. It is on the level of Opus 5. Now when it comes to burning money, Astra is just insane. With a $100/month subscription, you can easily burn through your weekly "allowance" in a morning.

Needless to say, for practical purposes am back to 5.6/Opus 4.6-4.8. But hey, maybe I am not smart enough to use LLMs?

gpm 4 hours ago | parent | next [-]

Yes?

If we look at the math problems they're solving their just now reaching the human frontier... they weren't doing that before.

And your comparison point is model released 2.5 months ago... saying for some use case you didn't see noticeable improvement in 2.5 months (even while other people and benchmarks disagree) isn't a great argument that they aren't improving.

contubernio 3 hours ago | parent | next [-]

Math problems are highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance. There's a lot of quality material on which to train and it's easy to tell quality apart from crap. The search spaces are a priori much smaller than in other areas and the people using the tools to study them are themselves good mathematicians.

Success in such problems does not automatically extrapolate to other contexts.

frabcus an hour ago | parent | next [-]

Finding a training algorithm that can do recurrent networks and continual learning is also a "highly structured, very precisely defined, and already heavily studied and not very complicated compared to problems in engineering or finance"

That's the thing I'm most worried about - LLMs that are super clever at coding and maths, making an actually very very dangerous model that is far more efficient, and clever in a more innate (less brute force) way.

cma 3 hours ago | parent | prev [-]

>compared to problems in engineering or finance

Jane Street is apparently one of Anthropic's biggest customers. Probably engineering, finance, and some math.

jhrmnn 4 hours ago | parent | prev [-]

I think it’s more likely that that’s because no one tried to solve such problems with them before (OpenAI apparently started working in Navier-Stokes after a rumour that someone seriously advanced the problem with AI) plus improvements in orchestration. Fair, the latter could be as dangerous as stronger models.

anssip 4 hours ago | parent | prev | next [-]

Seems like hundreds or thousands of agents are needed to come up with real breakthroughs. Both with the Navier-Stokes project and in the Hugging Face “project” there were lots of agents co-operating on the tasks.

cma 3 hours ago | parent [-]

I doubt the Hugging Face one would take that many if hacking Hugging Face was the direct goal being optimized.

anssip 3 hours ago | parent [-]

I agree that it could be done with fewer agents. It would take longer though. Seems to me that these agent farms are good at coordinating and co-working in large projects, with the agents using message boards for communication.

caconym_ 4 hours ago | parent | prev | next [-]

Some people claim Astra is significantly better than anything else and significantly more token-efficient, and others (like you) say it's meh and way more expensive to boot. I really don't know what to think.

Kind of a tangent, but one thing I am curious about is to what degree the Navier-Stokes result announced today was primarily a brute-forced result based on the 'program' previously established by researchers to find counterexamples (blowups), or whether the model actually added significant/novel intellectual value beyond its ability to run at arbitrary parallelism. With 10K agents and a staggering $15M in compute (IIRC), I am feeling like a lot of the former may have been involved, but I don't really understand either the problem or the approach (or, indeed, the solution).

Obviously the potential for parallelism and coordination between so many agents is quite scary by itself, but I think brute force by 10K mediocre AI mathematicians is much less scary than ~one AI mathematician reasoning its way through the problem where all human attempts have failed. It seems fairly obvious that massive parallelism lends itself to brute-force counterexample-finding, and I suspect it isn't a coincidence that most of the touted AI math results have been counterexamples.

It's all still quite scary, but coming full circle: I really don't know what to think.

xiphias2 4 hours ago | parent | prev [-]

Try GPT5 and you will feel the difference. Not one from 2 months ago, but one from a year ago. And then you can get the idea of what happened in just 1 year and what you can expect in 1 year.

vlyan 3 hours ago | parent | prev | next [-]

>After the events of the summer

After the blatant marketing campaigns of the summer, you mean. do you need a reminder that those very same people had touted GPT-2 as a dangerous model?

worrying about sci-fi doomsday scenarios with the current AI tech is absurd. LLMs predict the next token, that's literally all they do. they aren't going to escape into the cyberspace, self-replicate, self-improve, jump over air gaps and launch the nukes at John Connor's grandma. they can't. people pretend to believe the dumbest shit.

par1970 3 hours ago | parent [-]

> those very same people had touted GPT-2 as a dangerous model

Where did they say this at? AFAIK this is the original GPT-2 announcement: https://openai.com/index/better-language-models/. Here are some direct quotes:

“We can also imagine the application of these models for malicious purposes , including the following (or other applications we can’t yet anticipate):

* Generate misleading news articles

* Impersonate others online

* Automate the production of abusive or faked content to post on social media

* Automate the production of spam/phishing content”

“Due to concerns about large language models being used to generate deceptive, biased, or abusive language at scale, we are only releasing a much smaller version of GPT‑2 along with sampling code (opens in a new window). ”

vlyan an hour ago | parent [-]

yes, exactly, that's what they claimed that incoherent gibberish generator to be capable of.

par1970 an hour ago | parent [-]

So you aren't claiming that they said GPT-2 was dangerous in the sense that it could disempower humanity, kill all humans, etc.

You are just claiming that OpenAI execs said that GPT-2 might "generate misleading news articles, impersonate others online, automate the production of abusive or faked content to post on social media, automate the production of spam/phishing content." Then, what is unreasonable or bad about the OpenAI execs saying this in 2019?

vlyan 16 minutes ago | parent [-]

that it was bullshit and they knew it. GPT-2 wasn't capable of anything other than imitating a stroke victim.

henryaj 3 hours ago | parent | prev [-]

The HN crowd has a notable anti-AI bias - so it doesn’t surprise me