Remix.run Logo
TeMPOraL 10 hours ago

Your comment is already showing the mistaken, poisonous belief of security maximalism, that tries to reinterpret_cast everything into hacks and cybersecurity vulnerabilities.

Most of these things aren't "hacking". They're problem-solving and efficiently dealing with obstacles and random bullshit along the way. This, not "hacking", is what they're making their models "razor focused on".

Problem is, most normal computer use looks like hacking if you spin it that way, especially if you're not willing to question whether some of the roadblocks overcome weren't themselves an error. Not misconfiguration - an error, in humans making a decision to "secure" something more than it should be.

Now, this story was obviously a hack. But it wasn't malicious. It was an LLM given a Kobayashi Maru as a test, and solving it the Kirk's way. 20 years ago, we'd be impressed and be bringing up MIT prank stories.

(Of course, there is a legitimate reason to be alarmed. The flip side of "hacking" and "problem solving" being the same, is that these models can be used to cause mayhem if targeted properly, and they will eventually cause mayhem on their own, because alignment is an unsolved problem. Again, whether something is an obstacle or a sacred line not to be crossed, depends entirely on the values of the agent.)

jayd16 29 minutes ago | parent | next [-]

What is your definition of hacking if it doesn't include using leaked security tokens scraped from the web? Also, kirk 100% cheated.

qsera 6 hours ago | parent | prev [-]

>They're problem-solving and efficiently dealing with obstacles

They are problem solving as much as a falling rock is finding its path down a mountain.

NateEag 4 hours ago | parent | next [-]

I'd readily agree that they may be (probably are?) utterly unaware of what they're doing, with no spark of sapience.

However, I'm a sapient being employed as a software developer for my problem-solving ability.

If you gave me a Kobayashi Maru scenario as a challenge, I would probably come up with the idea of hacking out of the sandbox to find the answer.

If I was in a technical interview, I would probably even ask the interviewer if exploits are fair game, or if that's too far outside the box.

I highly doubt I'd find a new zero-day as quickly as these agents did.

I wouldn't say it's _impossible_ - I've found security issues before.

But I'm not a specialist, and I'd bet against myself.

If the agentic LLMs can consistently achieve something that's a bridge too far for me, then I don't know what to call that other than problem-solving.

I say this as an LLM hater who would push the "Nuke all LLMs" button the instant I had access to it.

Opus 4.8 and 5, at least, don't seem to me to be solving problems by deep, thorough understanding - my employers have compelled me to use Claude, so I've used them a lot to build things, and I constantly find both little and large hallucinations that scream "these are still missing something."

Maybe these new models are actually massively better, or maybe they're just the same kind of system 1 thinking done faster and harder.

The distinction is largely academic, though, for questions like "Can you keep these contained?", "Can you farm out arbitrary programming tasks to them and expect an acceptably mediocre answer?", or "Does it matter if these things are aligned?"

simianwords 3 hours ago | parent | prev | next [-]

incredibly naive comment. as if humans are materially different -- a question for which you would have no response to.

neuroticnews25 2 hours ago | parent | prev [-]

...efficiently?