Remix.run Logo
dschuessler 2 hours ago

Somewhat related: In 2018, Google DeepMind had already created AIs that were capable of beating professional gamers in StarCraft 2 (the sequel to Brood War): https://www.youtube.com/watch?v=cUTMhmVh1qs

qerghnui an hour ago | parent | next [-]

AlphaStar beat one retired professional by cheating.

AlphaStar won a showmatch against TLO. TLO was never one of the strongest players in the world. He had been retired for over three years by the time of the match. Google set the rule that their system would have human-like mechanics, but it played several times faster than any human, never issued a wasted action, had an inhumanly fast reaction time, issued commands with perfect accuracy using an API, and could see the entire map at once.

It was later released to the open ladder with more human-level mechanics. Even strong amateurs regularly beat it. I have beaten it myself. It was strong, but not even close to the level of the strongest human players. It had obvious and easily-exploitable deficiencies in strategy and building placement.

I think even the cheater version would have lost handily to Serral or any of the strongest players.

(It apparently beat MaNa as well as TLO, but those matches were never released to my knowledge. I see no reason to assume Google cheated less flagrantly in private than they did in public.)

benswerd 2 hours ago | parent | prev | next [-]

I predict LLMs will reach superhuman level and beat even that model in the next 12 months

orbital-decay 2 hours ago | parent [-]

Starcraft is APM-dependent. Unless the latency will improve greatly in frontier reasoning LLMs (which is unlikely), it will remain a bit like knitting with an excavator.

benswerd 2 hours ago | parent | next [-]

I predict latency will improve greatly in the next 12 months to more than 4x speed on current frontier tasks

loeg 2 hours ago | parent [-]

Yeah but Starcraft needs, like, 10-20x the APM these agents are doing.

benswerd 2 hours ago | parent [-]

I’m not convinced a lot of it can’t be solved with code mode.

Marine staggering for example seems like an ideal code mode task.

loeg an hour ago | parent [-]

Yeah. Some of it may just be "thinking" less rather than faster token generation.

jamiequint an hour ago | parent | prev [-]

Should be possible to play this with Jev.

adsfgoinoi 2 hours ago | parent | prev [-]

[dead]