Remix.run Logo
giancarlostoro 5 hours ago

I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?

I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.

> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.

- Sam Altman on AGI

chrsw 4 hours ago | parent | next [-]

I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.

azan_ 3 hours ago | parent | next [-]

I can call coworker right now and have conversation so frustrating that I wish I was talking to machine instead.

2 hours ago | parent | next [-]
[deleted]
dotancohen an hour ago | parent | prev [-]

Hell, I could do that with the wife at home.

usef- 3 hours ago | parent | prev [-]

I think they're claiming it's achieved by text models, not voice models, fwiw.

petilon 2 hours ago | parent | prev | next [-]

Here's another definition of AGI from Sam Altman:

https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...

Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?

Kevin Roose (New York Times): I probably would, yeah. Would you?

Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.

hackerbrother 2 hours ago | parent | next [-]

If the new AGI benchmark is "be Einstein/Feynman" then we've hit AGI.

jmalicki 28 minutes ago | parent | next [-]

What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?

The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.

block_dagger 37 minutes ago | parent | prev [-]

Wouldn’t that mean producing novel work like relativity and QED?

jmalicki 10 minutes ago | parent [-]

I would maybe argue that Einstein was the most LLM-like of great thinkers.

A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.

A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.

That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.

throwawayq3423 2 hours ago | parent | prev [-]

What is novel physics?

astro1234 2 hours ago | parent | next [-]

I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?

glenstein 28 minutes ago | parent | next [-]

This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades.

So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute.

adastra22 an hour ago | parent | prev | next [-]

There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever.

verelo an hour ago | parent | prev | next [-]

i wonder if we could train a modal, and omit all data prior to 1899, and see what happens?

astro1234 44 minutes ago | parent [-]

Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.

ShinyLeftPad an hour ago | parent | prev [-]

as always, as good as its prompt...

senderista 2 hours ago | parent | prev [-]

I assume solving one of the major open problems of physics?

adastra22 2 minutes ago | parent | next [-]

That’s already been done. I know of at least one novel result contributed by Claude to frontier physics. I’m sure there is more.

type_enthusiast 2 hours ago | parent | prev | next [-]

Would this be possible without it being able to run novel real-world physics experiments autonomously?

(Note: I am not suggesting we let it do this. Please don't, in fact)

petilon an hour ago | parent | next [-]

AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later.

auntienomen 13 minutes ago | parent [-]

Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition.

adastra22 an hour ago | parent | prev [-]

Why not?

colordrops 2 hours ago | parent | prev [-]

For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms.

petilon 2 hours ago | parent [-]

Solving "open problems" will push the field forward.

refulgentis 2 hours ago | parent | prev | next [-]

If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to.

adastra22 a minute ago | parent [-]

It’s better to call a spade a spade.

lenerdenator 4 hours ago | parent | prev | next [-]

I wonder if Altman's definition also includes taking on the same liability as a coworker would.

Probably not.

giancarlostoro 4 hours ago | parent [-]

Would 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford.

lenerdenator 4 hours ago | parent [-]

And that's the rub, isn't it?

If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit.

If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.

alex0015 4 hours ago | parent [-]

What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider.

If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.

Avicebron 3 hours ago | parent | next [-]

The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task?

alex0015 2 hours ago | parent | next [-]

It's not really a dream though? You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks. Another person can do the same thing but also keep a human in the loop. You and the other person may compete in the market for whatever your product or service is, and you both might do very well or one of you might do better than the other because of a whole host of different reasons. Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor, nor does it require them to offer the LLM service with a guarantee that it will never make any mistakes. So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money.

kbelder 3 hours ago | parent | prev [-]

The person responsible at that point is the sucker who fell for the dream.

degamad 3 hours ago | parent | prev | next [-]

> What do you mean by bearing no real responsibility for its actions?

If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.

It it makes a mistake and deletes your website from AWS, who is responsible?

If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible?

alex0015 2 hours ago | parent [-]

In all of these cases, it's you. It would be the same if you downloaded an open model, ran it locally, and it happened to make the same catastrophic mistakes. The consequence to the provider is that if they offer a product that does these things, people don't buy the product.

In general, the person whose job it is to provide the company with a working, non-adult website and not hack into other websites is the one who would receive consequences for failing to meet those expectations.

lenerdenator 2 hours ago | parent [-]

The problem I see here is that ultimately, you'll have capital wanting to replace workers like others have said, and have someone roughly equivalent to a manager or vice president driving teams of agents to achieve business outcomes.

These tools can push out more results than a human can hope to evaluate in a business-sensitive, or even realistic, amount of time. You have to take it at its word that it did things right, and there's no real fear of failure or consequence on the behalf of the agent.

zapkyeskrill 17 minutes ago | parent [-]

Of course, but "capital" is no stranger to risk management. I'm sure we'll see some spectacular failures, but most will handle this just fine.

fooqux 3 hours ago | parent | prev [-]

If future jobs are simply reduced to liability scape goats (or more appropriately reverse centaurs) for management to pin things on then I'm taking up goose farming.

ricardobeat 3 hours ago | parent | next [-]

I'm afraid you will be pushed out of the goose farming market by these new ultra-efficient farming bots.

lenerdenator 2 hours ago | parent | prev [-]

That's more-or-less what you are now, especially if you work at a company like Meta where 1) the guy at the top holds majority control of the company's shares and 2) keeps making massive, expensive mistakes either by accident or design.

ACCount37 2 hours ago | parent | prev | next [-]

ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them.

In practical terms? They usually get solved with a bigger badder LLM. "New ideas are needed?" Nah - ten times the params, ten times the test time compute.

ARC-AGI-3 was more of a failure in that regard than -1 or -2, because even on day 0, an off the shelf LLM with a harness could get 50%+. And messing with evals by forbidding "LLM with a harness" from scoring? Yeah no, that was just bad.

butterisgood 2 hours ago | parent | prev [-]

People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".

An AGI wouldn't struggle with that.

taneq 4 minutes ago | parent | next [-]

These are like saying someone isn’t human because they have a speech impediment or an auditory processing disorder.

AGI doesn’t mean infallible, it just means it can have a reasonable crack at things it hasn’t seen or done before.

nearbuy 2 hours ago | parent | prev | next [-]

The last version to fail on those questions was GPT 4.5.

Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".

the_gastropod an hour ago | parent [-]

Nooo. I just asked Claude (Sonnet 5 Medium) how many I's are in assassin, and it said 2. Granted, it got several words correct before this. But no, they still aren't great at this letter-counting thing.

nearbuy 6 minutes ago | parent | next [-]

I don't think we should count the lower tier models if we're discussing what the top ones are capable of. No one was suggesting that Sonnet is AGI.

SmashDan 8 minutes ago | parent | prev | next [-]

Anyone know why they aren't good at this?

steelframe 22 minutes ago | parent | prev [-]

Meanwhile Qwen3.8 27B got both the 'f's and the 'i's questions right.

jasondigitized an hour ago | parent | prev | next [-]

I'm going to go out on a limb and guess that there are plenty of savants who can't tell you how many r's are in strawberry.

slidehero 2 hours ago | parent | prev [-]

> gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".

this has been debunked too many times to bother rebutting. they struggle with those things because of the way they are.

it's completely irrelevant.

phlakaton an hour ago | parent [-]

If it allows you to distinguish readily between human intelligence and computer intelligence, it seems quite relevant to the question of whether computers have achieved something akin to human intelligence.

It may not be useful for anything else, but at least it can say that.

slidehero an hour ago | parent [-]

which just brings us back to the whole birds vs planes thing.

turns out that flapping wings is not the right way to unlock human flight.

computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.