Remix.run Logo
overgard 2 hours ago

Gary Marcus' has a good take on this:

https://garymarcus.substack.com/p/openais-amazing-but-vastly...

https://garymarcus.substack.com/p/two-critical-updates-re-as...

Not that there isn't something interesting in here, but lets be clear that we don't have enough information to evaluate this properly. And as always with these labs, BS takes a lot more energy to refute than it does to spread.

scarmig 2 hours ago | parent | next [-]

It's worth reading Marcus' first line, for the naysayers and flaggers on this post:

> Astra, a new model that OpenAI is testing internally, is amazing. No denying that.

aaroninsf 2 hours ago | parent | prev | next [-]

I find Marcus on this, something approaching sophistry and rhetorical showmanship in service of maintaining an ideological position, for reasons unrelated to the nominal intellectual clarity.

To sharpen that, I think he's (obviously) interested in maintaining his own brand as "thought leader" and this necessitates de rigeur defense of particular postures.

Sometimes this is easy because the facts warrant it; other times, a bit of rhetorical license is required to preserve nominal coherence and (at least, for the moment) hold certain lines.

This is one of the latter cases, and it's not subtle.

One of the celebrated properties of many intellectual advances or inventions in whatever domain is precisely that it appears obvious in hindsight. It is quite cynical to leverage consensus distrust of large AI players, warranted but also a popular social construction, to insinuate that these are not "real" advances or "real" hard problems, on the grounds they were in some sense cherry-picked.

Identifying the problems amenable to strategies on the table and intuitions (sic) about where bridges might be, is exactly the discerning work that is the core driver of almost all prior progress, but for celebrated accidents and flashes of insight. Anyone working in any challenging discipline knows that those are celebrated and told around campfires precisely because meaningful durable results arising like that is so uncommon.

These two articles make me think of nothing so much as my own durable reaction to the creeping goalposts of AI critics generally: that they often seem to me not unlike a water color cohort scoffing and jeering at the horse, because it got a D on its tensor calculus exam.

Marcus should be on guard against his own cynicism and take care that his assumptions do not prevent clear sight.

HardCodedBias 2 hours ago | parent | prev [-]

"Gary Marcus' has a good take "

I think that is an oxymoron.

neta1337 2 hours ago | parent | next [-]

How so? His predictions were accurate so far

energy123 2 hours ago | parent [-]

No they were not. These were his 5 predictions in 2022:

""" 1. By 2029, AI will still be unable to watch a movie and accurately explain the characters, events, conflicts, and motivations.

2. By 2029, AI will still be unable to read a novel and reliably answer questions about its plot, characters, conflicts, and motivations beyond what is stated literally.

3. By 2029, AI will still be unable to work as a competent cook in an unfamiliar kitchen.

4. By 2029, AI will still be unable to reliably create more than 10,000 lines of bug-free code from natural-language instructions or interaction with a nontechnical user, excluding simple assembly of existing libraries.

5. By 2029, AI will still be unable to convert arbitrary mathematical proofs written in natural language into symbolic form suitable for formal verification. """

There's still 3 years to go and he's already wrong on 4 out of 5.

sweezyjeezy 2 hours ago | parent | next [-]

Well I don't typically side with GM, but playing devil's advocate:

1. still not wrong? Unless it's just feeding the audio or screenplay I don't think you can feed AI a full movie in a single context window yet?

2. Not sure, but can you prove this wrong? Can you feed a full, unseen new book and get that kind of answer?

3. Not wrong.

4. I think he'd probably pull you up on 'bug free' - I don't think that frontier models can reliably write 10k LOC without _any_ bugs typically (not that humans can do this either).

Philpax 19 minutes ago | parent [-]

4. I think they can, especially if the problem statement is well-specified and, importantly, autonomously testable. Of course, specifying a problem that meets these requirements is non-trivial, but the claim requests _a_ counterexample :P

ducktective an hour ago | parent | prev | next [-]

Do LLMs generate deterministic or trustworthy answers?

an0malous 2 hours ago | parent | prev [-]

> There's still 3 years to go and he's already wrong on 4 out of 5.

Have these been tested or are you just guessing?

overgard 2 hours ago | parent | prev [-]

It can be annoying when someone you disagree with is frequently right!