Remix.run Logo
0x1ceb00da a day ago

[flagged]

antirez a day ago | parent | next [-]

So many mathematicians over the years tried hard and failed, but now Anthropic just for some PR magically did it? And this after LLMs obtaining different math wins? What is your logic here really escapes my understanding.

glimshe a day ago | parent | next [-]

The parent's absolutely nonsensical post highlights how polarized AI (as everything else) is today. I can understand someone being opposed to AI on moral, cost-benefit or productivity grounds. But we're seeing a lot of extremist "AI is good for nothing" posts out there nowadays.

pred_ a day ago | parent | next [-]

I don't think it is nonsensical at all. The author and his collaborator both appear to be bright people, so there's a good chance they had to offer non-trivial insights to guide the LLM, yet it's clearly in the interest of his employer to downplay whatever personal contribution they provided.

Edit: Now the OP is flagged/dead for some reason. You could disagree on their take (calling it a marketing stunt is maybe a bit much), but I think the argument is sound, so flagging seems counterproductive to the discussion.

zaptheimpaler 10 hours ago | parent | next [-]

Believing that AI played a very small role while we know this problem was open for decades with at least a few people taking serious cracks at it is just not a coherent logical position.

soerxpso 13 hours ago | parent | prev | next [-]

That doesn't really even diminish the contribution from Fable, if true. Droves of grad students have been provided the same sorts of non-trivial insights and turned up no results.

glimshe 21 hours ago | parent | prev | next [-]

I'm sure they had plenty of time to think about these insights without the LLM, as well as the many other mathematicians who tried to crack it over the years. Wether the LLM was simply an assistant or solved the problem entirely isn't as important as accepting than the LLM was the essential, previously missing piece in the solution.

witx 4 hours ago | parent | prev | next [-]

That's HN. No criticism allowed of llm whatsoever

minimaxir 16 hours ago | parent | prev [-]

> I think the argument is sound

They got flagged because it is a literal conspiracy theory that assumes bad faith.

latexr 19 minutes ago | parent [-]

> literal conspiracy theory

The strongest words they used were “marketing stunt”. Calling that a conspiracy theory is quite the stretch.

abc42 15 hours ago | parent | prev [-]

If you can call it polarized when a majority of people are just happily using the technology while a minority keeps spreading delusions and hatred.

raxxorraxor a day ago | parent | prev | next [-]

Those silly advertisers do everything for exposure and if that means digging yourself into a niche alleged mathematical theorem to refute it, it is what needs to be done!

Of course it would be really interesting how Claude approached this. Probably with some constraints regarding the input. And it would be interesting what these constraints were.

dist-epoch a day ago | parent [-]

Author probably doesn't want to show the prompt because they are now trying to find a bunch of other counter-examples with the same prompt

regularfry a day ago | parent [-]

The reasoning trace would be far more interesting, and that's not exposed.

YeGoblynQueenne 11 hours ago | parent | prev | next [-]

I mean we have no idea what happened exactly, how Fable was used, how many times it was run, whether earlier models were also tried, what was the prompt, how long it run for, etc etc. All we have to go by is a tweet.

Why not be skeptical about that?

CamperBob2 9 hours ago | parent [-]

What you're asking for is exactly the sort of thing that belongs in, and will appear in, a journal article. There will likely be a preprint on arxiv, so you might keep an eye out for that.

In any case, the fact that it was found by a commercial model means that the unfiltered reasoning trace isn't available even to the original author. So there are aspects of the problem-solving process we'll never see. Even if we did get access to the reasoning trace it wouldn't necessarily be definitive, given how these things work.

Hopefully it'll be possible to get the same solution from an open-weight model like one of the 3T heavyweights that are said to be coming up for release. If so, the chain of thought can be scrutinized in-depth.

YeGoblynQueenne 2 hours ago | parent | next [-]

No, I think it works the opposite way. Until there is an article somewhere that describes what happened, if there is one, all we have to go by is that some guy posted a counter-example for the Jacobian on X, with a vague allusion to using Fable and without any further information. Assuming and guessing anything about e.g. the method used at this point is just raising the noise level.

eru 4 hours ago | parent | prev [-]

The original author works for Anthropic. And even without that, Anthropic might find this important enough to dig up the trace?

jibal a day ago | parent | prev | next [-]

They gave their "logic", such as it is ... and it's utterly irrational.

Note that the "they" who published the counterexample on X is some rando mathematician (Levent Alpöge) working for Anthropic, not Anthropic the organization. He posted the counterexample in a tweet -- reason enough for "not disclosing the LLM chat session". There's no reason to think that it won't provided if asked for, but it hardly seems relevant.

gus_massa a day ago | parent | next [-]

> There's no reason to think that it won't provided if asked for, but it hardly seems relevant.

My guess is that the chat will look similar to a full transcription of a (multi month?) discussion between a few mathematicians. Full of dead ends and stupid errors (bit by the human and Claude) that would be embarrassing. We all know how bad it is, and we prefer to keep it behind the curtain.

0x1ceb00da a day ago | parent | prev | next [-]

> some rando mathematician (Levent Alpöge) working for Anthropic, not Anthropic the organization

Why do you trust a random stranger so much? Will you hand over your car keys to a random stranger? Sharing the chat will take 30s of their time.

> There's no reason to think that it won't provided if asked for

But they didn't provide it.

Liquid_Fire a day ago | parent | next [-]

OK, so the two options are:

A) Claude really produced this counterexample

B) A mathematician working for Anthropic solved a problem mathematicians have been working on for more than a century, and then credited it to Claude for PR purposes

If you believe B is more likely, why would you then believe a proof in the form of a chat log, when said chat log could itself have been faked by Anthropic way more easily than solving the mathematical problem in the first place?

fn-mote a day ago | parent | next [-]

The concern is that there could have been expert knowledge input, whose importance/worth we are unable to evaluate.

I don’t believe a mathematician produced the counter example secretly, but how much did they contribute to the result?

AI isn’t magic, so to evaluate the value delta, you need to know the value of the input.

hgoel 21 hours ago | parent [-]

While I agree that we need the inputs to properly evaluate what this means for LLM capabilities, I don't really believe that the amount of knowledge input matters much for the overall significance of the result.

These kinds of results are interesting for LLMs because mathematicians have been working on them for decades. If the result doesn't already exist, there's no way it's in the training data, and if mathematicians have been unsuccessfully tackling the problem for decades, it is believable that the use of a new tool made the result possible, even if guided by a great mathematician.

YeGoblynQueenne 11 hours ago | parent [-]

Yes, but the question is the extent to which the new tool was guided by the mathematician.

Are we talking a bicycle, powered by a human stepping on the pedals; or a rocket that will fly to the moon on its own with people inside?

Don't you want to know? I mean, doesn't everyone want to know?

hgoel 6 hours ago | parent [-]

Yes, of course, that would be interesting info to have. I just mean that from what we have already we can reasonably infer that the LLM played a role in the result being obtained.

If I am not mistaken, all of the flurry of novel results has come from existing mathematicians. This makes me suspect that the models aren't at the level where just any layman can get results. They require a skilled human in the loop to keep them on the rails and to properly explore the solution space.

YeGoblynQueenne 2 hours ago | parent [-]

Right, that's what I'm trying to understand, the extent to which models need expert guidance. I don't doubt an LLM was used, I just want to know- how.

Izkata 17 hours ago | parent | prev [-]

Weirder has happened: https://mathsci.fandom.com/wiki/The_Haruhi_Problem

InsideOutSanta a day ago | parent | prev | next [-]

If I came up with this counterexample, I sure as hell wouldn't give credit to Claude.

david-gpu a day ago | parent [-]

Indeed, the incentive is the opposite: hiding the fact that they used an AI would boost their own personal brand.

dooglius 19 hours ago | parent [-]

He worked for Anthropic, so such a claim would be very unconvincing

david-gpu 18 hours ago | parent [-]

Because only people who work for Anthropic use AI tools? What is the reasoning here? The guy is a mathematician, after all.

dooglius 14 hours ago | parent [-]

You have it backwards. Only people not working for Anthropic do not use AI tools.

eru 4 hours ago | parent [-]

Yes.

Though not everyone who works in the sausage factory still wants to eat sausages.

jibal 8 hours ago | parent | prev | next [-]

> Why do you trust a random stranger so much?

Why do you make false claims and attack strawmen so much?

dist-epoch a day ago | parent | prev [-]

> But they didn't provide it.

because they are using the same prompt to try finding other counter-examples. they are milking it

8bitsrule 8 hours ago | parent | prev [-]

A few months ago I asked a model how many primes are divisible by 35 with a remainder of 6. It confidently replied 'none'.

Counterexample: 35 + 6.

buzzin__ 7 hours ago | parent | next [-]

Kimi 2.6 gives the answer ""By Dirichlet's theorem on arithmetic progressions, since gcd(6,35)=1 , there are infinitely many primes of the form 35k+6 . So the answer is: infinitely many primes give a remainder of 6 when divided by 35.

buzzin__ 7 hours ago | parent | prev | next [-]

But, if the reminder is 6, they are not really divisible, are they? Try again with a sentence that actually makes sense: "How many primes, when divided by 35, give a reminder of 6?"

jibal 8 hours ago | parent | prev [-]

non sequitur

8bitsrule 7 hours ago | parent [-]

Perhaps ... but the lesson in trusting AI math was worth it.

jibal 3 hours ago | parent [-]

You're missing the point. The counterexample to the Jacobian conjecture is valid regardless of how it was discovered ... elsewhere on this page people are even suggesting that the Anthropic mathematician may not have actually used Claude and came up with the counterexample himself.

Trusting AI math is not an issue here. It's as if you had asserted that no primes when divided by 35 produce a remainder of 6 and the model said that 41 is a counterexample and then you complained about not being able to trust AI math.

P.S. As someone else noted, your AI query was malformed ... no prime is divisible by 35, nor is any integer divisible by an integer with a remainder of 6 -- divisibility implies a remainder of 0. So perhaps the AI simply took what you wrote literally.

0x1ceb00da a day ago | parent | prev | next [-]

[flagged]

j_maffe a day ago | parent | next [-]

Cheating how? The proof is in the pudding...

fn-mote a day ago | parent [-]

On one hand, yes.

The parent wants to be skeptical… nothing wrong with being suspicious of marketing claims, right?

I would be a lot less impressed if I found out the session was guided by an expert in the field who already had a good idea of where to look. For example, I don’t believe that the results Terrance Tao gets from an LLM are comparable to what I am going to get.

I’m not even saying I’m a skeptic. Just that there’s nothing wrong with keeping your eyes open and asking for details.

ceejayoz 19 hours ago | parent | next [-]

> I would be a lot less impressed if I found out the session was guided by an expert in the field who already had a good idea of where to look.

Why? The conjecture stood for over a hundred years. Plenty of such experts have tried.

"Fable makes experts able to solve things" is still a big story. I don't need to personally be able to do it for it to be a big story.

aeon_ai a day ago | parent | prev [-]

Because AI is only impressive if the average Taco Bell employee can guide it to address niche domain topics?

It seems obvious to me that you’d need someone to point at a thing and say “pay attention to that”, as a baseline, to have any results at all with the current architecture and technology

Redoubts a day ago | parent | prev [-]

Dude what?

William_BB a day ago | parent | prev [-]

There's a big difference between one shotting a counterexample using AI and using AI to find a counter-example by brute-force.

Both are impressive, of course, but they're hardly comparable.

eru a day ago | parent [-]

I'm not so sure. People had been trying to use brute force to find counterexamples before.

William_BB a day ago | parent [-]

AI is great at reducing the search space and using human-like reasoning (in a brute-force way) to carry out the brute-force search. I'm not surprised by this result. This is exactly what AI should excel at, with human guidance.

igravious a day ago | parent [-]

"using human-like reasoning (in a brute-force way)"

that's self-contradictory -- what brute force means is doing an exhaustive search of a search space (brute forcing it)

using human-like(?) reasoning means cutting down the search space by having some sort of insight or intuition which allows you to prune branches from the entire tree

YeGoblynQueenne 11 hours ago | parent [-]

LLMs don't search trees. They generate plausible proofs and a human has to check it's true. Repeat until.

tptacek 9 hours ago | parent | next [-]

That's not what happened here. This isn't a proof; it's a counterexample. The model was perfectly capable of verifying its correctness. You could have verified it by hand if you wanted; the verification is trivial. Finding it was the hard part.

YeGoblynQueenne 2 hours ago | parent [-]

>> The model was perfectly capable of verifying its correctness.

It's an LLM. It can't do that.

astrange 10 hours ago | parent | prev [-]

Agents can generate formal proofs that are checked with an oracle like Lean and can run in a loop.

YeGoblynQueenne 10 hours ago | parent [-]

What are the search algorithms then? DFS, BFS, A*, etc, can you name them?

tptacek 9 hours ago | parent [-]

Why are you asking them which basic computer science graph traversal algorithms a frontier model used?

YeGoblynQueenne 2 hours ago | parent [-]

OP:

>> using human-like(?) reasoning means cutting down the search space by having some sort of insight or intuition which allows you to prune branches from the entire tree

So according to the OP there's a search of a tree and it also uses pruning btw, so I want to know what search they mean. Why are you asking?

abc42 15 hours ago | parent | prev | next [-]

This anti-AI sentiment is getting borderline insane.

greyw a day ago | parent | prev | next [-]

The next few years are going to be difficult for you

gonzalohm a day ago | parent | next [-]

If we don't ask for proof when someone claims something, the next few years are going to be rough for everyone. We can't trust anything nowadays

mcosta 18 hours ago | parent [-]

They have provided a literal mathematical proof.

gonzalohm 15 hours ago | parent [-]

Proof that an LLM did it and that this is not yet again another marketing stunt.

Remember when Anthropic wouldn't release Fable because it would be "the end of cyber security as we know it"? Yet here we are

minimaxir 15 hours ago | parent | next [-]

This proof was posted by an individual researcher, Anthropic has not used it in any of their marketing.

Additionally, posting this during the World Cup would be the most inefficient way to do marketing.

latexr 12 minutes ago | parent [-]

> posting this during the World Cup

The World Cup is over. This tweet was the day after.

tptacek 13 hours ago | parent | prev [-]

Explain how the marketing stunt would work here? Tell me a story about how OpenAI spends money to pull off a similar problem by solving one of Smale's other unsolved problems. 80 years of mathematicians were unable to disprove the Jacobian Conjecture.

Is the marketing stunt that Anthropic has secretly built a world-class mathematics research group?

adw 4 hours ago | parent [-]

I mean, they kind of have, but mostly by finding an infinite money glitch.

0x1ceb00da a day ago | parent | prev [-]

Why?

minimaxir 16 hours ago | parent [-]

There will be more and more mathematical proofs provided by LLMs as they improve. If you assume every one is just a marketing tactic, you'll drive yourself insane.

anxoo 16 hours ago | parent [-]

you can believe a small, irrelevant lie and still go on happy. a very big lie sucks in more and more of your reality until you have to disbelieve your own eyes and ears.

anxoo 16 hours ago | parent | prev | next [-]

if you ask the chatbots for "list of top unsolved math problems", the JC comes in at a ranking of around #10 - #20. what, a problem that's been unsolved since 1939 was cracked because anthropic has an underground sweatshop of math Phds cranking out research, just so that they can slap "made by AI" on it? hell, maybe lizard people did it.

UltraSane 12 hours ago | parent | prev [-]

Finding a counterexample that humans have failed to find for 85 years is a marketing stunt?