370 points | 25h ago | Discuss on Hacker News | Back to Radar
> mommy, I heard you got cooked! I heard that a robot solved the math problem you worked on for your whole career! OOF!
My 8yo talks exactly like that. I could totally imagine him saying this, the same way, at the dining room table.
I asked ChatGPT "pretend you're an 8/9 year old today. how would you insult your mom about having her job be replaced by an AI?", and the responses it offered were:
> “Mom, AI took your job because apparently even robots were like, ‘Yeah… we can do this better.’”
> “Mom, congratulations! You got replaced by a computer. Even Siri has a job now and you don’t!”
> “Mom, AI took your job? Dang. I guess even a robot looked at your work and said, ‘I got this.’”
> “Don’t worry, Mom. You can still be useful… like teaching the AI how to make my lunch.”
All of these seem to have a vaguely Millennial flavor, aside from being pretty awkward and mechanical roasts. Trust the children and linguistic drift to be the best AI detector.
i used the free google ai: (deleted the examples... but they were vaguely close to what i hear my grandkids say.)
edit: neat, insta-flagged despite hundreds of non-ai comments that have never been flagged. i would have thought that hn would use some heuristics in their ai detection but i suppose not.
However, there is the fact that you had to know how to poke the machine so it plucks the right vocabulary out of the training data.
There is very clearly no theory of mind: no inherent internalised modelling of how a human of a particular age thinks, speaks and what they do or do not know. These are the most obvious cracks in the “LLMs are (or will be) the superintelligence” narrative.
Reminds me of the memes with the little girl making astute comments about the patriarchy to her father.
(Incidentally, if you asked me a few years ago I'd have predicted that reliable AI detection was impossible. I still suspect that it's impossible in general and that Pangram would break as soon as companies figure out how to make LLMs stop all sounding the same - but until then, it'll work fine.)
We are quickly moving to a world where all symbolic and numeric reasoning for economic purposes is performed by AI.
I don't know what that means in practical terms, but I agree that's the issue.
>I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.
Also, Ted Chiang's 2000 short story "Catching crumbs from the table" <https://np.reddit.com/r/singularity/comments/1wzu5gf/this_mi...>.
I guess it deserves respect as progress, but it just rubs me the wrong way. Like the machine did the absolute minimum to beat the previous mark.
Now, there are some critiques you can have of this. Namely, it is possible that these novel algorithms have significant trade-offs that make them almost never worthwhile in practice. "Fast" matrix multiplication algorithms are typically of this form. So perhaps this all points towards a deficiency in big O notation, which can be deceptive. But, for people who care about optimizing asymptotic complexity, it is still interesting.
https://hideoushumpbackfreak.com/algorithms/algorithms-stras...
I know, I know. Don't anthropomorphize AIs. They hate it when you do that.
-Boaz Barak
― Alfred North Whitehead, "An Introduction to Mathematics" (1911)
> Basically the paper is so horribly written that it’s impossible to read it without AI help
That's interesting and haven't seen this in all the coverage of this event.
It sounds horrible to wade through - like trying to understand someone else's messy code that still produces the correct output.
It's worthy of note that most humans, do not find most mathematicians understandable. As is frequently demonstrated in Calculus classes. Therefore it is arguable that even human produced results are not generally human understandable.
Why would anyone believe this (also) is not simply example N+1 of this is the worst it will ever be, as opposed to recognizing this as what will almost certainly prove to be an awkward moment, soon to be replaced by another order of magnitude of cleaner, clearer, more intelligible, etc.?
Ximm's Law: every critique of AI assumes to some degree that contemporary implementations will not, or cannot, be improved upon.
I don't think anyone is saying it can't or won't get better, but the question is how much better, on what timescale, and are there fundamental parts of the problem which will remain extraordinarily difficult to improve?
The comment I was responding to suggested a guarantee of an "order of magnitude" jump right around the corner. There is no guarantee of this, and if you view doomers as fools for having doubts, then we ought to look upon the folks who are sure of this sort of progress in the same way.
Already the unit distance proof was substantially human-edited (per Thomas Bloom). Then with the ten problems from Astra you started getting the citation issues. Then Navier-Stokes was a rushed 160 pages with barely any citations, and some of the related papers were called (by their "authors") the ugliest mess they've ever seen.
And now here we are. At least it seems that mathematical ability and communication with a mathematical audience are independent skills, and progress in the first does not imply the second.
This doesn't surprise me much, given two analogies: 1) many smart people are nonetheless horrible lecturers. (You can't quite get the opposite extreme, since to explain math well you have to be able to do it.) 2) AI writing in general hasn't improved. The models have annoying verbal tics ("honestly") and have no sense of which part of what they say is obvious and which is relevant.
This looks like an AI IPO PR powerplay, because at this point the proofs haven't been checked and it may not be possible for a human to check them - because proofs should be clear, not horribly written and noisy.
The noise is suspicious because it's the difference between brute forcing and cognition. A human proof won't just be logically correct, it will be cognitively distilled and coherent. It may still take years to understand it, but the logical flow will be straightforward, not obfuscated.
You want the path through the maze to be as short as possible and the map to be as clear as possible.
This sounds like the opposite. There may be a genuine path through the maze, but if it's too convoluted and takes too long it will be impossible to confirm.
I think the next step is to demand that proofs either be human-scale or they prove that a human-scale proof is impossible and the machine proof is as good as it gets.
I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
> This looks like an AI IPO PR powerplay,
Interestingly, the post has actually also an argument for this:
> Experience has shown that, even now, there will still be people explaining in patronizing tones why none of this is real and none of it counts. If such people were capable of being impressed by anything that happens in the empirical world, of updating on anything, they would’ve already been impressed and already updated several years ago, long before things had reached the point of an actual Mathocalypse.
> So, they’ll say, maybe the alleged solutions are not solutions at all, but just “AI slop.”
There's clear benefit in a babelfish that can coordinate disparate efforts, the only problem with the current iteration is giving credit to said efforts.
Google went quite far down the road to hell, but stopped short of taking credit for websites' content since the company understood that poisoning the well only goes so far. At this point, one can safely conclude that _Chat_GPT was an intentional attempt to squeeze out more data once they mined the internet dry.
Obviously they stole them from the Proof Fairy.
It's grapes so sour they could etch metal.
Also from what I can tell from the few fields I understand, the proofs aren't that long or complicated they are just terribly written.
The entire US federal budget for math research is something like $100M annually. And mathematicians in other countries are hardly making bank either. How does one reconcile how the market has historically valued mathematics with the cash-strapped frontier labs ploughing so much money into that enterprise?
Why should that make a material difference to the IPO? Because of the vibes, and investors are indeed all about the vibes.
The second paragraph is specific to OpenAIs behavior.
Note that the cool thing about rationality is that it does not depend on transitivity. If what companies do to attract investors works, then it is in fact rational of them, regardless of the rationally of investors.
If the models can do this, they're almost certainly good at just about everything, because the reasoning and creativity required to solve these problems will translate. And even if they were only good at this stuff, that's still a tremendously valuable thing, because quantitative reasoning and analysis is the bedrock for many, many industries.
oAI is gunning for the largest IPO in history at this point, and they might actually get there.
The "then" in your "if-then" bears a heavy load. Why would society assign so little economic value to pure mathematics if the skills for proving math theorems translate to massive value in "just about everything"? Would you expect top mathematicians to cure cancer if you transplanted them from the math department to a medical research lab?
I'd assume that the majority of people who study math take their skills and move onto some related STEM career that isn't pure math. Academia is incredibly small and competitive.
All of STEM relies on mathematical analysis, and new models are now superhuman at that. And yeah, I'd go a step further and say that the reasoning and creativity required to solve cutting edge math problems probably does translate to other tasks like interpretation of the law, or medical diagnosis, or accounting, etc., for the same reasons that I think most top tier mathematicians would excel at those tasks were they so inclined.
I would argue that "good at math" was a short hand for "good at X" because mathematicians were historically good at engaging with very complex ideas, distilling them and coming up with precise and concise answers they could validate by themselves.
Given that AI solutions are described as "psychedelical" and they rely on the outside source to validate the result, I don't think the same logic could apply to them.
https://en.wikipedia.org/wiki/Inter-universal_Teichmüller_th... seems like a counterpoint, but IANAM. (I am likely cherrypicking the far end of the bell curve re: straightforward here)
>Mochizuki and a few other mathematicians claim that the theory indeed yields such a proof but this has so far not been accepted by the mathematical community.
Proof can't be understood, proof doesn't matter.
Someone at OpenAI, please, work on this.
There is ongoing AI assisted work on this, and precisely the problematic gap in the theory (around Corollary 3.12).
As of now, a required step appears to definitely be missing, but it's not clear if this is a genuine / unrepairable defect.
https://zen.ac.jp/news/zmcpostevent0331e
Developers and people in CS in general seem to have gotten used to the idea that most productive SWEs don't need to exactly know how to produce assembly or trace every branch prediction or even most of the optimization the CPU (or even their compiler) is running. Mathematicians will get there.
But it doesn't follow that these proofs - or any proofs - are automatically on the far side of that limit.
The human usefulness of a proof depends entirely on its human legibility. Much of the value of proofs is in inventing new techniques and concepts and having new insights into relationships. Occasionally you get some game changing insight into practical physics or engineering. But that's rare.
Without that, proving or disproving a conjecture is an excuse for new and original thinking.
Compilers are not the same problem. The point of code is to produce reliable-ish consequences from various possible inputs. It's not a creative exercise in logical consistency, which is what maths proofs are, ultimately.
That's because there's a bunch of people who work on that stuff. Just because web developers don't care about it doesn't mean it doesn't exist. Utter magical thinking.
Chow, T. Y. (2008). A beginner’s guide to forcing (arXiv:0712.1320). arXiv. https://doi.org/10.48550/arXiv.0712.1320
> “All mathematicians are familiar with the concept of an open research problem. I propose the less familiar concept of an open exposition problem. Solving an open exposition problem means explaining a mathematical subject in a way that renders it totally perspicuous. Every step should be motivated and clear; ideally, students should feel that they could have arrived at the results themselves. The proofs should be “natural” in Donald Newman’s sense [13]:
> This term . . . is introduced to mean not having any ad hoc constructions or brilliancies. A “natural” proof, then, is one which proves itself, one available to the “common mathematician in the streets.””
https://en.wikipedia.org/wiki/Nothing-up-my-sleeve_number
?
In other domains I have seen first hand overwhelming evidence of how things that cause the AI to make mistakes also cause humans to make the same mistakes.
I wonder if the proofs being produced that are hard for humans to interpret are also hard for other LLMs to interpret.
In other words, I wonder if humans are still much better at compressing understanding into proofs than the best LLMs, and what it will take for LLMs to exceed them.
It kind of an explicit example of how the LLMs can be materially less intelligent than people, but still be more productive through scaling, and yet they also can't replace people because they are a categorically different kind of intelligence. It's like all the AI debates compressed into one example showing countwr-intuitive answers.
My mental model - for better or worse - is that intelligence has both shape and area, and LLMs are orders of magnitude smaller area but very different shape, and they have more intelligence area in the kind that humans have lesser of.
So yes very different, more in some material ways and lesser in others, but in total intelligence are still orders of magnitude lesser.
My gp comment was acknowledging that when you scale up many instances / brute force problems it confuses that "total area" claim a bit.
To follow the anthropomorphization... 1000 toddlers may have more total intelligence than a grown man, but does that matter?
The problem with these discussions probably/usually fold into differing/loose definitions of intelligence.
perhaps the chat-based ux has sort of fooled us into comparing these things to human intellegence. we don't really do this with chess, or other forms of ai, nor computers at large.
Isn't it fairly established that (generally [0]) manually written / optimized skill files perform a lot better than generated ones? Meaning that yes, this likely does hold.
[0] or to be specific, that the pecking order is: ai generated < human co/written < hyperoptimized for the specific model via some convergence process
Who do we demand this from? The AI companies? Or the mathematicians who are worried they will have nothing left to do?
As the old saying: great claims require great evidence.
In 1976, the proof of the Four Color Theorem was controversial because it was done with a computer examining over 1000 cases by brute force and was essentially not comprehensible by humans. But mathematicians ended up accepting it. So mathematics has a 50-year precedent of not requiring human-scale proofs. How is the current situation different?
(Disclaimer: Apologies if this sounds dismissive or argumentative. I genuinely think that the Four Color Theorem should play a role in these discussions and suspect that many people are unaware of the controversy over it.)
As AIs become smarter and smarter, there will be no amount of clarity that will make more complex proofs understandable to humans - this is an inevitable effect of the cognitive capacity gap.
Complaining about bad style can make some sense now (I disagree anyway), but it's an argument that will be dead shortly.
Maybe no human will fully understand a future proof, but they could fully understand a little piece of it. And many humans in aggregate could understand it, each with their own little piece.
There are many problems that aren't "interesting" in a mathematical sense but can still have valuable proofs. The entire field of formal methods in software development is mostly about this kind of problem. If I want to be sure that no input can cause out-of-bounds memory access, or that my superoptimizer found the lowest possible cycle count, I don't care if it's elegant, I just need some machine-readable proof I can run through my trusted verifier.
If this is true, then what's really the point of math? A lot of math is actually useful. A proof regarding cryptography, for example, would have practical application even if not understandable by humans.
Given the LLMs are struggling to explain themselves clearly, but also that these explanations cleared up somewhat by having an expert using one to query some of this research, but also the LLMs can solve problems faster than humans can read the proofs, it is possible for this category to have both examples of things that can be rendered in a human-comprehensible form, and also examples where it is not.
As an analogy: any human can check any single arithmetical calculation from a computer, that's not even particularly difficult. But a Raspberry Pi Zero can do those arithmetical calculations so fast that even if literally every single human was trained to do this at the level of the current world record holder, humanity as a whole could not keep up.
We have such tools for a large number of events that we're incapable of witnessing or understanding in "raw form".
We can't possibly read through thousands of raw records of recordings of individual's heights and make sense of it. But we can use Excel to calculate the average in seconds. We can chart the results and the image is perfectly comprehensible. We can trust the average calculation and the image because we trust the mechanism for transforming the data. We have what I would call a trusted path of provenance. We don't need to nor do we want to read the raw data.
Now we need things like this of the 2nd order. We need mechanisms of transforming trusted paths of provenance into things we can look at and easily verify.
We can do formal verification of code that would be hard to do by hand. We need formal verification tools for formally verifying formal verification tools. Eventually we will need multiple layers of this. As long as the chain and reasoning is intact, we should be alright.
It's kind of like with a horse. A horse is much stronger than us, but we can control it by pulling on two strategically connected pieces of rope.
There are mappings between complex numbers and 2D matrices allowing problems to be solved in either domain.
There's research into The Langlands Program looking to connect number theory and harmonic analysis.
There's research into Category Theory looking to define core concepts and relate them do different fields so that results in one field can be applied to another due to equivalence.
> I suspect that's possible without tripping over the halting problem. (But I can't prove it.)
I doubt that it is. For the language of proofs to be powerful enough to be able to express an arbitrary proof it would have be Turing Complete. Proving that a proof is the smallest proof of a given concept (i.e. there is no smaller human understandable proof) would then be proving the minimality of a program in a Turing Complete language.
This is basically how all AI approaches appear to work. They solve the problem you give them (in many cases) but in a very over-complicated way.
It's like they can't step back from the problem and realise that they need to simplify to make it work better. Nope, just keep hammering more code/proofs against the problem and eventually you'll hit the goal.
RL has a lot to answer for, I guess.
There is nothing wrong with that approach if it accomplishes the goal faster, i.e. you can work at the speed of an AI.
There’s not really a clear distinction between these things, in my opinion. Problem solving (and intelligence?) is a mix of search and compression. We like solutions that are elegant (high compression, simple search). But often what appears elegant to some is harder to appreciate for those without the same background knowledge or even the same amount of mental bandwidth (if you’ve ever worked with someone simply much, much smarter than you, you may intuit this!).
This basically describes every single PR at work for the past year. Diffs of 10k+ paragraphs of comments saying nothing. Just rubber stamp and move on, nothing else you can do.
Meanwhile you can use the model to help you out as Scott comments "Just now, however, Dana tells me that she’s been asking Astra all day to explain the new proof of the UGC to her and it’s been doing an amazing job and she’s starting to understand the construction."
If a journal receives a paper that is unreadable, it is outright rejected with the comment that it should be made readable, regardless the results in that paper. What is the point of having results if you are unable to convince your audience of those results? You could just as well just sit on them and never communicate them. What is harmful about this process?
Which is why it is obvious post-publication peer review is the only sane way to handle this. Formal peer review is not even remotely up to the task here.
Maybe there’s some variance or it depends on the reader. That said, like with a lot of other complaints about AI: human papers can be poorly written and poorly explained too. That’s always been the case. And sometimes a proof is just complicated and hard to expose nicely.
That's what a lot of us do nowadays, but instead of maths, it's code that looks sensible on the surface, but when you try to understand it's some kind of "alien" logic , names don't really make sense, etc.
Makes you wonder if AI can build upon such proofs.
If the AI cannot create proper abstractions, then how can it build a tower of abstractions?
Also, AI has limited context. At a certain point proofs may become so complicated that an AI cannot keep most of it in memory and will not reuse it for new proofs.
so then, if you don't understand the proof, how do you know it's a proof?
Welcome to lots of (most?) code PR's in the last year. Though for software at the PR level it has gotten better with the latest models.
And you're right, it's not really a theme around here, but there aren't that many mathematicians around HN. Check one of the maths forums and it'll be a common point of complaint.
When all you value is being the first, you have not time producing useful papers.
I think OpenAI wants the outputs to be unreadable. If outputs "too advanced for human comprehension" become the norm, then every interaction requires tokens. If someone can buy a day pass, get what they need, and leave, then there's no recurring revenue stream.
The he puts up preemptive straw man arguments against doomers. His blog has become a joke.
People might feel differently about AI if they were a part of the changes rather than being a helpless spectator.
That would have given grad students who've been grinding towards a PhD for years a fighting chance to see if they could leverage the model to push their work forward, rather than watching years of work potentially turn to dust via a tool they don't even have access to.
It wouldn't delay the progress of mathematics by any meaningful amount in the long run (an X month delay is nothing) for OpenAI to take this approach, and would help somewhat to preserve the health of mathematics as a field. Without it, the motivation for any young mathematician to devote years to a new problem must be sapped knowing there's an uneven playing field... an OpenAI team with access to colossal tools months before they'll ever be able to get access, willing to scoop anyone as soon as they can, perhaps without even taking the time to completely understand the proof.
I don't see any long-term benefit to OpenAI with their current strategy. This is an internal model; it's not available for sale at the moment. They've said they're not even going to bother claiming the Millenium Prize money for Navier-Stokes. It feels like kicking over hundreds of other people's chessboards just because they can.
I bet sharing unreleased models with outsiders safely ain't that easy.
Has anyone verified any of the proofs produced by OpenAI or is everyone just assuming that it just be true because the Lean code checks out? Couldn’t the Lean code just be formulated incorrectly?
> In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully.
For example, the statement of e.g. Fermat's last theorem in Lean should be understandable to anyone who played The Natural Number Game [0] and knows a bit of mathematics and programming. For the proof, you trust the compiler.
The statement of other theorems can be much more delicate, and the Lean formalization may require an extensive introductory section which will need to be carefully checked.
Then there are the cases where no Lean formalization is currently available, and all we have right now is an often impenetrable pdf in the OpenAI repo. I would not at all be surprised if some of those contained logical gaps.
Time will surely tell, but there are certainly doubts and lots people are very busy checking these results.
I believe so, even with all the usual safeguards properly in place: https://news.ycombinator.com/item?id=49672339
> is everyone just assuming that it just be true because the Lean code checks out?
Kinda? It's only been 24 hours since they dumped 722 manuscripts on the world, most of which are apparently basically unreadable, and only some of which come with a Lean proof, which in itself is not a joy to read afaik.
In other words, they save effort by wasting the effort of others.
I'm however puzzled by the number of proof claims without lean proofs. How does OpenAI have confidence in those, especially if, as noted in the blog posts, the papers are very hard to read?
The same problem almost every person on earth is going to have to reorient to in the next decade, which is: how do we eat and stay housed when we have no real economic value?
It's just that we lost one of the important ways to demonstrate understanding.
(To my credit, I have warned them since more than a year ago that we will reach this point.)
But are any of them going to be able to do any significant amount of studying maths without being homeless?
Why not? Maths studying correlates with high intellect, high intellect correlates with competitiveness in the job market. So their chances of finding a job that would allow to have much of free time (for doing math) without being homeless are pretty good.
Math isn't about collecting random theorems, progress in math is about gaining understanding of new systems, and the theorems are guideposts to aid in that understanding.
You can prove 1000 theorems and not really increase any understanding about a subject, but gain knowledge of 1000 random facts. For example, I can write down some complicated equation and ask you "does this have a solution in the integers"? And if you do a maze of very complex and tedious algebra to show that there is a solution, you would have proved a theorem, but you would not have done much to move math forward at all.
On the other hand, if you introduce some completely new technique, say you take my equation and turn that into an algebraic surface, and then you count some special curves that live on this surface using geometric ideas, and then you show that if the number of such curves is odd, there must be a solution in the integers, and in this specific case, it is odd, so there is a solution -- well, then you have really pushed math forward and people will celebrate your proof, even though no one really cares if the equation I wrote down has a solution in the integers.
For example, there is a long history of failed attempts to prove Fermat's last theorem driving algebra and number theory forward by introducing the concept of ideals, for example, and this concept ended up much more important than whether Fermat's theorem is true or false, which is not too much more than a piece of trivia.
Or for example, the recent proof of the Poincare conjecture relies on the machinery of the Ricci flow introduced by Richard Hamilton, who then applied it to solve a number of open problems, but Perelman was able to take it even more forward to solve Poincare. So Ricci flow was massively important machinery.
For this reason, we celebrate people like Gromov, who didn't really prove that many theorems but introduced amazing machinery -- for example, the h-principle, or Gromov Compactness -- these were ideas and math is about the ideas. The ideas are then applied, using laws of logic, to form theorems.
So mathematicians will need to mine these proofs to see if there are any new techniques - new machinery - being introduced, or if the AI just used the existing machinery more efficiently. Here too, we are just looking at AI as a form of search, which it is really good at, since there are so many thousands of papers and so many ideas, that there might be a connection between two areas that lead to a solution and the human mathematician, not knowing all known results, can't make that connection. In the future, we may wonder how anyone did math without AI, much like we would wonder how anyone can be a writer without access to a dictionary or reference work. Is the AI just searching through a catalogue of known ideas and connecting them or is the AI coming up with genuinely new stuff like Ricci flow or the h-principle?
What is interesting is seeing whether we can get AI to actually discover new machinery for us. That would be huge.
And then we need to find efficient ways to detect these ideas and describe them.
Really this is very exciting and opens up whole new workstreams for mathematicians.
Emotions can be funny.
1. Help understand, check, and explain the results.
2. Write new works explaining or refuting the new approaches and results in more lucid language.
3. Advance the field further.
I don't know why this is not obvious. Each step is intended to support human understanding, not to replace it. Any mathematicians who don't do these will be left behind, and if none do it, the field of human mathematics itself will become obsolete. All I am hearing so far is excuses.
#3 at this point might need more human intuition; but that might be a 2026 problem.
The "summarization" process can be multi-step. Initial steps develop new frameworks and prerequisite concepts. Later steps build upon these frameworks and concepts. Some can be simplified too if this can be done without loss of fidelity. The summarization comes last, and is meant to preserve sufficient context.
> I don't think anything about how to react to a paper being released in this state is obvious.
If it helps, exaggerate the condition a bit by imagining coming across a repository of alien knowledge.
If you don't understand why the paper is in its current form, how do you know what changes would damage the fidelity for the AI? I think your example of alien knowledge is misleading because it suggests the aliens don't need to understand the info - but in this case we intend to keep using AIs to expand this work. So if you transform it into something that humans think makes more sense without understanding why it's in a form that we hate now - you risk losing something in the form that you don't understand.
And if humans decide to give up on mathematics because the process of using the machine is so tedious, then there will be no value in automated theorem proving.
> resistance will fail utterly as long as AI itself can understand prior AI works
That's not the issue in the long term. Look further: where did these questions come from? They came from humans who had been exploring mathematics and noticed interesting patterns that cropped up again and again.
The people who posed these questions had a profound understanding of the background maths that led to them. Me - I don't. Take one of the results he quotes: "Positive solution to the Unitary Synthesis Problem". He may as well be one of your alien races yabbering to me in an alien language. The question is meaningless to me. The answer the AI came up with, correct though it may be, is equally meaningless. You may as well have told my dog how to find eigenvalues.
If we let AIs do all the understanding so we know nothing about math, there will be no human understanding that lets us find new questions, or understand the value of any questions AI might pose and find the answers to.
^^ half of the comments on this thread
It's looking to me like it's more of a slop PR problem than it is that these things are genius at math and will displace mathematicians. I am happy to be wrong but I strongly suspect the next few weeks to months will result in more and more of this work being exposed as slop.
These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
> These things are ok-ish to halfway decent at coding tasks with a ton of babysitting and still make tons of extremely simple errors almost constantly, why should math be any different?
1) AI is winning programming competitions, 2025 was probably the last year we've had human participant winning* 2) Math is different because there's formal verification.
* Of course competitive programming is different than enterprise programming, but competitive programming is closer to math.
My dislike lies in how it devalues skill + knowledge acquisition, disrupts art/software communities, and that it's now bloody impossible to tell the difference between someone who actually knows something from someone who merely regurgitates shit from their chatbot.
Usually 9 year olds imitate adults when they regurgitate such words in these circumstances. What a sad state of affairs.
The checks notes meta was retired ages ago btw.
And going from zero proofs to one proof (even a sloppy one) is a big deal regardless of whether it was written by AI or a human.
You might as well say "the obsession with sex has always held humanity back". Maybe, but it's complicated...
Many math problems are practically useless if you only care about the answer, the millennium prize about the Navier-Stokes equation is such a problem. The solution makes no physical sense, real life fluids don't follow the Navier-Stokes equations in such extreme conditions. But in the process of finding the solution, we may get insight into what will end up being really useful. The big mess that OpenAI produced is the solution no one really cared about, but it didn't deliver much of what people actually wanted.
One reason it is sometimes seen negatively despite being at least something is that it broke the incentive. Without the million dollar prize and with only the privilege of being second, people are much less likely to go for the insightful solution.
It would be fantastic if university and science was like "here is 100M$, play around and develop some 'understanding'". However the reality is that human societies are hierarchical and currently capitalistic which implies value creation and status building.
1) the funding bodies/agencies need proof of value that you're using the resources meaningfully to be able to assign resources
2) Humans are status seeking, power seeking, resource seeking and sexual reproduction seeking. If you hold a lot of power and make decisions, you have more of all of the above.
Highly recommend reading it. Very prescient for something written 26 years ago.
https://gwern.net/doc/fiction/science-fiction/2000-chiang.pd...
I also recommend "Exhalation", though that has nothing to do with AI.
What was it about the other problems that made them unsolvable? Was it just a time constraint, or are they just harder problems?
Probably some combination of: some of the 372 problems were easier than the rest; the AI got lucky on these 372; there were existing papers out there in the literature which proved especially helpful for these 372; and other similar factors.
In the past 2 years the AI's started solving math problems in roughly the order of "hardness" as ranked by humans.
https://cepr.net/publications/ai-didnt-steal-the-mathematici...
Frontier theft is just faster.
Also lean proofs are notoriously tedious and slow to write so this level of output is very likely to be from LLMs. The number of people able to understand this level of math and prove it using Lean is a few hundred at most.
We'd have to accept that there are hundreds of such mathematicians willing to take the OpenAI stock options for that conspiracy to be true, and that no one leaked being offered that at all, so personally I don't think that's possible. Which leads me to conclude that OpenAI's AI did indeed create the proofs, with some level of help from humans to guide it.
We all know it’s only a matter of time before lawsuits in this area become widespread particularly in lawyer happy America.
There are suddenly many new solutions to problems that have resisted sustained attacks (e.g. the Uniform Games Conjecture as detailed in TFA at some length). Where do you think they are coming from? Why is there suddenly a bunch of results to be stolen?
Gödel's incompleteness also constrains human and AI mathematicians. Both just strive to prove whatever can be proven in the system they are working in.
(That said, this is not fun, and I sympathise! I'm still a student and would like to avoid finance if at all possible. Just suggesting not to drop everything if you feel like you're learning in the process.)
(I'm now personally in the position of having to choose a PhD project, and this rapid change is interacting with making long-term plans really badly. Guidance welcome!)
Almost every single person who earns money for labor is, or will very soon be, in exactly the same position as you, many are just either unaware of how fast the change is coming or are deep in denial about it.
None of us know what to do about it other than hope that we find a peaceful political solution prior to the economic collapse.
Unfortunately, the current world is not set up for that, and until we get there, yeah things are gonna get pretty hairy. But if and when we get to the other side, I think the future will be glorious.
Pretty funny way of putting it. Presumably model X+2 will be able to explain these in elegant, human legible ways, though (as well as solve the remaining 95%)
https://arxiv.org/abs/2610.08144
> Autoformalisation is increasingly used to verify mathematical texts, including those generated by AI, as in OpenAI's announced proof of blow-up of solutions to the Navier-Stokes equations. In this process, an AI system translates the text from a natural language (NL) into a formal language such as Lean. Once this translation is done, the argument expressed in the formal language can easily be mechanically verified. The purpose of this article is to demonstrate why this process may offer no confidence in the original NL argument, owing to the various difficulties in performing the translation semantically faithfully. In particular, we highlight that the problem of resolving ambiguities in mathematical NL text, which is necessary in order to provide semantically faithful translation, is arbitrarily high up in the Solvability Complexity Index (SCI) hierarchy/arithmetical hierarchy (the SCI =∞). Hence, informally, providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem (which has SCI =1). To demonstrate the effect of this result we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof. In particular, we show that the formalised Lean proof does not correspond to the NL proof of blow-up of solutions to the Navier-Stokes equations.
And I don't think that paper addresses it, but if the LLM can find a bug in Lean and exploit it to prove something, there's a good chance it will find it and not report it. So if you've got some million-line proof in Lean, spit out by an LLM, you still can't quite trust it, even after validating the problem transcription.
(This is the same category of problem as the huggingface hacking incident, where the LLM finds and exploits an unintended cheaty loophole)
Why would it know it found a bug?
Similarly if the formal Lean problem statements (human generated) are correct translations into Lean (and the original NL statements are sound, which one would hope after decades), and no Lean bugs are abused by the proof (as defined above), then the proof is valid.
The NL/Lean discrepancies are super annoying and will make human analysis hard and fraught, but as many posters have found out the models themselves will gladly pick apart the NL-Lean translation for errors, and so my guess is that finding the discrepancies will not take too long. OpenAI really should have done a dynamic workflow over every lemma and step to ensure pointwise accuracy in the translation.
If you have any great skill at the upper level, you better work off-line/local private and not share, but for many the temptation will be too great.
And what the hell will the frontier labs have by then?
Maybe I'm overreacting, I'll have to screw my head back on before I can process this.
>If you’re still a proponent of that doomed worldview, still aboard the sinking ship, I encourage you in the strongest possible terms to read yesterday’s other great contribution to AI discourse, besides the OpenAI Mathocalypse dump: namely, Scott Alexander’s open letter to Steven Pinker. I feel some responsibility for this, as the person who first introduced Steven Pinker to the existence of the rationalist community, and who also first introduced Steven Pinker and Scott Alexander to one another (they had both been fans of each other’s writing).
Whole lotta yikes
Previously, due to the cognitive limitations of our brains, it was difficult to tell whether resolving a conjecture genuinely required years of dedicated effort, or if the answer was already hidden within existing human knowledge. Or whether a problem is just waiting to be searched out through ingenious or even brute-force means. With AI, we can now offload this search, especially across different areas of math, allowing us humans to focus our minds on discovering new mathematical structures and techniques.
Indeed, if one believes what Hilbert believed: we must know and we shall know, he should feel happy, as the belief has never been about who solves a problem, but about whether can advance our understanding of the universe. With the help of AI, we have a lot more possibilities.
Sadly, this is also what day-to-day work looks like for a lot of software engineers in industry right now. I spend my time reading and verifying thousands of lines of messy AI-written code. Compared to actually producing something, it's thankless, barely-credited, and unfun work.
To be clear, SotA models/harnesses are really good at making things that work one-shot, and their code golf and debugging game is insane.
When I go to implement, even with a good spec, I end up with code that's 80% of the way there in a fractal manner. The modular decomposition is 80% of the way to good code. The function decomposition is 80% of the way there. The computation structure and variable naming within functions is 80% of the way there. I can walk the AI through it and address each level of issues, but it's tedious as hell, and not clearly faster than doing it myself in some cases.
This is with omp/claude/codex, with Fable 5.1/Opus 5.5/6 Astra. The Chinese models do better at staying coherent, but they're a little less smart IME. I've tried many permutations of "fable planning, opus sub-agent implementation"; "pass a branch back and forth between opus and astra, as reviewers and feedback implementers". I haven't tried the "software factory with an architect and 2 juniors" thing. I haven't tried AGENTS.md beyond the /i-have-adhd, iso-24495 (lol at the Opus 5 induced PTSD), "No cleft constructions or Latinate absolutes" and general project orientation. Model character seems to change too fast to make agents worth it, and I see stuff about skills and excessive AGENTS reducing model capabilities.
Am I holding it wrong?
Matt Pocock says that he knows how, but it comes with a lot of context fiddling.
https://www.youtube.com/watch?v=v4F1gFy-hqg
"Software Fundamentals Matter More Than Ever"
This lecture is very impressive. However I haven't managed to reach AI enlightenment, yet.
I am honestly impressed by how many arguably highly intelligent humans are happy to throw out millennia of scientific methodology just because the world-burning probabilistic machine produced an output that tickles their dopamine brain.
Also glancing at the repo. Such an advanced machine, yet failing to output results with a common template. Every pdf is entirely different from the one before, making the work of the poor humans trying to find a sense in it even more difficult.
I don’t think it helps your argument if this is half of your post. The fact that the pdfs don’t follow a common template? That’s the complaint? Human papers also don’t follow a template.
If the only way to power a "technology" would be to kill an elephant every time you use it, then it's a flawed technology even if it solved fundamental problems of humanity. This is currently how GenAI sprawling data centers operate AND without solving any fundamental problem of humanity.
Also I don't care about A specific template, I just find it ridiculous that some trillion dollar company cannot put a line like "please keep papers' structure consistent" in whatever markdown file controls these things. It's sloppy, like everything OpenAI does.
There was a bunch public literature on these topics. They were major unsolved problems. Lots of humans would want to have solved them and tried. But here’s it an llm that has solved them.
It's already falling apart.
> apparently they tried the model on about 8,000 problems. So, right now it “merely” solves ~5% of the longstanding open mathematical problems that it’s asked about, the problems that whole communities have spent years on, after a single 3-hour attempt on them.
Back-of-the-envelope cost calculation:
3 hours of GPT-Pro compute per attempt / 5% success rate → 60 hours of compute per solved open math problem
At GPT-6 Astra API long-context pricing ($75/M output tokens), assuming 50 reasoning tokens/s, that's roughly $810 per solution:
60 h × 3,600 s × 50 tokens/s × $75 / 1,000,000 = $810
Of course, actual token usage is unknown. But even if off by 10x, $8100 is cheap for solving a longstanding mathematical problem.
“On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems.”
That leaves a lot of ambiguity. Is “result” exclusively “positive result shown here” or “all results”?
Since proving infeasibility is a valuable result, the unpublished problems must not have reached a valuable end state. Thus it’s an operator decision of when to turn the machine off and try another problem. I’d thought they had a three hour time box on this, but apparently not.
Sure yeah, today it's only a 5% solution (to the hardest problems), and we know it will top out somewhere.
But the cost reduction is just too good to ignore.
Terry Tao was saying that we need more mathematicians [0]. My paraphrase of his article is that we will need them to check the AIs, essentially (I missed things didn't I?).
But with a near 10,000X reduction in cost [1] today, it's really hard to argue for that when budgets are continuously tight.
[0] https://terrytao.wordpress.com/2026/09/24/were-gonna-need-a-...
[1] assume ~$150k/year for a mathematician of the caliber that can do this work. 2x that for overhead like healthcare, 401k, etc. Do this for 35 years of work = $10,500,000. Assume about 1 result of this caliber per mathematician. So ~$10M per result. OpenAi is saying they can do it for ~$1k, a 10^4 X decrease.
Isn't that more likely the easiest problems in the set? I suppose the hardest problems compared to those that have been solved.
Otherwise your points stand.
Sure, but we are approaching a time at which we will have to answer a "why" here. I.e. anything beyond checking and validation becomes of mostly artistic value at some point. And there is nothing wrong with that.
An early (easy to understand) proof that fits the bill for me is proof that the sqrt(2) is irrational: a very natural thing to think about —- for a unit square, what is the measure of the diagonal? And people were killed in ancient Greece over the proof, but it’s comprehensible with basic algebra.
And then you're off an order of magnitude of the efficacy of OpenAI. We don't know how much work of mathematicians is behind all these results, what's the real cost (which includes those problems they couldn't solve), etc.
https://www.merriam-webster.com/wordplay/bury-the-lede-versu...
In actuality both "lede" and "lead" are correct. "Lede" is mostly an American and recent respelling of the word "lead" and "lead" still remains the predominantly used term (while "lede" is predominantly an American journalistic spelling).
I heard they used lede to distinguish from (hot) lead. But since linotype has been dead for 50 years this is probably moot now.
Reading the account from Scott's wife, the proofs seem to resemble more a Zelda speedrun with a bunch of strangely exploited bugs than an actual thing.
The most important takeaway from this is that mathematicians finally consider us CS grunts their equals.
"It feels like something written by someone who’s on psychedelics. So much unclear and doesn’t make sense. Lots of name dropping of previous work without discussing why it can be used despite impossibility results. Basically the paper is so horribly written that it’s impossible to read it without AI help."
This makes one wonder how many of the proofs are actually valid, and how many are just impenetrable hallucinations.
I find it interesting that you state this as obvious. When ChatGPT was introduced four years ago, the argument was that language was fuzzy and probabilistic and that LLMs would thereby never work in math.
I think someone asked here a bit ago whether other search strategies using OpenAI levels of compute have been tried and think the answer is no (even 3 hours of present GPU computer I think is vast compared to anything available 10 years ago, say).
Comments are loaded live from Hacker News and are not stored by Mid or Real.
matt3210 24h ago on HN
dekhn 24h ago on HN
woah 23h ago on HN