MIDREAL

Sharing AI progress in mathematics

Comments

senderista 46h ago on HN
Good to see they're engaging with the mathematical community, even if they had to be publicly shamed into doing so.
sebzim4500 44h ago on HN
This is the opposite of what the mathematical community were asking for. I say this as someone who strongly approves of this approach.
senderista 44h ago on HN
Yeah I'm not sure they met them even halfway.
tim333 27h ago on HN
How so? I saw at least on mathematician say publish what you've got. What were they supposed to do differently?
karahime 46h ago on HN
Extremely unfortunate that gate keeping got to the point where they felt the need to ask for permission to share math.
bravoetch 45h ago on HN
In previous math sharing there was speculation about stealing human researcher's results or progress, via prompt inputs from those researchers, and sharing that as their own result. They're adjusting their process, and it seems ok.
whimsicalism 44h ago on HN
Let’s be very clear, the alleged “stolen results” were largely the product of another LLM, not de novo human work. Also, it was false - they did not steal the results.

No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.

Ancapistani 40h ago on HN
Can you show where it was discounted? Last I heard OpenAI was declining to deny it, presumably while they thoroughly confirmed.
whimsicalism 40h ago on HN
> “Following an investigation, we have confirmed that Buckmaster’s Codex prompts over the two months preceding this announcement and paper on September 8, 2026, could not have influenced the system in any way, including through training.”

https://openai.com/index/navier-stokes-solution/

youoy 36h ago on HN
Come on, they were working on this for more than 2 months. Dont fall for the corporate half truths.
whimsicalism 32h ago on HN
Buckmaster himself said essentially all the progress they made was from July 15th onwards (with a new model on the problem) with no real progress prior to then.

I think it’s really cope to claim this was a human result being stolen.

margorczynski 34h ago on HN
I really don't get how people are still continuing with this "stolen results" narrative after today. Like NS was kinda insignificant compared to treasure trove they released now, thinking that the LLM needs to "steal" from some human is simply coping.
robotpepi 33h ago on HN
After reading your comment one could even think that LLM's invented math from the ground up.
Leynos 32h ago on HN
HackerNews does love a good conspiracy theory.
robotpepi 33h ago on HN
> HN struggles with truth-seeking on these topics.

The facts are much more nuanced than how you're presenting them here.

binlog 33h ago on HN
What’s the nuance? Two researchers alleged theft. OpenAI investigated and confirmed that their research was not in the model’s training data. The world decided to take the first part as objective truth and ignored the second.
robotpepi 29h ago on HN
> OpenAI investigated and confirmed that their research was not in the model’s training data.

You mean when they say it was impossible to confirm anything but one day after it was 100% confirmed that there was no theft? And here I'm not even talking about all the ethical problems related to trying to scoop another group when you hear they're close to success, or how current solutions follow extremely closely human-generated ideas, or about the lack of relevant citations in OpenAI's paper.

Believing that OpenAI's claims have any substance cannot be explained by naivety alone.

keeda 27h ago on HN
They didn't try to scoop another group; they thought the other group had already solved it, so maybe their latest model could take a shot too... and the model solved the full problem when the other group had not! They found out after the fact that the other group had only solved an important sub-problem.

This was a low-key hilarious replay of George Dantzig and his homework problems: https://en.wikipedia.org/wiki/George_Dantzig

The controversy was whether they had plagiarised that other work on the sub-problem, which they categorically denied after an investigation. And yes, a few days to investigate something like this is reasonable for a company as big as OpenAI. Having seen how data infra is set up when petabytes of data are flowing about, there are thousands of entwined data pipelines to figure out. Not quite as easy as running a query on a sqlite DB!

> Believing that OpenAI's claims have any substance cannot be explained by naivety alone.

Yes, they could be explained by a GitHub repo full of proofs :-)

Or are you suggesting there were hundreds of researchers who just happened to be close to solving hundreds of these long standing open problems using Codex, and OpenAI swooped in plagiarized them all? ;-)

whimsicalism 32h ago on HN
glad you could present those facts in your comment, i know the HN character limit can make it hard
xpct 45h ago on HN
Just to be very clear: they aren't asking for permission, they are framing it that way because of the bad press.

There's no gatekeeping here!

lynndotpy 32h ago on HN
Right? Math is the most open of our academic knowledge institutions, by virtue of what it is. It's easy to get any math publication, and I am not aware of any other fields where an anonymous person can publish their work informally in an anime discussion and enter the annals of math knowledge.
potsandpans 21h ago on HN
hgoel 45h ago on HN
After how poorly OpenAI and Anthropic handled the previous cases, I approve of the more measured and cautious approach this time.

We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).

If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.

make3 43h ago on HN
It was OpenAI that fucked up the Navier Stokes explosion situation, not Anthropic
hgoel 42h ago on HN
I'm not referring to just that. For example, Anthropic published a half done report about some biology research that turned out to have already been discovered and patented, and IIRC both companies are guilty of claiming results without doing the basic diligence of citing the material their work builds on, effectively passing it off as entirely done by their AIs.
kzrdude 40h ago on HN
They didn't even follow the recommendations of the reference group. A few of them maybe, but this is still a dump of llm-written papers.
yesbutnotreally 32h ago on HN
The gate keeping, I'm afraid, will be now in the hands of various bubecks, responding directly to even more disgusting people.

With all the hierarchy present in mathematics, I would prefer it by far.

This thing named inappropriately "OpenAI" goal is just grabbing and monopolizing. Capitalists before could not really touch the deep of the human spirit with their filth, now they can.

ruffrey 28h ago on HN
This is a tricky one, but I do sort of agree with you. However OpenAI doesn’t give most researchers access to the models which produced the work. So the gatekeeping goes both ways, I think.
ravenical 46h ago on HN
https://github.com/openai/math
binlog 46h ago on HN
So happy this is shared on GitHub rather than some gatekeeping paid journal. Truly a new age for science.
fph 45h ago on HN
Most mathematical results are shared on Arxiv. Journals add peer review.
traes 45h ago on HN
GitHub is a significantly worse place to store important results than Arxiv. Of course, slop does not belong on Arxiv, buy slop should also not get published.
adverbly 45h ago on HN
End of an age for journals?
a57721 31h ago on HN
In a sense, because now journals will be flooded by LLM slop.
rafterydj 46h ago on HN
I don't know, this does not feel like the message hit OpenAI where it needed to hit, if this is their primary response.
osiris970 45h ago on HN
You want them to stop doing math research?
k2xl 46h ago on HN
Can someone knowledgeable about the subject outline the most significant portions of the results?
sebmellen 46h ago on HN
It’s fascinating to read through the reasoning traces: https://github.com/openai/math/tree/main/reasoning_traces

Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...

ndriscoll 45h ago on HN
> Thus at most one informative i. So cheater chooses arbitrary g_{v_i}, on exact duplicated input matches and passes, independent of actual satisfiability!

No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.

adverbly 45h ago on HN
> Look at one of their examples of an initial prompt

Interesting that its only an excerpt. I wonder what else they include but didn't share.

philipwhiuk 43h ago on HN
Attempts to edit the problem description on Wikipedia ;)

https://wikimediafoundation.org/news/2026/10/05/openai-rogue...

cubefox 38h ago on HN
These are not reasoning traces, these are summaries of excerpts of reasoning traces.
ed 46h ago on HN
Actual results: https://github.com/openai/math/blob/main/overview.pdf
gizmodo59 46h ago on HN
This is significant progress and released without all the drama. Some very important progress in Reinmann, Hodge and unique games theorem. Point the repo to your agent and ask for the significance! In a way this is probably 50-100 years of math progress by humans
fspeech 45h ago on HN
Math is the tool humans use to compress knowledge. So until we can comprehend it there really isn't much progress. Math theorems are tautologies, the truth of which are not dependent on proofs and proofs are erasable, at least classically. But the AI progress is exciting and AI proofs are a gold mine for humans (at least non domain experts) to explore.
gizmodo59 45h ago on HN
>So until we can comprehend it there really isn't much progress.

Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as

fspeech 45h ago on HN
If it changes how we think then yes it has an effect.
le-mark 44h ago on HN
But who will ask the questions or direct further research when humans no longer understand the state of mathematics? Llms lack the drive for homeostasis combined with the evolutionary drive for survival and reproduction thus to direct themselves. They could very easily spend an eternity down a rabbit hole when the warp drive equation was fairly close on another branch.
fspeech 45h ago on HN
Another way to state this: math theorems are like programs without side effects; it is immaterial whether a program without side effects is ever run. We study math for the side effects: it changes how we organize our thoughts.
warkdarrior 45h ago on HN
> Math theorems are tautologies

Proven math theorems are tautologies.

fspeech 45h ago on HN
True.
fspeech 45h ago on HN
FLT was no less a tautology before it was proved. We just weren't sure about it. Proofs only change us, not math.
binlog 45h ago on HN
What makes you think no one can comprehend this? It has been less than an hour since it dropped and there is already a ton of online chatter from people explaining the results, pointing out their favorites and more. Some of it is happening on this very thread.
fspeech 45h ago on HN
I didn't say that. I am responding to "In a way this is probably 50-100 years of math progress by humans." I am actually very excited about AI proof and I am working overtime in my own way to try to comprehend as much as I can.
gpt5 45h ago on HN
Math is far more than that. If you can solve prime factorization for example, suddenly you can listen and interfere with almost every private conversation on the internet.

We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.

fspeech 45h ago on HN
This doesn't contradict what I said. But I do appreciate the fact AI can produce side effects not just humans. I made it sound like only human knowledges matter. That's too narrow.
caaqil 45h ago on HN
> until we can comprehend it there really isn't much progress

Who is "we" here exactly?

fspeech 45h ago on HN
Whoever wants to study the result.
caaqil 45h ago on HN
> Whoever wants to study the result.

Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.

Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.

fspeech 44h ago on HN
AI is very helpful with understanding AI proofs. Agent swarms produce messy proofs overall but locally they are excellent and can teach anyone who wants to study them. No one controls math (in a material way funders do control an aspect of practicing math). Still, theorems are already true before we prove them. The difference a proof makes is whether it convinces the reader.
yieldcrv 45h ago on HN
Academics have been treating it that way because they had no other choice, and its been a waste of everyone’s time and often times taxpayer resources

Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades

Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants

fspeech 44h ago on HN
If you actually looked into how agents proved FLT you would be even more amazed by the fellow human beings who are able to keep all this in their heads! I for one can only begin to grasp the scope with AI and scripts.
fspeech 45h ago on HN
I think it would be helpful to people who want to understand what a formalized proof is to read Thomas Hales on this: https://www.math.stonybrook.edu/~bishop/classes/math536.S24/...

He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.

fspeech 42h ago on HN
Fun challenge: find the formal definition of simple_closed_curve in the essay and tell me if you believe you learned anything about a planar curve.

BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.

AIblemblio 35h ago on HN
It is progress on another / the next evolutionary later: A AI/AGI/ASI system.

Which either replaces us in the long term, augments us or makes us better (gentherapy).

traes 45h ago on HN
Not to pick on you specifically, but as someone who spends a lot of time unproductively reading AI math discourse it's truly shocking how incapable all the supposed math enthusiasts are of spelling Riemann.
xpct 45h ago on HN
I just did a quick search on this and apparently the misspellings are German surnames as well:

https://en.wikipedia.org/wiki/Reimann

https://en.wikipedia.org/wiki/Reinmann

conformist 45h ago on HN
Yes sure but they are different surnames and pronounced differently.
xpct 45h ago on HN
I didn't mean to oppose OP's point, I just found it interesting as a non-German speaker!
traes 45h ago on HN
I just don't understand how it happens. If they had ever taken an intro to real analysis class they would learn to spell his name. If they were just parroting what an AI told them... shouldn't they still just say his name? An individual could just be dyslexic or mistaken but it seems to be a substantial volume. I guess they just don't care enough about it to commit the correct name to memory, only remembering the "pattern" of the name and filling in the spelling via guesswork?
NewsaHackO 45h ago on HN
People just don’t spell that seriously buddy, especially when it is so immaterial to the point.
traes 45h ago on HN
My point is it is crazy to make public claims about how important or not important a mathematical result is when you can't spell Riemann. Yes, it technically doesn't matter, but it betrays a damning lack of familiarity with introductory mathematics.
xanderlewis 44h ago on HN
You’re (as Claude would say) absolutely right, and I suspect the original commenter has no idea what they’re talking about.
justanotherjoe 25h ago on HN
Then ignore them? I dont get these types of mysteries. It's easy enough to find math majors these days to ask them their opinions on things. There're so many of them. You most likely know some from your highschool. They'd probably say the same things though, or even freak out harder.

While it takes a good math person to make breakthroughs it's much easier to find someone who has a feel of whats important/hard and not. Even a mediocre math major/master is far more authoritative than an expert at adjacent fields (CS,physics). Or to listen to webdevs 'ai skeptics' or whatever on the internet.

ndriscoll 45h ago on HN
Maybe they skipped straight to Lebeg integrals.
raegis 38h ago on HN
Thanks for the laugh!
thunspa 35h ago on HN
very good lol
jryb 44h ago on HN
Autocorrect might be doing it
paulhebert 43h ago on HN
My last name is Hebert.

There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.

Even in situations where they just read it or I just said it.

I’ve had Herbert soccer trophies, health insurance cards, etc.

The mind fills in a lot of blanks and doesnt always get them right.

jbaber 43h ago on HN
I sympathize. -- Not Barber
Agentlien 36h ago on HN
My last name is Kvick - the Swedish word for quick. I live in Sweden. I get a lot of people thinking my name is spelled Kvik, Quick, Kwick, ... My father once got a mail addressed to Mr. Kvack (Swedish for quack, like a duck).
ShadowOfThePit 6h ago on HN
Fuck, I read Herbert at first. Fast reading is just guessing half the words!
neutronicus 42h ago on HN
iPhone would be my guess
lanyard-textile 45h ago on HN
They're mathematicians, not linguists :)
traes 45h ago on HN
The mathematicians don't make this mistake, I assure you. I doubt there is a math professor on planet earth who would spell Riemann as Reinmann. In fact, I imagine no one who has ever heard the name pronounced would do so.
pixl97 44h ago on HN
Uh oh, no true scottsman....
vector_spaces 44h ago on HN
It's just a name you write so many times as a math undergraduate or first year graduate student due to the number of load bearing results and objects named for him. The name even appears in lower division coursework

Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.

quacktopia 40h ago on HN
I regularly heard lots of names during my 3 year math undergrad degree. I couldn't remember how to spell many of them then and can spell fewer now. I imagine most people on the course were similar, and me nor no one I knew were dyslexic.

Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something.

Google existed then and now and we could look them up if needed.

gpm 43h ago on HN
One of the best mathematicians I've had the pleasure of learning under added 6 to 7 and got 15 during a lecture... I assure you they're capable of misspelling last names too.

Mathematicians aren't exactly known for being well rounded.

jwilber 43h ago on HN
Pure mathematicians getting arithmetic wrong is a bit of a meme it’s so common. I don’t think it’s the same as a misspelling of a popular mathematician, not that either are indicative of much tbh.
senderista 39h ago on HN
google "Grothendieck prime"
scrame 42h ago on HN
Oh god, that reminded me of my linear algebra teacher in college. Doing matrix multiplication by hand and ending up with 9x6 = 45, and then having to direct him to the cell with the wrong number. Loved the topic, hated the class.
bananaflag 39h ago on HN
Still, enough misspell Lebesgue as Lebesque.
bootsmann 37h ago on HN
If they’re mathematicians they have written this name down about a 100 different times throughout a standard Real Analysis course. Riemann was foundational in that field.
jere 44h ago on HN
“How many Ns in Riemann?”
sdenton4 43h ago on HN
Let's figure it out! First, draw little boxes over all the letters. Then add up their areas. Then make the boxes smaller and repeat, until you have forgotten how to count.
tim333 31h ago on HN
Human brains seem to have somewhat similar failure modes to LLMs and how many 'r's in strawberry.
zone411 45h ago on HN
There was A LOT of drama about this release.
cyclopeanutopia 37h ago on HN
> Point the repo to your agent and ask for the significance!

Wow, this comment really shows how low this community fell.

robotpepi 33h ago on HN
> This is significant progress and released without all the drama.

I feel gaslighted.

enoether 46h ago on HN
Unique Games Conjecture [0] is a seminal conjecture in Complexity Theory, and is an underlying assumption for many, many inapproximability results. A valid proof is a big deal!

[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...

impossiblefork 45h ago on HN
Yeah, that's one of the big things of TCS. I think I see that as bigger than that Millenium Prize problem.
davemp 43h ago on HN
TCS being theoretical computer science? I have not seen that acronym before.
jhanschoo 42h ago on HN
Yes, TCS is theoretical computer science, I commonly use that acronym too.
gregdeon 45h ago on HN
This was the biggest highlight for me as well. Astounding...
inkysigma 45h ago on HN
I also don't think there was general consensus on which way this would resolve prior to this (or is that a little out dated?) unlike some of the other major problem resolutions. I heard rumors that there would be a big result in TCS and speculation it would be UGC that or P neq PSPACE but I'm still a bit shocked.
amluto 44h ago on HN
I'm really glad that OpenAI is formalizing these things, because I'm not convinced that their current internal frontier models are particularly good at writing down their thoughts in English. From the (probably awesome) Unique Games Conjecture/Theorem paper, the first two sentences of section 1.1 start to define the problem:

> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).

I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.

1. e is maybe a name of a list.

2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.

3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.

4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.

So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.

Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.

If this were my paper, or if I were trying to train a model to write math, I'd want something like:

A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.

A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).

danbruc 35h ago on HN
[…] a finite edge set E = (V × V) […]

E ⊆ V × V

amluto 32h ago on HN
Oops, that’s what I meant.

In this particular case, though, I think my typoed version may be equivalent. An edge with no constraints has the same effect as no edge at all.

I definitely messed up the constraint definition, though: u_e and v_e refer to vertices, not edges. That’s what I get for writing it with minimal proofreading.

Catloafdev 45h ago on HN
This is a pretty hilarious thing to read juxtaposed with AGMAI's requests.

Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"

open592 45h ago on HN
Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?

goalieca 45h ago on HN
Don’t paste your research into these AI because they will train on it and then scoop you.
esafak 45h ago on HN
I think that happened after word of the project reached OpenAI and they allocated resources to it.
binlog 45h ago on HN
Use whatever is published as the new base for your research. Use AI tools to help you going forward.
xpct 45h ago on HN
In other words, you've already taken a gambit with the first half of your PhD, now take a second gambit, praying that you have something to publish by the end of your PhD.

It has to feel awful to be in this position.

torben-friis 45h ago on HN
Could be worse, imagine having years of experience in a profession these things can now handle by themselves.

:)

jltsiren 44h ago on HN
It's worse for those who will face the job market (or maybe tenure review) in the next 2–3 years. If you have more time before you have to justify your continued employment, you can pivot to something else. Theorems and proofs may be cheap now, but there is a lot of value in figuring out what is worth studying.
caaqil 45h ago on HN
> what do I do?

Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.

aaraujo002 45h ago on HN
This happens all the time, even without AI. Other researchers or PhD students can publish the same results before you. I say that based on my experience during my PhD.
CaptainNegative 42h ago on HN
Yes and no. It's pretty normal for students to get scooped, sometimes even multiple times (ask me how I know...).

It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.

There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.

dcl 45h ago on HN
This has always been a challenge for PhD students and researchers, it's just far more likely to occur now it seems. Getting scooped doesn't feel good, but it's a signal you've been thinking about things other people care about.
dekhn 45h ago on HN
Let me give you some perspective: my entire phd was made obsolete by CRISPR. It was a wonderful thing.
thimotedupuch 45h ago on HN
Interesting. If you don't mind, could you please share a little bit about that ? You already finished your dissertation ? It was about the works of Doudna and Charpentier ?
dekhn 45h ago on HN
No, back in the late 90s and early 00s, people were trying to engineer custom nucleases and transcription factors, my work was on doing molecular dynamics simulations to optimize TF sequence specificity (similar to engineered zinc fingers) for gene therapy. I wrote up my dissertation and published it in 2001, and then went off to find enough compute, IO, and smart people to make it happen (https://research.google/blog/groundbreaking-simulations-by-g...).

My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.

vasco 43h ago on HN
So it's not the situation they described at all, you didn't waste time during the PhD having to scramble to change topics as it happened after you were done.
globular-toast 33h ago on HN
That, and also it's just a completely different approach which might later on turn out to be useful. People should remember that artificial neural networks were developed decades before they were useful. People were doing all kinds of other approaches to ML like support vector machines before advances in hardware made deep neural nets feasible and therefore interesting again. ANNs were never obsoleted by SVMs.
dekhn 29h ago on HN
Actually, I'm pretty sure SVMs were obsoleted with ANNs (not just in terms of UFFs).
boznz 43h ago on HN
For every door that shuts another one opens - great if you're not a cabinet-maker.
vinyl7 45h ago on HN
Look forward to being obsolete I guess
bobmarleybiceps 45h ago on HN
I think eventually companies won't get as much stuff that's usable for marketing, so they'll stop investing so much into ai for math, so eventually cheap and poor graduate students will be able to do relevant work again without worrying about getting scooped by a company with a million GPUs :-/
moralestapia 45h ago on HN
That would be unfortunate but the world does not owe you anything and is not going to stop for you. Which is also a valuable thing to learn in your 20s (ideally earlier).
hgoel 45h ago on HN
It could still be interesting if your approach to the problem was different to theirs.
claaams 45h ago on HN
Don't worry, if you use openAI and get lucky they might offer to share credit with you for your work.
yieldcrv 45h ago on HN
Yes, and?
glitchc 45h ago on HN
Perhaps consider switching to a more applied field. Experiments in the physical realm hold value, especially if you document the process.
ex-aws-dude 44h ago on HN
That’s always been a thing, it’s called “getting scooped”
vouaobrasil 43h ago on HN
Killing with knives has always been a thing. Now, we have the machine gun.
ex-aws-dude 42h ago on HN
The scoop gun
bamboozled 44h ago on HN
Ask OpenAI for money when you don't have a job or future?

I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.

vouaobrasil 43h ago on HN
> Let's hypothetically say I'm a PHD student who is half way through my studies and I have a halfway written version of one of these "preprints" - what do I do?

I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.

s3graham 42h ago on HN
You might enjoy https://asteriskmag.com/issues/15/so-you-think-you-could-be-... if you didn't see it recently.
pratikdeoghare 43h ago on HN
> what do I do?

Very hard question.

Your work makes you one of the very few people who really understands the problem and solution and its significance.

netsec_burn 42h ago on HN
Verification is equally important, if not more so.
porcoda 40h ago on HN
As others said, it's not that this isn't a phenomenon that is unique to now. It happens. I had to pivot a bit of my dissertation near the end because at a random conference I spoke with a researcher from another continent and realized one of my ideas was already out there in some form. I just missed it since it was in a conference proceedings outside the usual set I looked at. So, I had to scramble to adjust and still come up with something novel. I survived, and defended, but it wasn't that much fun at the time.

What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.

I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.

rubikscube09 38h ago on HN
math will just be black boxed away. no one will "need" to understand it.
katatue 39h ago on HN
At my (German) university, the solution would have been to change from a "cumulative" thesis (which requires peer-reviewed papers) to a monograph-style one. Because the general rule was that your undertaking must be novel ar the time you submit your thesis topic, not necessarily at the tome the thesis itself is submitted (years later).
DCKP 36h ago on HN
I have had this conversation with my PhD students yesterday. I am 100% sure that all of their problems can be solved by publicly-available models now (I solved a case of one myself as a test, it took 15 minutes). So the challenge for them is to see how much they can accomplish in their allotted period, and still pass a defence on at the end of it all. The PhD defence is going to become all about a test of understanding, not a test of quantity of publication.

A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.

rfgplk 36h ago on HN
Frankly, I believe PhDs don't have to be novel in the absolute sense, only new research to the student. I can't remember how many times I've effectively invented something from a clean room approach only to realize someone already published something years ago or that my "new" algorithm has a name. So this really shouldn't be holding anyone back.
fooker 32h ago on HN
The purpose of a Phd is to train researchers, not to solve one small problem.

Some math PhDs would spend a year or so doing things with AI and lean, and graduate. And keep doing more math afterwards.

Some others, with more stubborn advisors, will keep trying to find a gap where there's no AI progress.

CS subfields go through this every ten or so years.

ghm2180 31h ago on HN
The follow up question then naturally would be how do phd advisors with people whose fields are in someway premised on making breakthroughs in theoretical fields that AI can solve work deal with it?
Davidzheng 28h ago on HN
If you had halfway to one of these papers you would be anyhow be in the top echelons of math phds so probably you have less to worry than most!
ks2048 45h ago on HN
I think they should put human names on the papers as someone who has reviewed the result, even if just a preliminary review. (I’m assuming they didn’t just pipe their model output directly to the internet and these had some amount of review?)
xpct 45h ago on HN
Presumably they don't because they're training the audience (us) to trust the machine, not its verifiers, even if they were included.
alexgoodhart 45h ago on HN
I appreciate you acknowledging the politics behind this. OpenAI does not intend to be a software company for long, they intend to be scientific infrastructure. They'll want to be faucet from which pours embryos, orbital calculation, geothermal/substructural rating, and etc.
robotpepi 33h ago on HN
> they intend to be scientific infrastructure.

beyond naive.

agnosticmantis 45h ago on HN
Long term this will be the only reasonable author list: Chad G. Peter {1}, Mat H. Lean {2}.

1: Author 2: Verifier

/s

chiwilliams 44h ago on HN
There are competitive reasons that they don't want to share all the people on the team.
chrisjj 44h ago on HN
> I think they should put human names on the papers as someone who has reviewed the result

Assume the empty list you see is complete. :)

make3 43h ago on HN
I think it's a damned if they do, damned if they don't scenario. I think they would be afraid to seem like they're claiming that the selected scientists deserve the praise of an invention or something, when it's "just the machine who did the work"
ks2048 43h ago on HN
Yeah, that's probably right. But, it's a "new world" - invent a new standard - "Written by Claude Foo X.Y; initial review by John Doe".

With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).

procedurecall 42h ago on HN
Looking at the quality of the writing in these documents (or at least, the ones that relate to problems I've worked on), I do not think humans reviewed them.
kzrdude 40h ago on HN
And that means that the recommendations of the reference group has not been followed. The only improvement here is that it is version tracked? No authors, no careful write-ups with exposition.
rubikscube09 38h ago on HN
the reference group can recommend all they want, no one will review 700 plus papers.
aaraujo002 45h ago on HN
The Advisory Group states in its recommendations [1]:

"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."

To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?

[1] https://agmai.org/general-sep29/

osiris970 45h ago on HN
Comical ask
medler 45h ago on HN
The rest of that document makes a pretty compelling case for why this is a bad practice
esafak 45h ago on HN
I fear that professional mathematics will wither, and there will be nobody left to digest the AI results of the future, leaving us unable to challenge the AI.
pixl97 40h ago on HN
Then make 2 AI's and force them to challenge each other.
warkdarrior 45h ago on HN
The latest posts from Terry Tao on Mastodon effectively ask for an AI to explain its results to human mathematicians.

> "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"

https://mathstodon.xyz/@tao/117395269325940185

binlog 41h ago on HN
There are a dozen+ AIs available to you that can do that right now.
jhrmnn 45h ago on HN
It all hinges on the definition of “progress”. The debate of the past month is all about questioning whether formally proving outstanding unproven theorems without human understanding constitutes progress. This is quite different from solving diseases.
andriy_koval 39h ago on HN
without humans understanding who are currently losing entitlement. Regular Joe could never understand or claimed to understand high math.
mattr03 45h ago on HN
I don't think this is a reasonable take at all. Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art. Nothing is gained from OpenAI solving all these problems but taking jobs from mathematicians. Other than advertising for OpenAI at least. It could not be more different from having AI work on disease research etc.
bravoetch 45h ago on HN
> Most work on maths has no real benefit other than to further human understanding of maths - it's more like an art.

It's been a while since I was reminded of this xkcd: https://xkcd.com/435/

zeroonetwothree 43h ago on HN
Most math isn't just proving novel famous results. Just like most of software engineering isn't writing code.
fph 45h ago on HN
...but we're not talking about diseases. Publishing an AI-generated Navier-Stokes solution does not save lives. (And, in fact, it harms some.)
Yamata 39h ago on HN
It harms lives? How so?
fph 22h ago on HN
Many mathematicians I know are shaken or depressed by these news of machiines that might put them out of business. Especially those who don't have tenure yet.
tchalla 45h ago on HN
Why did you leave out the entire quote?

> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.

To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.

aaraujo002 45h ago on HN
Maybe, but it still seems like an excuse. OpenAI has a proprietary model capable of solving these problems and is willing to share the results with the mathematical community. So basically, the ask is to just not use the model and leave the problems unsolved?
adrian_m 45h ago on HN
The ask is to let mathematicians outside of OpenAI use it, ie. at least wait until the model is released.
agnosticmantis 44h ago on HN
Which mathematicians though? Only fields medalists? Grad students? Any hobbyist wanting access?

These models are too expensive for broad access unfortunately.

Jtarii 44h ago on HN
ChatGPT pro is accessible to literally anyone who has a job and lives in a developed country.
Jweb_Guru 43h ago on HN
Literally every single one of these papers was developed with an internal model that not even most OpenAI employees have access to.
anthonyrstevens 29h ago on HN
Why is this important? Do you not think that at some point, probably sooner rather than later, math researchers across the world will have access to similar capabilities?

This is sort of like discounting putting humans in space because only a few nations have the means to actually do it at the moment.

strange_quark 44h ago on HN
I don’t think that’s sufficient. If they want to be good stewards of mathematical research and not just doing marketing, they need to at the very least tell us the datasets and any techniques they used to train this model. IMO that’s probably just as if not more valuable than a dump of un-reviewed results.

They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.

blurbleblurble 41h ago on HN
It's honestly shit marketing that only cultivates spite and erodes their whole brand. This is a total ego trip.
TeeWEE 43h ago on HN
The work is not having AI poop out these docs the work is validating them and publishing them. OpenAI doesn’t care and wanted to race to publish potential findings and let others review it which is lazy and selfish
perching_aix 45h ago on HN
The trope you're drawing a parallel with has a (to me) compelling counter though: there being a cure for every disease wouldn't stop people from getting sick.

How this maps back to math, idk.

bmitc 45h ago on HN
Advocating purely for progress and not humanitarian value is how we'll all get enslaved.
pavitheran 45h ago on HN
From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”
password54321 45h ago on HN
Oh cool, we will all now have a math genius on our computer.
an0malous 45h ago on HN
Well, on their computers. But you can rent them for a price.
binlog 41h ago on HN
An open source model will reproduce it 6 months later
robotpepi 33h ago on HN
which you an run IF you have the hardware. who knows how heavy these models are.
jrflo 45h ago on HN
It was using their internal math model, so not yet for us
password54321 45h ago on HN
I used future tense. It was implied this will be available.
scrlk 45h ago on HN
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
inferencecoder 44h ago on HN
It doesn't imply that, it's just measuring the amount of compute.
bigmadshoe 44h ago on HN
But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?
inferencecoder 40h ago on HN
Not necessarily, could be agent swarm with low N
bigmadshoe 40h ago on HN
At some point that stops being a "swarm" and just a handful of subagents. E.g. 6 agents running for 30 minutes each isn't really a swarm in my eyes.
orlp 45h ago on HN
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

timjver 45h ago on HN
>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]

That doesn't sound right

orlp 45h ago on HN
Oops, edited.
machomaster 45h ago on HN
They did say that. "3 hours of ChatGPT Pro thinking compute"
orlp 44h ago on HN
Yes, what does that mean?
mh- 39h ago on HN
It means the level of effort that a ChatGPT Pro plan summons when thinking. For 3 hours.

If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.

pixl97 44h ago on HN
Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
orlp 44h ago on HN
I'm not denying that, but I'd still like to know what that cost.
Jtarii 45h ago on HN
That estimate is obviously going to conveniently ignore all the failed runs.
mathisfun123 45h ago on HN
With so many results in so many different areas no way they even remotely spot checked well enough.

Prediction: one of these is wrong and this (publicity stunt) will backfire.

Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.

bravoetch 45h ago on HN
What does a backfire look like? It's ok to be wrong in the science/math world.
mathisfun123 45h ago on HN
of course in science/math it is but it's not okay if you're a business selling supercalifragilisticexpialidocious infallible intelligence.
bravoetch 45h ago on HN
Do they claim that's the case? I don't think they do.
mathisfun123 45h ago on HN
does company A making product B claim that the product is robust and consistent? is this a serious question?
zamadatix 44h ago on HN
If you were waiting for companies to sell engines that never break down you'd still be stuck pre industrial revolution while the rest of the world has been to space.

The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.

anthonyrstevens 29h ago on HN
2023: "AI is useless" 2026: "One of these solutions to dozens of previously-intractable problems at the frontier of human knowledge MIGHT be incorrect"
jojva 45h ago on HN
You have not read their readme:

> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.

mathisfun123 45h ago on HN
i have and i'm exactly saying that if it comes to pass one of them is wrong it's going to backfire. ie yes that's my exact point/bet.
stevenhuang 44h ago on HN
I don't think anyone would particularly care if only one of them is wrong, if most are correct.

If they are all wrong, that's when it would backfire.

orlp 44h ago on HN
It's likely that way more than just one of these is wrong. But even if it turns out 80% is wrong this is still 100+ results...
dekhn 45h ago on HN
I'm a software engineering/biology/ML guy who loves when clever math ideas get turned into real solutions (https://en.wikipedia.org/wiki/Compressed_sensing). I am curious if any of the results have immediate applications in any kind of engineering or science.

It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.

brandonpelfrey 44h ago on HN
Same. I have agents analyzing the papers here to see which of these are actually new approaches, novel application of unusual approaches, etc. If there are new intuitions and ideas, those are the mostly powerful reusable components by my estimation.
OutOfHere 43h ago on HN
Please share your findings.
qnleigh 37h ago on HN
I was talking with a number of physicists this evening, and of the papers who's problem statement we understood, I don't think any of them will have near-term applications. For most of them, they were results that I think the physics community believed to be true, and the paper provides the first rigorous proof. For several it was news to me that they weren't already theorems! These results are important advances in mathematical physics, but they don't tend to have much impact on experiment.

From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.

autuni 36h ago on HN
It feels odd to me that they wouldn't prefer applied problems. Seems like an easy way to profitability. Probably based on what attributes they're looking for in a problem when picking them.
hnfong 36h ago on HN
You can't really profit from proving theorems of applied problems (that are widely regarded to be true). Those who need to apply those theorems on real problems would have already done so (and if they don't work in some cases, well, congratulations... you found the counter example!)
kidel001 27h ago on HN
At first this was my take as well. That plus, well, maybe they are just on a serious PR kick with maths. But I am starting wonder if they have determined, or strongly suspect, that the road to exponential model improvement must first be paved with extraordinary improvements in math. Like in some sense this seems like a test case for where their true intensions might go: vast improvements in the efficiency / size / speed of models and their training. Hard to imagine trusting the models in all those spaces without first trusting them / training them to address new or unsolved math.
qnleigh 13h ago on HN
This is a really interesting observation, especially given their current valuation and circumstances. Probably the verifiability of mathematics is enabling this progress, and making progress on science and engineering is not so straightforward. You can be sure that this is a top priority for them, especially now that there is concrete evidence that LLMs can achieve super-human performance at math. A model that achieves similarly super-human performance at things like material or drug design would be incredibly valuable. In fact, maybe this is why Anthropic hasn't been making such enormous strides in mathematics - they have been prioritizing biology instead.
kevinwang 45h ago on HN
wow
prideout 45h ago on HN
This includes a proof of Barnette's Conjecture, which is one of the graph theory conjectures that I tried attacking with SOTA models a few months ago. I like it because it is easy to understand with a basic knowledge of graph theory. I spent quite a bit of time on it and failed. Their proof looks approachable at first glance.

https://github.com/openai/math/blob/main/preprints/Paired-st...

an0malous 44h ago on HN
Any idea what made OpenAI successful where you weren’t?
kulahan 44h ago on HN
Trillions of dollars might be a bit of an advantage.
martinky24 40h ago on HN
Trillions?
KyleTheDev 28h ago on HN
Quadrillions?
ForHackernews 44h ago on HN
They ingested all of his sessions with their SOTA models from a few months ago. ;)
digitaltrees 44h ago on HN
The fact that this is plausible should be deeply disturbing and disqualifying for openAI. The fact that they may prevail and win is a travesty of our failed system.
zeroonetwothree 44h ago on HN
What does "win" mean? There is no prize for this, and having someone discover a proof benefits us all.
breezybottom 44h ago on HN
Sure there is. A job, tenure, professional respect, Fields medal.
vuurmot 44h ago on HN
The prize is a tenure for the researcher, and in OpenAI's case, a higher valuation when they IPO?

In this case, the tenure is gone, and OpenAI has increased their valuation

digitaltrees 41h ago on HN
Win means being able to monetize the intelligence they have created by exploiting the past present and future collective intelligence of humanity to amass wealth and power without regard for the debt they owe
fnordpiglet 44h ago on HN
Disqualifying for what? If you develop a proof you aren’t competing for something, you’re expanding the frontier of knowledge. It’s a binary state of the world, either it’s proven or not. Prevail and win what exactly?
digitaltrees 41h ago on HN
Disqualifying for participation in civil society and the social contract. Why do they get to participate in and receive economic benefits, be shielded from liability, and effectuate their will to amass more power and influence such as monopolization of computer, training data, capital other resources. I have multiple founder friends that have been told firms are allocating less capital because they are reserving it for the OpenAI and anthropic IPOs.
andriy_koval 27h ago on HN
> Disqualifying for participation in civil society and the social contract.

US/world deviate more and more from social contracts. Money matters way more, and OpenAI wins here.

IsTom 36h ago on HN
The use of supposed little ways multiple people pushed the envelope in their sessions with no attribution whatsoever.
AIblemblio 35h ago on HN
Its a capitalistic issue, not a company/technology issue.

If we would have discovered this breakthrough of LLM/ML on scale in a non capitalistic world, we would all work together advancing it faster than it goes right now for the benefit of humanity.

And I don't live forever (at least for now) i def want to see were this road is heading.

Its a conflict of interest for sure, a cnflict of the future of a lot of humans

permalaise771 33h ago on HN
There's that great Ted Chiang quote: "Most of our fears about technology are better understood as fears about capitalism."

I am always looking for leftist writing imagining a positive vision for AI. Is there any which you'd recommend?

nl 31h ago on HN
By leftist do you mean progressive or left wing economically?

There's quite a lot of progressive positive writing.

On the economic side the Australian Council of Trade Unions statement is about a positive future: https://www.actu.org.au/speeches-and-opinion/joint-statement...

digitaltrees 32h ago on HN
I hear you and would have what would probably feel like pedantic push back (we are not in capitalism as much as the unchecked end result of unregulated capitalism that becomes monopolistic corporatism), but it’s hard for me to see how this would be different in mercantilism, feudalism, or even communism as the human tendency to seek and hoard resources is universal for some fraction of people so as long as that confers an advantage then AI would be used by the designers in antisocial self enrichment.

I’d love to hear more thought experiments if you have some so my failure of imagination or lack of awareness can be overcome.

AIblemblio 30h ago on HN
My personal system would be based on resource points: Define the amount of resources our planet has in a sustainable fashion, everyone gets the same and can use them how they like.

Technocracy had this already in form of Energy accounting.

Unfortunate something like communism sounds similar and just because we have seen that it didn't work due to technology issues (planning ahead without necessary information is hard?) and no gain of function which would push people, the basic idea is similiar to energy accounting.

Another thing this system needs might be a way for the system to protect itself.

I do think so that a society as diverse as ours will continue struggling with this as long as we do not give abundance resources to everyone or educate/indoctrinate people the 'right' way.

One thought I have regarding AI: IF it happens to slow, people will get used to the status quo and inequality and we will see a future of a handful rich people and a lot more poor people. IF it happens too fast, people might be more desprate to standup and demand something better.

whamlastxmas 44h ago on HN
Their internal model is allegedly like 4x as capable as the publicly available ones
sebzim4500 44h ago on HN
Presumably it's mainly the better model, I don't see much evidence of a particularly advanced harness based on the reasoning traces that they provided.
an0malous 43h ago on HN
That’s what I was wondering. Thanks.
seanmcau 44h ago on HN
Probably the model OAI used that is strictly better than whichever SOTA - 3 months model OP used?
redanddead 30h ago on HN
Nah.

The model is probably comically big and inefficient but big enough

Finally, size really does matter!

ZephyrBlu 29h ago on HN
The internal model they used to solve the Navier-Stoke's problem was significantly better than the public Astra model, and they also used 10,000 agents.

Astra wasn't even released 3mo ago. It would not be surprising in the slightest that a public model from 3mo would not be capable of solving this problem, but an internal one from current day would be.

andriy_koval 27h ago on HN
my speculation is that they have math-specialized model retrain, so it doesn't need to have all world info in weights, but can focus on math RL training.
zzzeek 43h ago on HN
I'm going to guess the ability to hold a million individual details in an attention space at once, compared to the typical human capacity for about six or seven
TeeWEE 43h ago on HN
Did you validate the proof? Who did?
jboggan 41h ago on HN
I've been messing with that problem since 2002. I'm curious if you were trying the dual spanning tree direction (which is what the purported proof is using) or working with cycle construction on the original graph. I was working heavily with edge-Kempe swaps but couldn't quite get there.

I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.

kingstnap 45h ago on HN
Some of these are interesting ngl.

109. Integer multiplication below n log n

Surprising that this is possible.

158. The Euclidean plane cannot be colored with five colors.

Only 6 and 7 remain!

376. Universal computation in forced Navier–Stokes flows.

Morning coffee proven turing complete

mFixman 45h ago on HN
> We give a deterministic algorithm that multiplies two n-bit integers in O(n (log n)^(1−κ)) worst- case time, with κ = 2^(−182).

LMAO, I don't think I ever saw such a small number in a CS result.

sobellian 45h ago on HN
I am fully braced for it to be a https://en.wikipedia.org/wiki/Galactic_algorithm

Very surprising result though! Multiplication is easier than sorting.

senderista 44h ago on HN
It would be absolutely unbelievable if such an improvement were practical.
zeroonetwothree 44h ago on HN
Then 'n' means kind of different things for sorting vs. multiplication though. For example for sorting we assume constant time comparison, which doesn't make sense inputs of O(n) bits
sobellian 42h ago on HN
If you sort n k-bit items for a total time of O(nk logn), that scales more poorly in n than multiplying n-word integers. Of course if k is constant you can do radix sort, but I genuinely don't know under what conditions radix sort is more/less galactic than this multiplication algorithm.
adgjlsfhk1 40h ago on HN
this alg is way more galactic than radix sort. radix sort often wins in the hundreds of elements. the nlogn multiplication requires numbers with more digits than atoms in the universe (although that could probably be brought down a lot)
sobellian 40h ago on HN
Ah thanks for pointing this out, for some reason I had always equated radix sort and bucket sort (with 2^k buckets) in my head. But I learned today that this isn't true!
kingstnap 45h ago on HN
Yeah its ridiculously small, but any improvement on n log n is wild.

Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?

Which low and behold ->

130. Fourier transforms below n log n.

xyzzyz 45h ago on HN
They also separately give algorithm for Fourier transform over complex number faster than O(n log n)
saalweachter 44h ago on HN
Wikipedia just told me there's a galactic algorithm for integer multiplication in O(n log n) based on FFT so I'm guessing those two proofs are related.
pfdietz 40h ago on HN
Multiplication is a lot like convolution, so the connection is natural.
rubikscube09 38h ago on HN
multiplication is implemented w the fft
anon-3988 44h ago on HN
It fascinates me that there's something like this in something as solid and rigid like matrix multiplication. What causes something so rigid to break apart and "leak" at very large scale? Why does the "optimization" appear to be very, very small? Why does galactic algorithm exists? I can't imagine long division suddenly breaking apart after a billion digit, the structure seems very stable? I have heard before that matrix multiplication is apparently optimize-able at very, very large scale.

Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?

adgjlsfhk1 43h ago on HN
One way to think about it is that the classical algorithms are the ones that are fast for small numbers. Galactic algorithms often work for small inputs, it's just that to be faster you need big inputs. A common case of this is a requirement that log(n)<<klog(log(n)). If k=100 then this algorithm will take huge sizes to win
HarHarVeryFunny 28h ago on HN
Can anyone ELI5 to make it make sense?

It seems n would have to be unimaginably large for this to make any difference. What changes about multiplication / FFT at large enough size ?

I guess nobody expected that it did before this result.

anabab 11h ago on HN
even though the constant is tiny, this breaks the long held assumption that nlogn is the minimal possible bound, so further research seemed unfruitful. this discovery will now trigger more research in this area, and probably more optimal algorithms will be discovered.

Like matrix multiplication, the common assumption was that it cannot be improved past n^3. Then Strassen broke the barrier (with a more significant constant) which caused intensive research - and now we are around n^2.3..2.4.

zeroonetwothree 44h ago on HN
Integer multiplication is very unexpected, I think most people believed in the n log n lower bound!
tootie 44h ago on HN
Note that these are all preprints. None are verified.
FuckButtons 42h ago on HN
Other than the by the lean certificate you mean.
measurablefunc 41h ago on HN
Lean has bugs & proofs of ⊥ that have gone undetected previously.
jaykru 40h ago on HN
many of these are not accompanied with leanslop
mi_lk 45h ago on HN
Curious if Sébastien Bubeck still work at OpenAI? He came out quite dirty after Navier-Stokes drama
yewenjie 45h ago on HN
A lot of these seem to be proving conjectures rather than finding counterexamples, a lot of people used that to claim that these models are not really smart/creative etc.

That copium didn't last for what, three months?

sebzim4500 44h ago on HN
Don't worry, more copium will be delivered. TBF so far it's still only solved the easiest of the millennium problems.
justanotherjoe 25h ago on HN
That's sort of how problems work.
zeroonetwothree 44h ago on HN
I acknowledge I am impressed how quickly it moved beyond just counterexamples.
redox99 45h ago on HN
The stochastic parrots have predicted the next token once again.
anthonyrstevens 29h ago on HN
How deep must your head be buried in the sand to trot out this comment, on this thread.
simianwords 26h ago on HN
the post was clearly sarcastic
connor11528 45h ago on HN
will this make the math for building data centers work?
sashank_1509 41h ago on HN
No that’s gonna happen when they take your job
plaidfuji 38h ago on HN
Underrated joke
foota 45h ago on HN
From their github: "The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking." That's pretty crazy.
dyauspitr 41h ago on HN
Crazy because that’s almost nothing right?
foota 41h ago on HN
Yes
johnisom2001 28h ago on HN
This is what really made me think to my self, "holy shit". I can't believe not more people are noticing and talking about this. Unless perhaps they didn't actually read the README, and are just talking about what they heard from someone who also didn't read it?
oblio 8h ago on HN
People don't believe this because it's false, those 3 hours are just the last session connecting all the dots. They've tried a lot more, probably thousands of hours, in total.
foota 2h ago on HN
Honestly? I don't think so. If anything I think they would obscure this using averages (e.g., maybe they solved some hundred problems and the easy ones took a few minutes and the hard ones took a couple of days), but if they have a model that's as good at math as they are as coding I don't see why it'd be impossible.

I think novel mathematics is intrinsically harder than coding (much in the way that computer science research is harder than coding), but the models are scary good at coding.

zone411 45h ago on HN
A quick check shows that this list claims to fully solve 90 of the top 500 open problems in math (https://proofatlas.ai/open-problems/).

The highest ranked would be:

| 22 | Hilbert’s tenth problem over ℚ |

| 29 | Unique Games |

| 31 | Anderson-model extended states |

| 37 | Spacetime Penrose inequality |

| 48 | Nonexistence of Landau–Siegel zeros |

| 52 | Baum–Connes |

| 78 | Abundance |

| 80 | Hadwiger |

| 87 | Bose–Einstein condensation |

| 92 | Two-dimensional entanglement area law |

anematode 45h ago on HN
Dear lord that website is laggy
manquer 44h ago on HN
At this rate solving P=NP is going to be easier than solving front end perf …
m_mueller 44h ago on HN
wait, maybe this is the same problem....

with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...

mswphd 40h ago on HN
worth mentioning that "NP" is not "non-polynomial" but "non-deterministic polynomial (time)". If NP was non-polynomial time then NP != P would be trivial (and in fact, P != EXP is known by the time hierarchy theorem).

Non-deterministic can be explained in several ways. One is in terms of a hypothetical "nondeterministic Turing machine" with certain non-physically realizable properties. The easier way is that a NP problem gets as input not only the problem instance x, but a "witness" w, that may depend on the problem instance. This witness generally makes the problem of deciding the problem instance straightforward (e.g. for SAT, x is the SAT instance, and w is a description of how to set the variables so that it is true).

sebzim4500 29h ago on HN
I reckon I could tell you in polynomial time whether a div was vertically centered, not sure if I could write the CSS in polynomial time.
echelon 43h ago on HN
Please let P=NP, Please let P=NP

Whomever is running this simulation, please.

osti 42h ago on HN
It's math, the result shouldn't be different just because it's a different sim.
brookst 40h ago on HN
Depends how fundamental the variables are. If we can code a sim for a topos[1], why can’t we be in such a sim?

1. https://arxiv.org/pdf/1012.5647

mertyildiran 40h ago on HN
Well if the fundamental constants or hidden variables of the universe are shifting because of his comment then it can change the outcome.
NooneAtAll3 30h ago on HN
unless mechanism behind our universe dynamically alters our logic on the fly to be artificially self-consistent
qarl 26h ago on HN
To be fair - there are statements in math that are independent of the axioms. For those statements, the universe you find yourself in can pick either version (true OR false) and still be consistent.

See also: noneuclidian geometry and axiom of choice.

sm-silversight 41h ago on HN
Why?
adrianN 41h ago on HN
Being able to solve NP hard optimization problems would enable progress in many areas of science and technology. For example it would allow us to find poly-sized Lean proofs for theorems efficiently, since proof verification can be done in polynomial time.

It would also be amusing to annihilate nearly six decades of proofs that assume P!=NP.

manquer 41h ago on HN
Could also break the basic principles underlying most encryption approaches. I would rather have my bank account not stolen and internet working
echelon 40h ago on HN
I've had enough Internet for one lifetime.

As long as we also get low order polynomial solutions to important problems, it'll be worth it.

Besides, unencrypted wifi was funny.

mswphd 40h ago on HN
to depress you even more, it is consistent with everything that we know that P != NP and that cryptography does not exist. So there is a worst of both worlds, and we cannot rule it out.
charcircuit 41h ago on HN
Even if P=NP it doesn't mean that the P approach will be better than the heuristic approach we already do today.
adrianN 40h ago on HN
Of course, if we get ridiculous polynomials it doesn't mean much in practice. People who hope for P=NP generally hope for nice polynomials O(n^3) or something like that at worst.
black_knight 39h ago on HN
Leans proof checker is not polynomial time, unfortunately. It is super exponential. Basically, because it can verify the result of any function it can prove to be total.
adrianN 37h ago on HN
Oh that’s unfortunate.
sebzim4500 31h ago on HN
That's fine, we just change the problem from "find a lean proof of length < f(n)" to "find a lean proof that can be validated in time < f(n)".
vector_spaces 44h ago on HN
Not to mention it's got that signature Claude Clutter UI design
p-e-w 44h ago on HN
Interesting how perceptions differ. My first thought was “Wow, that’s well designed for a math website”.
zone411 44h ago on HN
Except that Claude wasn't used.
sourcopolo 43h ago on HN
Probably Copilot then
optimalsolver 44h ago on HN
Was anyone in the math community aware of the inbound tsunami at the beginning of the year?
thrance 43h ago on HN
I predicted, over 2 years ago, that theorem proving would fall way before other problems that people believe are harder.

https://news.ycombinator.com/item?id=41072330

mag7269 42h ago on HN
Fucking even called LEAN the “hottest shit under the sun”—which it is. You, legend you!
bice 41h ago on HN
There was a Wired Magazine article from either the late 90s or early 2000s that made a prediction that this sort of thing would eventually be possible, likely within my lifetime. I believe the context was "distributed computing" models of the time, like SETI.

I've never been able to find that article as an adult, but I would love to know who wrote it.

schoen 41h ago on HN
Some candidates suggested to me by an AI:

Gina Kolata allegedly in the New York Times in 1996 on the Robbins conjecture (noting that computers had started to contribute to math research in some sense), and a longer piece in Math Horizons by her the following year ("Computer Math Proof Shows Reasoning Power"). I didn't immediately find the NYT article, so I don't know if it might be a hallucination.

John Horgan in Scientific American in 1993 (https://www.scientificamerican.com/article/the-death-of-proo...). There's also a retrospective on the topic by the same author in Scientific American in 2022 (https://www.scientificamerican.com/article/should-machines-r...).

Natalie Wolchover in Quanta (but reprinted in Wired) in 2013 (https://wired.com/2013/03/computers-and-math).

I was involved in some distributed computing stuff in the late 1990s and early 2000s and I don't really remember people in that community talking about proofs but there may have been a "if we had a mechanical proof-checker, could we do distributed searches for valid proofs that it would accept?" conversation somewhere at some point. There were definitely volunteer distributed computing projects working on pure math; I remember the Optimal Golomb Ruler search (https://en.wikipedia.org/wiki/Golomb_ruler). So, that could possibly have shaded over into "can we find proofs this way too?". At the time it probably would have been based on brute force searches through proof space rather than clever optimization, though.

The idea that you can lexicographically list all proofs in some formalism and then mechanically determine if any is valid is quite clear from Gödel's construction of the function Bew in "On Formally Undecidable Propositions", but he points out that you don't know where to stop because you don't know how long a valid proof would potentially have to be (so "is this a valid proof of this claim?" can be decided mechanically in a limited time, while "is there any valid proof of this claim?" can't be! maybe the shortest valid proof is 49 steps long but you eventually stopped checking after looking at all 7-step proofs, or something).

bice 40h ago on HN
Thanks! Yea, I have used various LLMs to dig for this article, as well as Google search multiple times over the past 20 years. The article I'm remembering was 100% prior to Nvidia's CUDA release in 2007. My best guess is that it was from late 90s, but possibly early 2000s.

The article I'm remembering was not just about mathematics, but indeed all of physics and related fields. I believe it speculated that eventually distributed computing models could essentially take the world's mathematics and physics formulas and various datasets that we believe to be accurate with high degrees of confidence, and then look for patterns or trends, and then from those trends, mathematicians and physicists would be able to investigate further. Not dissimilar to Folding@Home and SETI@Home.

Keep in mind, that this is the best I can remember from 30 years ago, and I've thought about it so frequently that I am certainly misremembering some of the details. Anyways, it's always been this really compelling possibility, and I wish I could find that article that inspired me so long ago and re-read it! :) I really think it was Wired, but it's possible it was Popular Mechanics, or even an expert guest on TechTV who gave an interview. Hard to say for sure, but I've always thought it was a Wired article.

Appreciate your help though!

schoen 39h ago on HN
Oh, I remember hearing about "computer scientists" or something that would attempt to determine physical laws on the basis of empirical evidence, possibly also in that timeframe. That might be another thing to look for. I'm sure that's something people were writing about.

Edit: with the noun-noun compounding being different from the usual interpretation here, like "scientists who are computers" rather than "scientists who study computation"! Maybe "computerized scientists" or something.

pillefitz 39h ago on HN
Ted Kaczynski,the Unabomber, made the same prediction 30 years ago.
AnotherGoodName 42h ago on HN
Lots. To give an example Terrance Tao was lambasted skeptics on this site for stating it in 2024.

https://unlocked.microsoft.com/ai-anthology/terence-tao/

" I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.

Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?

We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."

He's pretty damn smart that guy.

mianos 42h ago on HN
> He's pretty damn smart that guy. This is probably the understatement of the year. I am literally ROFLing.
mertyildiran 40h ago on HN
Terrance, Reinmann and Hebert walks into a bar...
pseudohadamard 42h ago on HN
And do any of them actually matter? Will the fact that Noodleheinz's Third Postulate now has a proof affect anyone?
bice 40h ago on HN
It's really impossible to predict which discoveries will "matter", have a direct impact on other fields, or an impact in making other mathematics or physics discoveries.

Only after a world's worth of experts look at these results and then mull over if and how their own fields are impacted by this new info will we be able to answer this question.

I'm reminded of a great TV Show, James Burke's Connections. Where discoveries in one area of science would revolutionize or fundamentally change a completely different area. https://www.youtube.com/watch?v=XetplHcM7aQ&list=PL5HjoPOFFC...

It can take decades to really know the full significance. You know, the whole "We stand on the shoulders of Giants", well the Giants just grew a few inches all at once.

aureianimus 41h ago on HN
I was at the workshop that resulted in the Leiden Declaration in Fall 2025. The majority vibe was that this was inevitable, but hard to predict whether it would be in one year or 30 years.
efficient_dairy 41h ago on HN
I guess they showed this to the advisory group they created. I guess the group tried reading the work for a day and they could only think of telling them to release the results to the community. I now understand why the group had this suggestion.
k2xl 44h ago on HN
Result 003 (Quasi-Riemann Hypothesis), from my reading of mathematicians reactions, is a landmark discovery.
adgjlsfhk1 44h ago on HN
yeah if it holds up, is the biggest result in number theory in 200 years
JoshuaZ 43h ago on HN
Number theorist here. This is a massive big deal, and would likely be a Fields Medal for a human if a human had done it. But it is an exaggeration to say it is the biggest result in 200 years. At a minimum, it is hard to argue that it is a bigger result than the proof of the prime number theorem in 1896 (which this is a strengthening of), or Riemann's original 1859 paper where he laid out the zeta function and its analytic importance, or Dirichlet's proof of infinitely many primes in arithmetic progressions which is the late 1830s.

But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.

asdfologist 42h ago on HN
How about 100 years?
howunfortunate 42h ago on HN
(unrelated: love your username)
JoshuaZ 42h ago on HN
Yeah, completely reasonable to argue that.
AmazingEveryDay 42h ago on HN
What is your favourite unsolved problem in number theory which if solved, would be more important than 1896 prime number theorem?
JoshuaZ 42h ago on HN
Generalized Riemann hypothesis.
gavagai691 41h ago on HN
I am also an analytic number theorist, and I disagree. Not only do I think Fields Medal is an understatement (Fields Medals have been awarded for far less than proving quasi-RH + no Siegel zeros), I don't think it is unfair to say that this is a bigger deal than the 1896 proof of the PNT.

As for Riemann's memoir, it's hard to compare. You could argue that was "just" noticing a connection (between number theory and Fourier analysis) that nobody had noticed before; in fact this is the kind of thing AI is extremely good at. I'm being a little cute here.

I think if a human had proven just these two results in the form of a uniform zero-free region for L(s,chi) from nothing as OpenAI did it would not be unfair to say that it would be the single greatest advance in math (easily dwarfing Wiles' FLT), and it would instantly put them in the ranks of greatest mathematicians of all time. Unlike something like Navier Stokes there wasn't a semblance of a research program, experts basically considered this hopeless and would have said the chance of seeing a proof in our lifetime was near zero.

For some comparison, Yitang Zhang's bounded gaps result might have gotten him a Fields Medal if he was not disqualified by age. When it was floated that he might have proven Siegel zeros don't exist, it was considered (by experts) clearly a much bigger deal. This result blows that out of the water (it's a way better version); at least analytic number theorists I talked to thought it was plausible but unlikely that Siegel zeros would be eliminated in our lifetime but thought RH was basically hopeless.

qnleigh 37h ago on HN
> it would be the single greatest advance in math

Did you mean to not qualify that? That is a bold statement indeed.

margorczynski 34h ago on HN
Thank you for the detailed explanation. From what I'm reading from a lot of mathematicians there's at least a dozen of results here that are field-definining and worthy at minimum of a Fields medal.

I guess the biggest news are not the discoveries themselves but how they were found and that math is going through the biggest revolution as a field since almost ever.

JoshuaZ 34h ago on HN
While some of my work is in analytic number theory, much is in other subareas, so it is possible I should defer to you on this.

It seems to me less than PNT in terms of what can we actually do with this. Many different areas of math use PNT, and from my standpoint, PNT is helpful not just for what it implies directly but because it lets us make really good heuristics about whether some sets are infinite or not, and what their rough size is. (Granted, one can do that also mostly via Chebyshev). For those purposes, this doesn't really enter in. Similarly, PNT feels like a statement at least I can say explain to my mother without any technical details. This isn't that. But that may also be my own biases of wanting things to cash out to very concrete statements about the integers.

I agree that one striking element is how no one saw this coming. This isn't building on an existing research program, which itself is remarkable. And last night, before I went to bed, I saw a conversation between a bunch of analytic number theorists who seemed to think there was potentially some slack in the quasi-RH argument, which if that's the case means this is going to go even further.

handle584 30h ago on HN
1896 PNT is basically 1859 Riemann + a trig inequality.

1830 Dirichlet's result is qualitative only, it shows infinitude but not the asymptote in terms of the zeros for it predates Riemann.

To me this is the first substantial step after the 1896 PNT, and we really do not see much progress in the whole 20th century. Personally so far there are only two people worth mentioning,

- Euler, introduces the real zeta function and Euler product, establishes the functional equation at (half?) integers.

- Riemann, introduces complex analysis ideas to the zeta function.

And of course this result if it is true. This is first to penetrate the critical strip, which nobody had any idea how to approach for over a century and a half.

omoikane 41h ago on HN
Did you mean this one?

https://github.com/openai/math/tree/main/preprints/The-Quasi...

I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.

https://github.com/openai/math/tree/main/preprints/The-Quasi...

mertyildiran 40h ago on HN
Funny that it says "written with human assistance" instead of saying "written with AI assistance". So we're assistants to the machines that we have created.
drnick1 37h ago on HN
In the same way that the driver is the assistant of a car?
NooneAtAll3 30h ago on HN
train engineer an assistant of a rail-following machine
magicalist 44h ago on HN
> the top 500 open problems in math

At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?

> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.

reasonableklout 42h ago on HN
Now I'm curious if there is such a site or article that ranks open problems based on votes from human mathematicians.
zone411 44h ago on HN
By category in the top 500:

  +----------------------------------------------------+------+---------+-----------------+
  | Category                                           | Full | Partial | Matched / total |
  +----------------------------------------------------+------+---------+-----------------+
  | Geometry and topology                              |   25 |       7 |         32 / 74 |
  | Algebra, representation and category theory        |   17 |       2 |         19 / 53 |
  | Analysis and PDE                                   |   11 |       6 |         17 / 40 |
  | Number theory and arithmetic geometry              |    4 |      13 |        17 / 117 |
  | Probability, ergodic theory and dynamics           |   11 |       5 |         16 / 37 |
  | Combinatorics and discrete geometry                |    7 |       2 |          9 / 34 |
  | Theoretical computer science                       |    4 |       4 |          8 / 57 |
  | Mathematical physics                               |    5 |       1 |          6 / 19 |
  | Applied and computational mathematics              |    2 |       2 |           4 / 8 |
  | Quantum information and computation                |    2 |       1 |          3 / 17 |
  | Cryptography, coding, information and optimization |    1 |       1 |          2 / 26 |
  | Logic, foundations and set theory                  |    1 |       1 |          2 / 18 |
  +----------------------------------------------------+------+---------+-----------------+
  | Total                                              |   90 |      45 | 135 / 500 (27%) |
  +----------------------------------------------------+------+---------+-----------------+
trebligdivad 42h ago on HN
I'm curious if they'll find any fun crypto maths holes/bugs.
errpunktjose 42h ago on HN
they are already lol
ajkjk 40h ago on HN
The most interesting for me were the faster matrix multiplication, integer multiplication, and FFT. Maybe just cause they're easier to appreciate.

There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.

nautilus12 45h ago on HN
Have any real mathematicians working on these problems reviewed any of these and determined if they are just gobbledegook or not?

The ones with lean proofs could still be formulated incorrectly

applicative 45h ago on HN
Why didn't they just give mathematicians access, so they could at least understand and write up the results in publishable form?
stevenhuang 45h ago on HN
zamadatix 44h ago on HN
I think they meant "access to the model" rather than the results.
oh_no 40h ago on HN
I'm seeing a lot of this and it makes no sense, the internal model solved these but give one of the papers to Astra and Opus and I'm sure it will have no problem recreating it.
zamadatix 31h ago on HN
I don't think they mean this from a "verify this paper" perspective.

How valuable it would actually be to share the model with other mathematicians vs just have OpenAI's mathematicians churn out and clean up results isn't very clear to me though as they don't say how much effort it's requiring from their team to prompt and clean these up vs how much it's bound by "time to run the model" or similar.

NotOscarWilde 45h ago on HN
As a TCS/scheduling person, this one is definitely of lesser importance than UGC, but it has been an open problem since the book of Garey and Johnson in 1979:

A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]

Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:

Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.

That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.

[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...

keeganryan 42h ago on HN
The largest I've seen [1] is an exponent of 10^12, which I suppose still counts as polynomial time.

I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).

[1]: https://github.com/openai/math/blob/main/preprints/Determini...

senderista 40h ago on HN
That's why I grimace when I see pop-sci descriptions of P as "all problems that can be solved efficiently".
nl 39h ago on HN
It's efficient, but somewhat slow..
sebzim4500 31h ago on HN
To be fair, there is a pretty strong correlation between a problem being in BPP and being efficiently solvable in practice.

There are some exceptions of course (graph isomorphism was solved in practice when the best theoretical algorithms were still exponential) but in general once people find a n^100000 algorithm it soon turns into a n^3 algorithm with reasonable coefficients.

senderista 27h ago on HN
well nobody uses AKS right?
thedreammachine 40h ago on HN
Is it mostly an artifact of the proof or does the algorithm actually need anything close to it?
algorias 37h ago on HN
The runtime looks very weird. The +2 can and should be dropped. This reduces my confidence that the bound is tight. Who knows how the model came up with that expression.
jrflo 45h ago on HN
Glad to seem them changing their tact with the whole NS debacle. Hopefully we can all focus on the results now rather than the surrounding drama.
binlog 41h ago on HN
They didn't change anything lol. The news cycle has just moved on.
curtis-jm 45h ago on HN
You can read the papers here: https://hub.valency.io/collections/openai-math
againstapples 45h ago on HN
As an AI "doomer" can I ask the non-doomer people here how you interpret the significance of results like these, and what kind of progress you expect to see in the next 1-5 years?

Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?

never_giveup 45h ago on HN
Try using AI for your work, whatever you do. You will quickly understand the limitations.
ggreer 44h ago on HN
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.

Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.

psvv 42h ago on HN
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

ggreer 42h ago on HN
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
psvv 42h ago on HN
My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.
ggreer 40h ago on HN
Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

psvv 39h ago on HN
I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.

The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.

ggreer 38h ago on HN
Ok, share links to the conversations with both models. I asked ChatGPT Astra 6 medium effort and it said, "Much less than they used to." and provided stats showing how accurate they are.[1]

1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...

psvv 31h ago on HN
https://share.google/aimode/iKkrZtVYo4DSielWs

https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...

I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?

This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)

So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.

If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.

So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?

In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?

I guess we'll just have to see. I wish you good luck with your wagers.

ggreer 27h ago on HN
You are extrapolating from the mistakes made by free versions of smaller models to claim that frontier models struggle with easy problems. This is an obvious mistake in reasoning because as you can see from my shared Astra conversation, frontier models don't have the same limitation. (They can count letters and they know they can count letters.)

Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.

Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.

The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.

Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.

psvv 27h ago on HN
I didn't save the original query so I asked again this morning and ended up getting the same response -- points for consistency, though it might have been more reassuring with the correct answer.

I don't think my point is really landing so I'll try once more and then give up.

Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.

The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?

I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.

silverlake 29h ago on HN
This is so strange I tried it on Gemini Flash: "Yes, but significantly less than before." When you read beyond the first word it explains where LLMs might fail and why.
raegis 38h ago on HN
Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.
onidj 39h ago on HN
I just asked Opus 5.5 if any AI driven advances in mathematics have been announced in the last day or so and it gave me a summary of this OpenAI announcement. https://claude.ai/share/319b437a-c1f2-4119-8dc5-45d36545fed9
tripledry 36h ago on HN
Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)

> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.

And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...

tmp10423288442 35h ago on HN
The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.

Showing 400 of 1480. Read the rest on Hacker News

Comments are loaded live from Hacker News and are not stored by Mid or Real.