1316 points | 46h ago | Discuss on Hacker News | Back to Radar
No clue why I'm downvoted for this, HN struggles with truth-seeking on these topics.
I think it’s really cope to claim this was a human result being stolen.
The facts are much more nuanced than how you're presenting them here.
You mean when they say it was impossible to confirm anything but one day after it was 100% confirmed that there was no theft? And here I'm not even talking about all the ethical problems related to trying to scoop another group when you hear they're close to success, or how current solutions follow extremely closely human-generated ideas, or about the lack of relevant citations in OpenAI's paper.
Believing that OpenAI's claims have any substance cannot be explained by naivety alone.
This was a low-key hilarious replay of George Dantzig and his homework problems: https://en.wikipedia.org/wiki/George_Dantzig
The controversy was whether they had plagiarised that other work on the sub-problem, which they categorically denied after an investigation. And yes, a few days to investigate something like this is reasonable for a company as big as OpenAI. Having seen how data infra is set up when petabytes of data are flowing about, there are thousands of entwined data pipelines to figure out. Not quite as easy as running a query on a sqlite DB!
> Believing that OpenAI's claims have any substance cannot be explained by naivety alone.
Yes, they could be explained by a GitHub repo full of proofs :-)
Or are you suggesting there were hundreds of researchers who just happened to be close to solving hundreds of these long standing open problems using Codex, and OpenAI swooped in plagiarized them all? ;-)
There's no gatekeeping here!
https://proofsandprompts.com/2026/10/07/on-openais-release-o...
We cannot have them rushing to publish amidst tons of confusion, rumors of threats/scooping and outright plagiarism of existing work (by failing to cite said work).
If they're going to participate as scientists in these more rigorous fields, they're going to have to match that level of rigor, not lower it to the disastrous low that ML research publication is at.
With all the hierarchy present in mathematics, I would prefer it by far.
This thing named inappropriately "OpenAI" goal is just grabbing and monopolizing. Capitalists before could not really touch the deep of the human spirit with their filth, now they can.
Look at one of their examples of an initial prompt: https://github.com/openai/math/blob/main/reasoning_traces/re...
No idea what it's so excited about, but it's cute that it "is." I for one welcome having access to a math buddy 24/7 that's way above my level but also always "willing" to talk at where I'm at.
Interesting that its only an excerpt. I wonder what else they include but didn't share.
https://wikimediafoundation.org/news/2026/10/05/openai-rogue...
Not really? We are at a point if an AI today can solve it, it can be stepping stone of understanding something deeper to tomorrows AI and it continues. Sort of like our limitations doesn't matter. Obviously there are many scenarios in this recursive loop but saying it isn't much progress is not how I view this as
Proven math theorems are tautologies.
We are not far away from the moment where these models will be restricted, and sharing the results will be done more carefully.
Who is "we" here exactly?
Right. Before all the AI disruption, pure Math traditionally welcomed anyone who wanted to study its esoteric proofs, right? I remember all the excitement of the average Math enthusiast casually reading Wiles' proof over coffee.
Bottom line is, the relevant people can still understand the generated proofs. The disorienting part is they are a little slower than they'd like, but they'll get there.
Look at that, taxpayer funding was cut and a private sector solution came in just the nick of time, far accelerating the holding patterns we’ve been in for decades
Humanity doesn’t need all iterations towards the blueprints, the blueprint is good enough, we all stand on the shoulders of giants
He spent years formalizing his sphere packing theorem because the proof (human produced) was already beyond the ability of peer reviews. Now his formalization effort likely can be easily reproduced by a model. However one should read his experience about what a formal proof is: often the problem is the statement not the proof. The example he gave is the Jordan curve theorem. It's actually quite challenging to formalize the concept of a planar curve (there are space filling curves). So it is not necessary that someone can look at a formal statement and say aha it is about a planar curve, unlike FLT where there is not much problem in recognizing what the statement is about.
BTW the essay is eminently readable for anyone interested in math. Hales wrote it in favor of formalized math and to educate his peers and students about it.
Which either replaces us in the long term, augments us or makes us better (gentherapy).
While it takes a good math person to make breakthroughs it's much easier to find someone who has a feel of whats important/hard and not. Even a mediocre math major/master is far more authoritative than an expert at adjacent fields (CS,physics). Or to listen to webdevs 'ai skeptics' or whatever on the internet.
There’s about a 1 in 10 chance when someone spells it (or says it) they say Herbert.
Even in situations where they just read it or I just said it.
I’ve had Herbert soccer trophies, health insurance cards, etc.
The mind fills in a lot of blanks and doesnt always get them right.
Yes, you can be an amateur mathematician who manages to avoid such classes and may not know those results or objects, but if you haven't read and written down the name enough to avoid habitually misspelling it, you are outing yourself as a meat proxy unless you are dyslexic.
Mostly we wrote initials in our notes and the exams didn't ask about them. The lecturers only wrote initials on the chalkboard after maybe writing the name once when introducing it the first time. We weren't there to do mathematical history and remember names or something.
Google existed then and now and we could look them up if needed.
Mathematicians aren't exactly known for being well rounded.
Wow, this comment really shows how low this community fell.
I feel gaslighted.
[0] https://en.wikipedia.org/wiki/Unique_games_conjecture [1] https://github.com/openai/math/blob/main/preprints/The-Uniqu...
> A Unique Games instance has a finite vertex set, a finite alphabet K, and a nonempty list of oriented constraints e = (u_e,v_e,π_e), where π_e is a permutation of K. A labeling a satisfies e when a(v_e) = π_e(a(u_e)).
I'm sorry, what? I admit it's been quite a few years since I've thought about the Unique Games Conjecture, and I never dug that deeply, but this part is very, very elementary graph theory and notation. So let's unpack it.
1. e is maybe a name of a list.
2. The elements of that list are tuples, where each tuple is (a vertex, a vertex, a permutation). So e indexes into the list and u_e is the source vertex for the e-th constraint in the list called e. Thanks.
3. a is a labeling. I'm fairly confident that, by "a labeling", they mean that e is a function from vertices to colors, where the colors are the elements of k.
4. That vertex coloring a satisfies the list e, when, for, um, an index e into e, a(v_e) = π_e(a(u_e)). But this isn't for all e, it's for some e, and the goal is to count them.
So maybe e isn't a list? Maybe e is a constraint that is represented as a tuple, so e = (u_e,v_e,π_e) and u, v, and π aren't sequences at all but are, in fact, the trivial unpacking functions that unpack the pieces of the tuple.
Reading this stuff is pointlessly painful, and it's extremely easy to make mistakes when being sloppy like this.
If this were my paper, or if I were trying to train a model to write math, I'd want something like:
A Unique Games instance has a finite vertex set V, a finite edge set E = (V × V), a finite alphabet K of possible vertex colors, and a nonempty list of oriented constraints. Let Π be the set of permutations of V. Each constraint e is a tuple in E × E × Π, where we write u_e ∈ E for the first element, v_e ∈ E for the second element and π_e ∈ Π for the third.
A vertex coloring a : V → K satisfies e when a(v_e) = π_e(a(u_e)).
E ⊆ V × V
In this particular case, though, I think my typoed version may be equivalent. An edge with no constraints has the same effect as no edge at all.
I definitely messed up the constraint definition, though: u_e and v_e refer to vertices, not edges. That’s what I get for writing it with minimal proofreading.
Basically "Here you go, have fun with this, fuck all your demands, by the way we're gonna be releasing the model stay tuned!"
Seems like a lot of PHD students are doing to have to pivot the entire structure of their PHD studies? Or just produce something which is already written by OpenAI?
It has to feel awful to be in this position.
:)
Precisely what all NLP researchers and the ML community at large did in the last few years: embrace the frontier and realize that attention is all you need.
It's much less common for students to have their entire thesis direction removed from under them, as might be the case for someone working in fine-grained complexity assuming 3SUM has no subquadratic algorithms, or working assuming ~UGC. Both of which (publicly) seemed like perfectly valid research directions until a couple hours ago.
There might be stuff to salvage from their conditional results anyway, but this is not your average scooping.
My approach would require custom engineering for every different sequence we'd want to target. With CRISPR, you just "program" the system with a guide sequence, you don't need to do massive engineering to solve a protein design problem.
I guess the only answer is to adapt with the tools. If we can't do that, then yeah, we're in trouble.
I did get my PhD...before AI. And my honest advice (to myself back then, even) would be: quit the PhD, become an electrician, and work hard to buy a tiny house in the middle of nowhere to watch the world burn in this madness.
Very hard question.
Your work makes you one of the very few people who really understands the problem and solution and its significance.
What isn't so normal is the probability and ease by which this kind of thing can happen today versus decades ago when I was in school. As OpenAI said, it only takes a few hours of compute to do what likely was much more than a few hours of human effort. The only reason this kind of scooping/overlapping was rare was mostly a function of how fast other humans could do the same work. With machines, that totally changes the relative pacing between the human trying to learn how to be a researcher and the machine that can grind out results.
I'm less worried about the phenomenon of overlap and scooping and such. I'm more worried about the long-term impact on fields (not just math), especially considering the early stage students and researchers entering the pipeline now. I'm not sure what happens to disciplines when that pipeline stalls.
A much bigger issue is: What is the point of any research mathematician publishing anything now? I really hope that one positive effect of all this will be to finally topple the awful peer review model we currently have, with the biggest publishers gatekeeping with extortionate fees.
Some math PhDs would spend a year or so doing things with AI and lean, and graduate. And keep doing more math afterwards.
Some others, with more stubborn advisors, will keep trying to find a gap where there's no AI progress.
CS subfields go through this every ten or so years.
beyond naive.
1: Author 2: Verifier
/s
Assume the empty list you see is complete. :)
With a deluge of results, having some human expert vouch that it even might be worthwhile would help. (e.g. see the link on HN yesterday, "Two Room-Temperature Antiferromagnetic Semiconductor Candidates" - I see it not worth looking at unless a subject expert vouches for it).
"We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models."
To me, this is a take against progress so that mathematicians can keep their jobs. What would we do if, instead of math, we were talking about diseases? Are we going to keep diseases around so that doctors can keep their jobs too?
> "I believe that AI can contribute positively in all of these directions [NB: exposition, community building, new directions of study]"
It's been a while since I was reminded of this xkcd: https://xkcd.com/435/
> At present, some frontier AI labs are testing advanced mathematical problems on proprietary models that remain inaccessible to the broader scientific community. Our recommendations are formulated with this practical context in mind. However, ideally, they would not do so. We want to state clearly from the start: we do not endorse this practice, and we ask them to stop testing advanced mathematical problems on proprietary models.
To me, the issue is that the models are proprietary which are only accessible to a few people in 2 digits. It's not about progress but access.
These models are too expensive for broad access unfortunately.
This is sort of like discounting putting humans in space because only a few nations have the means to actually do it at the moment.
They won’t because they don’t care and the only way it got this good is something along the lines of they trained on every mathematician’s codex sessions even if they opted out because they consider the thinking traces or output and metadata fair game.
How this maps back to math, idk.
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
That doesn't sound right
If folks are going to analyze this claim with a critical eye, I'd be zeroing in on "average" rather than acting like this measurement is somehow unclear.
Prediction: one of these is wrong and this (publicity stunt) will backfire.
Edit: don't tell me about lean. For lean to function as a proof certificate you need to represent the theorem correctly. Again: good luck doing that across such a broad swath of problems.
The question never if something works 100% of the time but how often it breaks and how that fits the need well. Solving one of these problems is a massive undertaking and accomplishment for the best minds, solving hundreds in a month but being wrong about 10% or something would likely not be the death knell you believe it to be.
> Some of the unformalized results could have issues. We will endeavor to fix any such issues quickly. We are also exploring community-hosted repositories for these materials.
If they are all wrong, that's when it would backfire.
It's fine if not, but it'd be great if even just one of these helped us solve a long-running problem.
From what I can tell, all of the physics results here are quite mathematical. But I am very curious how the internal model they used would perform on more applied problems.
https://github.com/openai/math/blob/main/preprints/Paired-st...
In this case, the tenure is gone, and OpenAI has increased their valuation
US/world deviate more and more from social contracts. Money matters way more, and OpenAI wins here.
If we would have discovered this breakthrough of LLM/ML on scale in a non capitalistic world, we would all work together advancing it faster than it goes right now for the benefit of humanity.
And I don't live forever (at least for now) i def want to see were this road is heading.
Its a conflict of interest for sure, a cnflict of the future of a lot of humans
I am always looking for leftist writing imagining a positive vision for AI. Is there any which you'd recommend?
There's quite a lot of progressive positive writing.
On the economic side the Australian Council of Trade Unions statement is about a positive future: https://www.actu.org.au/speeches-and-opinion/joint-statement...
I’d love to hear more thought experiments if you have some so my failure of imagination or lack of awareness can be overcome.
Technocracy had this already in form of Energy accounting.
Unfortunate something like communism sounds similar and just because we have seen that it didn't work due to technology issues (planning ahead without necessary information is hard?) and no gain of function which would push people, the basic idea is similiar to energy accounting.
Another thing this system needs might be a way for the system to protect itself.
I do think so that a society as diverse as ours will continue struggling with this as long as we do not give abundance resources to everyone or educate/indoctrinate people the 'right' way.
One thought I have regarding AI: IF it happens to slow, people will get used to the status quo and inequality and we will see a future of a handful rich people and a lot more poor people. IF it happens too fast, people might be more desprate to standup and demand something better.
The model is probably comically big and inefficient but big enough
Finally, size really does matter!
Astra wasn't even released 3mo ago. It would not be surprising in the slightest that a public model from 3mo would not be capable of solving this problem, but an internal one from current day would be.
I am now very interested in the explicit calculation of Hamiltonian cycles in the non-bipartite case, and/or the calculation of their absence. If P=NP I think that's going to be a great route of attack.
109. Integer multiplication below n log n
Surprising that this is possible.
158. The Euclidean plane cannot be colored with five colors.
Only 6 and 7 remain!
376. Universal computation in forced Navier–Stokes flows.
Morning coffee proven turing complete
LMAO, I don't think I ever saw such a small number in a CS result.
Very surprising result though! Multiplication is easier than sorting.
Like there is somehow redundancy in a fourier transform that makes it sub Linearithmic?
Which low and behold ->
130. Fourier transforms below n log n.
Does anyone have an intuition to what causes it? What happens at these large scale (or very small)?
It seems n would have to be unimaginably large for this to make any difference. What changes about multiplication / FFT at large enough size ?
I guess nobody expected that it did before this result.
Like matrix multiplication, the common assumption was that it cannot be improved past n^3. Then Strassen broke the barrier (with a more significant constant) which caused intensive research - and now we are around n^2.3..2.4.
That copium didn't last for what, three months?
I think novel mathematics is intrinsically harder than coding (much in the way that computer science research is harder than coding), but the models are scary good at coding.
The highest ranked would be:
| 22 | Hilbert’s tenth problem over ℚ |
| 29 | Unique Games |
| 31 | Anderson-model extended states |
| 37 | Spacetime Penrose inequality |
| 48 | Nonexistence of Landau–Siegel zeros |
| 52 | Baum–Connes |
| 78 | Abundance |
| 80 | Hadwiger |
| 87 | Bose–Einstein condensation |
| 92 | Two-dimensional entanglement area law |
with non-polynomial side being represented as the frontend programmer's constant need for more performance to do the same task...
Non-deterministic can be explained in several ways. One is in terms of a hypothetical "nondeterministic Turing machine" with certain non-physically realizable properties. The easier way is that a NP problem gets as input not only the problem instance x, but a "witness" w, that may depend on the problem instance. This witness generally makes the problem of deciding the problem instance straightforward (e.g. for SAT, x is the SAT instance, and w is a description of how to set the variables so that it is true).
Whomever is running this simulation, please.
See also: noneuclidian geometry and axiom of choice.
It would also be amusing to annihilate nearly six decades of proofs that assume P!=NP.
As long as we also get low order polynomial solutions to important problems, it'll be worth it.
Besides, unencrypted wifi was funny.
I've never been able to find that article as an adult, but I would love to know who wrote it.
Gina Kolata allegedly in the New York Times in 1996 on the Robbins conjecture (noting that computers had started to contribute to math research in some sense), and a longer piece in Math Horizons by her the following year ("Computer Math Proof Shows Reasoning Power"). I didn't immediately find the NYT article, so I don't know if it might be a hallucination.
John Horgan in Scientific American in 1993 (https://www.scientificamerican.com/article/the-death-of-proo...). There's also a retrospective on the topic by the same author in Scientific American in 2022 (https://www.scientificamerican.com/article/should-machines-r...).
Natalie Wolchover in Quanta (but reprinted in Wired) in 2013 (https://wired.com/2013/03/computers-and-math).
I was involved in some distributed computing stuff in the late 1990s and early 2000s and I don't really remember people in that community talking about proofs but there may have been a "if we had a mechanical proof-checker, could we do distributed searches for valid proofs that it would accept?" conversation somewhere at some point. There were definitely volunteer distributed computing projects working on pure math; I remember the Optimal Golomb Ruler search (https://en.wikipedia.org/wiki/Golomb_ruler). So, that could possibly have shaded over into "can we find proofs this way too?". At the time it probably would have been based on brute force searches through proof space rather than clever optimization, though.
The idea that you can lexicographically list all proofs in some formalism and then mechanically determine if any is valid is quite clear from Gödel's construction of the function Bew in "On Formally Undecidable Propositions", but he points out that you don't know where to stop because you don't know how long a valid proof would potentially have to be (so "is this a valid proof of this claim?" can be decided mechanically in a limited time, while "is there any valid proof of this claim?" can't be! maybe the shortest valid proof is 49 steps long but you eventually stopped checking after looking at all 7-step proofs, or something).
The article I'm remembering was not just about mathematics, but indeed all of physics and related fields. I believe it speculated that eventually distributed computing models could essentially take the world's mathematics and physics formulas and various datasets that we believe to be accurate with high degrees of confidence, and then look for patterns or trends, and then from those trends, mathematicians and physicists would be able to investigate further. Not dissimilar to Folding@Home and SETI@Home.
Keep in mind, that this is the best I can remember from 30 years ago, and I've thought about it so frequently that I am certainly misremembering some of the details. Anyways, it's always been this really compelling possibility, and I wish I could find that article that inspired me so long ago and re-read it! :) I really think it was Wired, but it's possible it was Popular Mechanics, or even an expert guest on TechTV who gave an interview. Hard to say for sure, but I've always thought it was a Wired article.
Appreciate your help though!
Edit: with the noun-noun compounding being different from the usual interpretation here, like "scientists who are computers" rather than "scientists who study computation"! Maybe "computerized scientists" or something.
https://unlocked.microsoft.com/ai-anthology/terence-tao/
" I expect, say, 2026-level AI, when used properly, will be a trustworthy co-author in mathematical research, and in many other fields as well.
Then what? That depends not just on the technology, but on how existing human institutions and practices adapt. How will research journals change their publishing and referencing practices when entry-level math papers for AI-guided graduate students can now be generated in less than a day—and with the far better accuracy of future AI tools? How will our approach to graduate education change? Will we actively encourage and train our students to use these tools?
We are largely unprepared to address these questions. There will be shocking demonstrations of AI-assisted achievement and courageous experiments to incorporate them into our professional structures. But there will also be embarrassing mistakes, controversies, painful disruptions, heated debates, and hasty decisions."
He's pretty damn smart that guy.
Only after a world's worth of experts look at these results and then mull over if and how their own fields are impacted by this new info will we be able to answer this question.
I'm reminded of a great TV Show, James Burke's Connections. Where discoveries in one area of science would revolutionize or fundamentally change a completely different area. https://www.youtube.com/watch?v=XetplHcM7aQ&list=PL5HjoPOFFC...
It can take decades to really know the full significance. You know, the whole "We stand on the shoulders of Giants", well the Giants just grew a few inches all at once.
But yeah, this is still a very big deal. Among other things, it will drastically improve all sorts of Rosser-Schoenfeld type results for the PNT and that's just a start. For comparison, I have a paper form 2018 where this result would cut 3 pages out and make the full result cleaner and much tighter, and there are likely hundreds of papers like this.
As for Riemann's memoir, it's hard to compare. You could argue that was "just" noticing a connection (between number theory and Fourier analysis) that nobody had noticed before; in fact this is the kind of thing AI is extremely good at. I'm being a little cute here.
I think if a human had proven just these two results in the form of a uniform zero-free region for L(s,chi) from nothing as OpenAI did it would not be unfair to say that it would be the single greatest advance in math (easily dwarfing Wiles' FLT), and it would instantly put them in the ranks of greatest mathematicians of all time. Unlike something like Navier Stokes there wasn't a semblance of a research program, experts basically considered this hopeless and would have said the chance of seeing a proof in our lifetime was near zero.
For some comparison, Yitang Zhang's bounded gaps result might have gotten him a Fields Medal if he was not disqualified by age. When it was floated that he might have proven Siegel zeros don't exist, it was considered (by experts) clearly a much bigger deal. This result blows that out of the water (it's a way better version); at least analytic number theorists I talked to thought it was plausible but unlikely that Siegel zeros would be eliminated in our lifetime but thought RH was basically hopeless.
Did you mean to not qualify that? That is a bold statement indeed.
I guess the biggest news are not the discoveries themselves but how they were found and that math is going through the biggest revolution as a field since almost ever.
It seems to me less than PNT in terms of what can we actually do with this. Many different areas of math use PNT, and from my standpoint, PNT is helpful not just for what it implies directly but because it lets us make really good heuristics about whether some sets are infinite or not, and what their rough size is. (Granted, one can do that also mostly via Chebyshev). For those purposes, this doesn't really enter in. Similarly, PNT feels like a statement at least I can say explain to my mother without any technical details. This isn't that. But that may also be my own biases of wanting things to cash out to very concrete statements about the integers.
I agree that one striking element is how no one saw this coming. This isn't building on an existing research program, which itself is remarkable. And last night, before I went to bed, I saw a conversation between a bunch of analytic number theorists who seemed to think there was potentially some slack in the quasi-RH argument, which if that's the case means this is going to go even further.
1830 Dirichlet's result is qualitative only, it shows infinitude but not the asymptote in terms of the zeros for it predates Riemann.
To me this is the first substantial step after the 1896 PNT, and we really do not see much progress in the whole 20th century. Personally so far there are only two people worth mentioning,
- Euler, introduces the real zeta function and Euler product, establishes the functional equation at (half?) integers.
- Riemann, introduces complex analysis ideas to the zeta function.
And of course this result if it is true. This is first to penetrate the critical strip, which nobody had any idea how to approach for over a century and a half.
https://github.com/openai/math/tree/main/preprints/The-Quasi...
I thought it was interesting that it said "This paper was written with human assistance", unlike this other Quasi-Riemann Hypothesis preprint that didn't have the same disclaimer.
https://github.com/openai/math/tree/main/preprints/The-Quasi...
At least put a disclaimer for the ad for this site, and maybe disclose how you came up with a total ordering for "top" open problems (vibes)?
> How problems are ranked. LLMs compare pairs of problems. A reliability-weighted model combines those judgments into the ranking, with calibration across model families. The model-family weights are OpenAI 1.00, Claude 1.00, GLM 0.95, and DeepSeek 0.90. These are modeling choices, not measured probabilities of correctness.
+----------------------------------------------------+------+---------+-----------------+
| Category | Full | Partial | Matched / total |
+----------------------------------------------------+------+---------+-----------------+
| Geometry and topology | 25 | 7 | 32 / 74 |
| Algebra, representation and category theory | 17 | 2 | 19 / 53 |
| Analysis and PDE | 11 | 6 | 17 / 40 |
| Number theory and arithmetic geometry | 4 | 13 | 17 / 117 |
| Probability, ergodic theory and dynamics | 11 | 5 | 16 / 37 |
| Combinatorics and discrete geometry | 7 | 2 | 9 / 34 |
| Theoretical computer science | 4 | 4 | 8 / 57 |
| Mathematical physics | 5 | 1 | 6 / 19 |
| Applied and computational mathematics | 2 | 2 | 4 / 8 |
| Quantum information and computation | 2 | 1 | 3 / 17 |
| Cryptography, coding, information and optimization | 1 | 1 | 2 / 26 |
| Logic, foundations and set theory | 1 | 1 | 2 / 18 |
+----------------------------------------------------+------+---------+-----------------+
| Total | 90 | 45 | 135 / 500 (27%) |
+----------------------------------------------------+------+---------+-----------------+There's also one that says that forced Navier-Stokes can implement universal computation (so, is Turing complete). I don't think any of these are resolving open problems per se, but they're interesting for other reasons.
The ones with lean proofs could still be formulated incorrectly
How valuable it would actually be to share the model with other mathematicians vs just have OpenAI's mathematicians churn out and clean up results isn't very clear to me though as they don't say how much effort it's requiring from their team to prompt and clean these up vs how much it's bound by "time to run the model" or similar.
A Polynomial-Time Algorithm for Three-Machine Unit-Job Scheduling [1]
Since some people talk about small numbers that pop up in integer multiplication results, here a completely different number appears:
Theorem 1.1. Let an explicitly listed finite directed acyclic graph specify the precedence constraints on n >= 1 nonpreemptive unit-length jobs on three identical machines. There is a uniform deterministic algorithm that constructs a feasible schedule of minimum makespan. Given also an integer deadline 1 <= T <= n, it decides feasibility exactly and returns a schedule whenever the answer is affirmative. Both tasks can be performed in O((L + 2)^150020) steps on a deterministic multitape Turing machine, where L is the total binary input length.
That is some crazy exponent -- plus an interestingly old computational model to boot; not something that is natural to most of us. I have no capacity to check its correctness today, but I hope it is true purely for the exponent.
[1]: https://github.com/openai/math/blob/main/preprints/A-polynom...
I'm sure all of these super small or large constants will improve over time, but it's still amusing. It is entertaining to see the exponents directly rather than have them hidden as n^c or epsilon or O(1).
[1]: https://github.com/openai/math/blob/main/preprints/Determini...
There are some exceptions of course (graph isomorphism was solved in practice when the best theoretical algorithms were still exponential) but in general once people find a n^100000 algorithm it soon turns into a n^3 algorithm with reasonable coefficients.
Like do you see the technology plateauing at the current level, do you expect progress will continue but only in mathematics, I'm interested to know why others are not concerned?
Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.
It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.
What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.
But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.
The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.
1. https://chatgpt.com/share/6ac5e4cc-02f0-83e8-8f05-99a7ea2bf9...
https://chatgpt.com/share/6ac63b6a-481c-83e9-a8fa-a13ce7402d...
I used whatever the default free model and thinking time was. If progress was really as fast and continually cheaper as some worry it is, wouldn't we expect free models by now to know (or even perform) what frontier models were capable of as much as 2 year ago?
This deep in the "comparing logs" tangent we risk missing the point. It's not what exactly frontier models are capable of at this particular point in time. But that there's entire categories of problems that seem easy to us which LLMs really struggle with. We've stumbled on several just a few replies into casual conversation. (Can they count? Can they know if they can count? Can they reproduce results? How quickly do new capabilities filter into free models? And that's just what's come up naturally, if we wanted to pick adversarial examples there's more to choose from.)
So while there's a number of difficult problems that are easy for LLMs (like bulk generating lean proofs), there are plenty of things where progress is not so impressive.
If LLMs can struggle so much with such easy problems, what hard problems have we yet to discover that they'll struggle with? The fact that no one knows, 5 years in advance, what those problems will be does not mean the chance of them is zero.
So far progress on the things LLMs are good at is fast and easy. It's like fire in a room full of oxygen. But once the low hanging fruit is gone, and the oxygen is out of the room. How fast will the fire burn through steel walls?
In my opinion it's a mistake to look at only rate of progress on one type of problem (whether it be what LLMs are good at OR what they're bad at) and assume progress on all tasks will progress at that rate indefinitely. Isn't there a saying about exponential curves, in nature, all being sigmoids eventually?
I guess we'll just have to see. I wish you good luck with your wagers.
Many people in this thread have made claims about limitations of frontier models, but I'm the only one who has shared a conversation with one. Everyone else is either sharing conversations of smaller models making mistakes, or they're making claims about frontier models but not linking to examples of them falling over. If frontier models were so easily fooled, you'd think someone would link to a conversation showing that.
Why look at the rate of improvement of free models when you can look at token pricing? Back in 2023, GPT-3.5 cost around $20 per million tokens. Astra costs half that.
The worry is not that smaller free models will replace people's jobs. The worry is that future models will. We are talking about the capabilities of frontier models because those put a lower bound on the capabilities of future models. Extrapolating from smaller models is a waste of time, as you can interact with the frontier model to figure out its capabilities and limitations.
Also the timestamps on the shared conversations show that you asked Gemini 10 hours after ChatGPT, which means you asked it after your comment claiming you asked both models.
I don't think my point is really landing so I'll try once more and then give up.
Let's say frontier models today have no problem counting letters, I never really disputed that but only asked about it. It seems based on the other replies in this thread, it's a bit of a "who you ask" kind of thing, but let's grant that they have no issues with it now.
The first version of chatgpt was released 4 years ago next month. Which is not quite 5 years but close. In that time we've just barely managed to get spelling down. If we extrapolate that rate of progress forward 5 more years, are you still afraid for your job?
I think we're all more likely to lose our jobs from a downturn in the economy caused by the capex/debt bubble bursting than being made redundant by AI. (And the continual pricing reductions only seem to make this result more likely.) Hopefully neither happens and in 5 years we'll all still be gainfully employed.
> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.
And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...
Showing 400 of 1480. Read the rest on Hacker News
Comments are loaded live from Hacker News and are not stored by Mid or Real.
senderista 46h ago on HN
sebzim4500 44h ago on HN
senderista 44h ago on HN
tim333 27h ago on HN