• 4 hours

    It’s like firing 5,000 gallons of raw sewage against a wall and then picking out the pieces of undigested sweetcorn to put in your soup.

  • 14 hours

    This seems to be the mathematical equivalent of using AI to submit 200 unverified pull requests to an open source project and then telling the maintainers it’s their job to figure it all out.

    The one mathematician saying he’s not going to spend hours of his time to verify if a slop report is true, let alone do it for dozens of reports, reminds me of all those projects making rules that unverified AI pull requests will be trashed.

    • 10 hours

      I think OpenAI has retracted some of them already. For OpenAI it’s a really low-risk thing: if it’s wrong, then just retract the paper. If it’s right though, the fame all goes to OpenAI. Meanwhile, people who spent their whole life researching on this topic, needs to confirm this manually for OpenAI.

    • I mean, this is a step worse actually. We’ve already seen mathematicians claiming that these systems have actively scooped them. Lots of academics are using these systems regularly.

      At this point, every time I see a paper being “published” by an AI company, I’m wondering who they stole the result from.

    • 13 hours

      Also makes me wonder if it found and exploited (or got caught out by) some bugs in the Lean programming language…

      (Not saying Lean is buggy, but finding bugs seems more likely to me, as a programmer who knows not much about mathematical proofs since I haven’t looked at anything like that since uni, about a decade ago)

    • 12 hours

      At least half of these are formally verified. Although they did retract 3 papers. Out of about 700

      • 12 hours

        No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.

        OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…

        • Even worse, the software-based proofs are not the same as the plain text ones that they’re giving out.

          There’s literally no reason to trust them, because it’s completely divorced from the actual text.

        • 11 hours

          Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.

          Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.

          • It’s interesting that the one person who actually seems to know what they’re talking about is the one getting downvoted. AI is bad at many things for many reasons, but that doesn’t mean we should just assume that anything derived from AI is automatically slop.

  • 13 hours

    Just throw that garbage in the bin, where it belongs.

      • 8 hours

        It is sarcastic. But it is also what OpenAI tried.

        The reason they just drop these and say nothing is because these are unverified slop.

        OpenAI’s method of verification is rewrite them as Lean proofs that can be verified by computers. Two things have happened so far:

        1. The Lean proof proved the slop wrong, easy, retract.
        2. The Lean proof, being generated by a LLM, is susceptible to hallucinations. For example the Navier Stokes problem Lean proof turned out to be slightly different than the original natural language proof, because the LLM tried to bend a condition to make the proof compile.

        But, in case nobody found any problem, OpenAI gets to claim credit for the discovery until someone can review and prove something’s wrong.

        You can say these proofs are “Schrodinger’s correct”

    • 11 hours
      • L=RL=BPL
      • multiplying two numbers can be faster than nlogn
      • matrix multiplication in n^2.25
      • 3sum is subquadratic
      • Hilbert tenth problem over rationals
      • Unique game conjecture - now theorem

      To name but a few. Each of these alone can be ground breaking.

      Edit: sorry, the 3sum result was not from this batch. it happened around the same time so it got jumbled in my brain, but this one was proved by Claude.

        • 11 hours

          my guess is that probably these problems aren’t that well known to the general public? unlike Navier-Stokes, etc.

          if the likes of Riemann Hypothesis, P vs NP, etc. got proven i bet everyone will hear about it immediately.