• threeonefour@piefed.ca
    link
    fedilink
    English
    arrow-up
    95
    arrow-down
    1
    ·
    7 hours ago

    This seems to be the mathematical equivalent of using AI to submit 200 unverified pull requests to an open source project and then telling the maintainers it’s their job to figure it all out.

    The one mathematician saying he’s not going to spend hours of his time to verify if a slop report is true, let alone do it for dozens of reports, reminds me of all those projects making rules that unverified AI pull requests will be trashed.

    • 42yeah@eviltoast.org
      link
      fedilink
      English
      arrow-up
      15
      ·
      3 hours ago

      I think OpenAI has retracted some of them already. For OpenAI it’s a really low-risk thing: if it’s wrong, then just retract the paper. If it’s right though, the fame all goes to OpenAI. Meanwhile, people who spent their whole life researching on this topic, needs to confirm this manually for OpenAI.

    • a_non_monotonic_function@lemmy.world
      link
      fedilink
      English
      arrow-up
      10
      ·
      4 hours ago

      I mean, this is a step worse actually. We’ve already seen mathematicians claiming that these systems have actively scooped them. Lots of academics are using these systems regularly.

      At this point, every time I see a paper being “published” by an AI company, I’m wondering who they stole the result from.

    • Venator@lemmy.nz
      link
      fedilink
      English
      arrow-up
      10
      ·
      6 hours ago

      Also makes me wonder if it found and exploited (or got caught out by) some bugs in the Lean programming language…

      (Not saying Lean is buggy, but finding bugs seems more likely to me, as a programmer who knows not much about mathematical proofs since I haven’t looked at anything like that since uni, about a decade ago)

    • nialv7@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      arrow-down
      6
      ·
      edit-2
      5 hours ago

      At least half of these are formally verified. Although they did retract 3 papers. Out of about 700

      • Crozekiel@lemmy.zip
        link
        fedilink
        English
        arrow-up
        14
        arrow-down
        1
        ·
        5 hours ago

        No they aren’t, about 40% of them claim to have been self-verified. But when actual mathematicians looks at them, it isn’t even clear if the “proof” included is proving the thing the paper claims to prove. It all has to be looked at with great scrutiny to find out if any of it even has merit.

        OpenAI just dropped a bunch of busy work on the entire field of mathematics that may or may not turn into anything at all…

        • a_non_monotonic_function@lemmy.world
          link
          fedilink
          English
          arrow-up
          4
          ·
          4 hours ago

          Even worse, the software-based proofs are not the same as the plain text ones that they’re giving out.

          There’s literally no reason to trust them, because it’s completely divorced from the actual text.

        • nialv7@lemmy.world
          link
          fedilink
          English
          arrow-up
          3
          arrow-down
          3
          ·
          5 hours ago

          Saying “self-verified” is massively downplaying it. They are formalized and checked in Lean4 1, which is a programming language used by mathematicians today, to mechanically check their proofs to rule out human error. In other words a theory being stated in Lean generally means it’s more likely to be correct than one stated in mere human language.

          Now, Lean, like any piece of software, has had bugs. A while back someone exploited a Lean bug to “prove” the Collatz conjecture 2. So it is possible that AI agents found a bug and used it. But the vibes I got from mathematicians in the field is that that’s not very likely.

      • gole@lemmy.zip
        link
        fedilink
        English
        arrow-up
        4
        ·
        1 hour ago

        It is sarcastic. But it is also what OpenAI tried.

        The reason they just drop these and say nothing is because these are unverified slop.

        OpenAI’s method of verification is rewrite them as Lean proofs that can be verified by computers. Two things have happened so far:

        1. The Lean proof proved the slop wrong, easy, retract.
        2. The Lean proof, being generated by a LLM, is susceptible to hallucinations. For example the Navier Stokes problem Lean proof turned out to be slightly different than the original natural language proof, because the LLM tried to bend a condition to make the proof compile.

        But, in case nobody found any problem, OpenAI gets to claim credit for the discovery until someone can review and prove something’s wrong.

        You can say these proofs are “Schrodinger’s correct”

    • nialv7@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      4 hours ago
      • L=RL=BPL
      • multiplying two numbers can be faster than nlogn
      • matrix multiplication in n^2.25
      • 3sum is subquadratic
      • Hilbert tenth problem over rationals
      • Unique game conjecture - now theorem

      To name but a few. Each of these alone can be ground breaking.

      Edit: sorry, the 3sum result was not from this batch. it happened around the same time so it got jumbled in my brain, but this one was proved by Claude.

        • nialv7@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          4 hours ago

          my guess is that probably these problems aren’t that well known to the general public? unlike Navier-Stokes, etc.

          if the likes of Riemann Hypothesis, P vs NP, etc. got proven i bet everyone will hear about it immediately.