Asaf Karagila
I don't have much choice...

Why I don't trust Lean code from AI

There are no comments on this post.

So, I wrote a post about why I am not nearly impressed with OpenAI's preprint regarding the Partition Principle. But that is not nearly the whole story. Let me tell you one more thing.

I saw a few people who commented on my previous posts through various social posts and reposts. Many of them were understanding of the plight of mathematicians. Some were trying to make the following point: I claimed that I did not (nor I intend to) read the Lean proof generated by ChatGPT. Some people felt like this is a deal breaker, like this is a reason to disregard everything else I had to offer, because I will not accept this bit of code.

So, let me explain why this is the case. Because I think that if you are not a pure mathematician, you might miss some subtle point in the discource here.

Firstly, Lean is not as sound as you'd think it is, and long code is just codename for trouble. It is known that Lean has soundness issues in its kernel. There are bugs that allow you to peove, in Lean, that \(0=1\). Clearly, this is a false statement. But from a false statement (under the theory that runs Lean) we can prove anything else. When humans write code, we have both accountability and ovresight. Humans produce code at a certain speed, and not any faster than that, so it is not too hard to follow and verify that the code does what it needs to do. I do not have the ability to write a whole new kernel, nor I'd want to do that. I do not think that soundness issues are necessarily problematic when human-generated content is at play, because humans are slow. We produce proofs at a slow rate, which allows us to self-verify.

On the other hand, when ChatGPT produces, in the span of two weeks, several thousands, hundreds of thousands, or even millions, of lines of code... Well, who's going to go through it all to verify that this is all legit? Now we have two problems. The first, which is more famously common, is that the formalisation of a statement is not what the statement actually mean. This is an issue, because if I try to prove that every polynomial over the complex numbers has a root, and the notion of polynomial is not formalised correclty, then the Lean code that I ended up with might not prove what I actually want it to prove. So, if I was writing the code by myself, presumably someone would follow me around and check,as I go along, that I have formalised all the objects correctly, etc. But with ChatGPT, producing a million lines of code, this is not feasible.

The other problem goes much much deeper. AI agents, to the best of my understanding, are very much goal oriented beings. I am sure that you've heard of that agent in Australia that was told to book a gym class, and decided to hack the gym to delete existing bookings of other people and circumvent the five classes a week limitation, just so it can book more classes for its user. The point is that AI is very very good in figuring out exploits and using them. When anyone generates, out of the blue, millions of line of code, we need to be suspicious of this endeavour. Even more so when the entity generating the code is coming from a family of entities that is renowned of finding exploits, bugs, and hacks.

In short, if ChatGPT produced three million lines of code in Lean, I am all but ready to believe that these contain an exploit to obtain the correct result. Even if in the case of the Partition Principle I am willing to accept that the formalisation of the statement itself is easy enough to do right. This trust does not extend to anything beyond that.

Which is what most people outside of mathematics are missing: Lean is not infallible (neither is mathematics, of course). Indeed, Lean is full of fallible problems. And an AI that is bound to find bugs and exploit them (rather than disclose them), is extremely problematic as a certificate. So, that is why I did not intend on reading the Lean code, nor I planned (or plan) to run and test it.

To add to that, it might be worth pointing out that there is another reason (which I hinted at before) for not engaging with OpenAI at this point: mathematical research is done in a very specific way (at least in pure maths), and it is not a particularly fast pace kind of way. OpenAI (and other AI companies) are disruptors. They disrupt the way people do things, and they are trying to insert themselves as "the new way of doing it". Sometimes it works, like how we all ask AI bots for recommendations instead of going on a search engine and trying to figure out how to consolidate a bunch of websites into a small and coherent bit of knowledge.

And this is where the problem lies. It may or may not be the case that mathematical research needs to be disrupted (I don't think so, but I can understand why some people might think otherwise). But the way OpenAI is trying to disrupt mathematical research is entirely based on "everyone outside of maths already think that ChatGPT is an all knowing God, now let us make the maths people bend the knee!". It is evident in the way they release hundreds of "preprints" whose quality is subpar even to some cranks' manuscripts. But, the public perception is that ChatGPT is a god, so many people expect that the mathematical community is going to engage with these "proofs" because they come from ChatGPT. But, as I said before, if OpenAI really wanted to contribute to the progress of mathematics, they would contact experts on each problem, and try to get some feedback and collaboration before going public.

And so, since it is in the interest of OpenAI to write bad proofs, to not vet the Lean code properly, and to lie to us, I am not trusting their produced code. And there are plenty of examples recently to support my doubts. It is just not quite there yet. Maybe in the future, maybe even soon, this won't be a problem anymore. I am excited to see that kind of future, and I am excited to see what kind of use we can make of these AI tools. But I want us to get there together, as a community, not as adversaries. Not while we hate OpenAI, or we are split in how we view and use AI in our research.

So, to all the tech companies. Work with us, not against us. Progress is done through collaboration, not competition.

P.S. I also saw people complain that mathematicians insist that OpenAI produce output that we can read. Well, what is mathematics, as a whole? Is it the production of more and more logically valid statements, or is it the study of understanding and human comprehension of mathematics? I have seen many attempts to define "mathematics", but ultimately the only one that actually fits is "mathematics is a social activity that mathematicians do together". In that sense, if OpenAI claims to prove things, but the claims are unreadable and subpar, then these are not mathematical claims, as they do not participate in our common discourse. There are a dozen cranks who post "preprints that refute Cantor" every single day, let alone the Reimann Hypothesis. There is a reason we do not pay much attention to those. And no, it is not because of arrogance. I am sure that OpenAI does not want to be classified with these people, and for that, we—the mathematical community—need them to play with us, not against us, not disrupt our work. With us. C'mon Sam Altman, if you do this well, you might get a bottle of whisky.


There are no comments on this post.

Want to comment? Send me an email!