A lot of math is extremely specialized, to the extent that only a handful of other experts in some field have any experience with those mathematical ideas, with most of them not even yet present in the published literature. It's really not a stretch to claim that it's pretty dubious when the AI decides to use these highly specialized tools after it has trained on chat logs where these techniques were being discussed.
Also, if we just take "high-quality" input data, which these chats would certainly be classified as, then the models are more than large enough to memorize everything verbatim. Spitballing some numbers, research literature suggests that LLMs are optimally trained with around 20 training tokens per parameter (fairly confident on this figure), that a DNN parameter encodes around 4 bits of data (less confident here) and I found sources in the 1-4 bits of information per token range (least confident here). So, fairly conservatively I would estimate that a model has the capacity to fully memorize around 5% of its training data, presumably high-quality data is a lot less than that.
If there's one thing I'm absolutely confident in, it's that Sam Altman personally goes to great lengths ensuring that ethical standards are upheld at his company.
Presumably because this was something Levent did in his spare time and because it was not obvious that this work would eventually lead to a breakthrough.
> Why did Tristan use OpenAI's models when it should have been known was a potential outcome?
I'm sure in the past he had less cynical feelings about OpenAI and their penchant for academic fraud.
> I understand they wanted a normal math collaboration but presumably what Levent brought was his resources (as far as I can see Navier-Stokes is not his speciality)
I think you're not giving the guy enough credit in saying that his contribution came down to having an API key for Anthropic models.
> Normally these things are hashed out formally beforehand to avoid the sort of thing now happening.
How would that have helped? That agreement (which may well still exist) would not have involved OpenAI.
So what do you think his contribution was? His preprint record shows no research on fluids - and the statement says that the first LLM-generated proof Tristan received from Levent was 'the most horrendous I have ever read.' Levent is out for mathematical scalps whether it is in his field of expertise or not, and he has the resources to do it. And I am not saying he is not a very clever person, but the idea that you can bring yourself up to the forefront of research in PDEs, in particular NS, and contribute new ideas in less than a year is implausible.
I guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max
There's a spectrum between "barely scientific entertainment" and "dry technical reference". Also, it's not fully a zero-sum tradeoff, great authors have written serious textbooks that are quite entertaining to read, and there are pop-sci books that do a great job at covering advanced material.
I certainly would only trust my textbook as the authoritative source, but i can see that in the absence of an expert teacher it's nice to have something that can critique a proof. Imo the fact that it's hard to verify that your own proof is correct is one of the main barriers in self-studying math, especially if one is at a level where one is not completely fluent in applying the various techniques. This also applies to judging answers to open-ended questions in any other field.
I'm a private math tutor specializing in exactly this sort of material, and I agree with this very strongly. Knowing what "counts" as a proof is one of the most common gaps I see in students who come to me after self-studying, and most students do need some back-and-forth with an expert to really get that skill down. I imagine that LLM's could be very helpful for this if they were used judiciously!
reply