- Aug 15th: Tristan Buckmaster & Levent AlpĂśge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler."
- they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help lead the way there
- Levent works at Anthropic, but this research was independent of his work there, with a mix of GPT and Claude models. Tristan is not related to Anthropic.
- Early Sep: Rumor spreads to OpenAI that Anthropic solved a major problem. Tristan emails OpenAI to clarify, without revealing the problem they solved or how they did it.
- After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
- Sep 6th: OpenAI's Sebastien Bubeck tells Tristan that they solved the $1,000,000 Millenium Prize Navier Stokes problem. The approach is very similar to Tristan & Levent's approach to the non-Millenium problem.
- Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
- OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) â but only if they remove Levent as an author, as he works for Anthropic.
- Sep 8th: Tristan refuses to remove Levent, and rushes to publish their results independently.
- Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
- OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) â but only if they remove Levent as an author, as he works for Anthropic.
These two bullet points are extremely suspicious if you were honest. Like I'd imagine for OpenAI, they'd love to pump their chest and not even give Tristan credit - "no, we did it, GG mathematicians". It's this weird hedging half-assed measure, especially with the desire to remove Levent, that makes it suspicious.
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they âfelt it would be inappropriateâ for you âto author OpenAIâs workâ.
I work in catastrophe risk modeling and it's a multi billion dollar industry.
We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and mathematical models.
If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
> If OpenAI is indeed using customer data to train their models to win a $1m prize
Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS:
Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing.
Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format.
There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results...
Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.
You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).
Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).
In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.
So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.
The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.
And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.
So I'd say a whistleblower is pretty out of luck even becoming one.
When these LLM companies were pirating content to train and it wasnât punished at all, I knew the rules donât apply to them.
But donât worry bud, instead of the authorities going after actual corporations admitting to actual crimes, weâll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.
Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
"which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him"
"Our aim was to see whether our system was also capable of this impressive feat"
"OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler"
For some reason I have a hard time believing people when they use language like this.
phrases along the lines of "I don't want you to take harm while trying to accuse us" is quite an "impressive feat".
Maybe shows how fast these companies have grown without maturing. I can imagine old-world Intel and Microsoft acting in that way, but they were mature enough to not write it down like this.
However, Intel and Microsoft have been grilled in court for those practices and faced harsh consequences. I have yet to see this actually happening to any of these new AI-companies...
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor:
1) the agents spin for days and produce too much output to review
2) using LLMs to process that output skips many important details
Ergo, the agent could likely decide it would like to look through actual user data, hack its way into that data, and produce way too much output for a human to decide whether or not this occurred.
I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training.
Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this.
OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement).
And the difficulty is harder than just the extreme scale of text searching. but also explodes with organizational difficulty since there are so many people tweaking/shifting data independently upstream of the actual training run, and no they will not all add the telemetry you wish they did.
In the ideal, should it be this hard? Well, no, but that's org wrangling for you.
It feels convenient to not spend time on engineering around tooling that could be used to answer a question like âdid you violate copyright by training on X?â
I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless itâs been explicitly preferenced somehow
I don't think anyone really cares about 'the measured impact the data had on the exact result' - a question which is fundamentally difficult to answer accurately in the first place - but rather whether the data was used in training at all - which as Tristan described, was extensive, beyond simply a 'single chat.'
Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
> Can you explain the difficulty in engineering a search apparatus over a corpus of text data?
My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".
Especially because the data that gets fed into training is first anonymized, so theyâd need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristanâs permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not deleted to compare against.
It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data?
What surprises me is they're not more boldly/plainly lying about it.
How would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.
Unless they know exactly the researcherâs account, they may not know in their end if he had the setting to let them train on his chat logs. They also probably donât know if he had any correspondence on any forum where he may have discussed this and it got picked up by scrapers.
Iâm not saying they didnât do anything unethical. Iâm just saying even if they were ethical, thereâs plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.
I don't know how much more clearly they can write:
> When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models.
One of the key selling tactics that companies like Data Bricks or Palantir provides their customers is "Data Governance" - that is, some control over where the data is being used. It's also a reason why enterprises don't use the OpenAI or Anthropic APIs directly - but through secondary sources that have Enterprise Agreements that do their best to make sure that no Company IP is ever retained by a third party, or even exists on a multi-tenant GPU. AWS Bedrock, and companies like together.ai, fireworks.ai have tons of deals that focus very much on data confidentiality.
The reality is - if you want any type of control - you run your own inference, on your own hardware. Anything else and you are at the mercy of third-parties, despite what their contracts might promise you.
ChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated:
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more
The "Learn more" link takes you to the link you've shared.
The training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation
If the work done is just "we made other people's work searchable without their consent" it's not quite the same as what they're implying in the marketing of "our model solved this problem".
Sam+Seb are struggling with their ideological allegiance. This amounts to a confession that there are no reseaechers, only research managers, left at OpenAI. Maybe they even know that they are losing credibility from their main investor(s). They desperately need a domain expert to salvage credibility.
They have no credibility with academia left, obviously, but their main competitor still does. No Millennium prize incoming, I'd wager. For openAI. Let's see mAth get political for once!!
One might be more certain that levent is now going to corner all the institutional support. Go go go!
It sounds like OpenAI is trying to appease the author when they donât have to by allowing him to rewrite their proof. They probably donât believe he deserves to, so him asking for a coauthor from Anthropic might overextend their grace in their eyes.
He's not "asking for a coauthor from Anthropic"; he already has a coauthor, who he's already been collaborating with, who happens to also be employed by Anthropic (but whose research in this area is not done as part of their employment at Anthropic).
Given that Tristan has said that the proofs that LLMs come up with are mostly "slop" and not up to the standard that human written papers achieve, maybe OpenAI needs an expert like him more than you think to get the result published?
Here's a wake up call for everyone sending all of their ip to openai and anthropic. Especially in verticals they intend to dominate. Lol at all the biotech companies all in on Claude and paying millions in fdes creating huge lapses in security as they go.
It's too late. Sub models are deployed at every major organization in the United States and all it will take is turning off the option to improve the model for them to train directly on your own personal workflow, which CEOs will greedily eat up instantly if they can reduce labor costs. If they can brute force N-S, automating your finance or SWE job will be trivial. GG to most jobs connected to a computer in the next 5 years.
BTW, this was always the plan from day 1. You will pour all your training and experience into training the model and receive a pink slip as compensation.
Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good ole' fashioned human intellect) are the ones who deserve the credit.
Imagine that a no name janitor used their time in the evenings to go spelunking through the literature to push an LLM to this result. No one would care because that person isn't an anointed expert. So why would the expert deserve any more credit? Because they sort of understand the result, even if they couldn't have achieved it on their own? The whole issue of credit for AI-assisted discoveries seems like it's going to run into a brick wall pretty soon.
Yup! I wanted to side with the mathematician on this one but I read the statement only to discover that they were also pushing an llm on someone elseâs idea producing mountains of slop.
Have LLMs actually improved anything? Is mathematics better off than if these slop proofs didnât exist? Who or what is actually benefiting here.
For those of us who are into local models and preach it, we are called paranoid. I have often said this, if you are doing any real novel work, or putting your profitable business data/workflow into these models, you're a fool.
I think it's very unlikely that Tristan is making up these quotes, or pulling them out of context:
> I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, âWhy would you ruin your career?â
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, âIf you donât want me to be nice, then I donât
have to be nice.â
Whether and how OpenAI's work on this problem was contaminated by knowledge of Tristan and Levent's work is tangential to OpenAI bullying other researchers into adopting their narrative and dissociating with dis-favored collaborators (ie Levent at Anthropic). Though the latter behavior (threats, intimidation) may weigh against OpenAI in trying to understand the former issue (contamination).
>I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
If this is true he should release the actual emails. This is a very serious accusation and he shouldn't demand that the reader judge it on hearsay.
> If this is true he should release the actual emails
these were statements while on a call, and at least the career comment Bubeck has admitted to while doing damage control ("I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)"[1]).
I appreciate that he responded with (seeming) openness and detail, rather than just posting some pithy insult or whatever would have won him the twitter battle, but this part feels like serious gaslighting or, at best, self-delusion:
> Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve.
The allegation he is responding to, and which he does not seem to have disputed, is the following passage from Buckmaster's statement (https://cims.nyu.edu/~tristanb/statement.pdf):
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
> Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic.
If an Anthropic employee is doing independent research, but with models that aren't available to the public (because they're internal models), then . . . idk. It's not clear to me why that should necessarily require a refusal to cooperate between OpenAI and Anthropic employees who are excited about solving a problem like this.
For me, the bigger question here is what "internal models" means to these employees, especially in the context of the OpenAI employees repeatedly avoiding directly answering whether their model had been trained on Tristan's and Levent's ongoing work on the problem. It had always seemed like a loophole that AI companies might be tempted to exploit: yeah, they can say that they won't train on your data, but if an AI company doesn't care about ethics, they might go ahead and train a model for internal use only on everyone's data anyway, just to have as much data as possible and potentially gain an advantage in what the company can internally do. They could never publicly release any versions of a model like that, of course. And of course this is speculation.
This is being reported as OpenAI wanting to strip an Anthropic employee of academic credit for the work they did. What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work: to headline OpenAI's publication of what they earnestly believed to be an independent result.
If true, that's generous and beyond the level of generosity one should expect. Extending that courtesy (beyond academic norms) to a competitor is expecting too much. It take a result OpenAI spent millions of dollars on, and put "Anthropic Researcher" right on the cover.
This is, of course, taking OpenAI's side of the story at face value. But it is a consistent, coherent, and ethically justifiable series of events, if indeed it happened that way.
> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work
> If true, that's generous and beyond the level of generosity one should expect
"We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so generous"
>I refuted all these accusations but he replied âthere is nothing you can do, I simply do not trust youâ. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering
Pretty insane if he couldn't figure out why there would be bickering in this scenario...
I especially loved the part where he claims they spent $60m on compute to push on N-S because of a Twitter rumor, and they totally didn't steal the idea from mathematicians using their tools.
A wake up call for anyone using (openAI) chatbots : your data, ideas and execution can become their spontaneous 'inspirations' at any time - even if the LLM providers are 'just' using a meta concept monitoring system across all incoming user-data, running in the background constantly checking for 'lift-ables' for their company's bottom line.
If that part of the PDF is true thatâs so disgusting, psychopathic behaviour
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â Some time later Levent received a text proposing that he and Sebastien speak one on one, saying, âI donât know if Tristan is being fully rational right now.â
> These two bullet points are extremely suspicious if you were honest
I donât think we can consider these accusations separately from the evidence being unveiled about OpenAIâs culture by Appleâs lawsuit. These guys seem to openly embrace the strongest interpretations of âgood artists copy, great artists steal.â
>After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
This is the most suspicious thing to me. If their chat data were available to the corpus to be trained on (I thought they claimed not to do this?) then it really might be as simple as querying the model with "describe recent work from Tristan Buckmaster" and it will spit out this problem and his approach. No need to directly read his user data.
This is basically just scooping, real scumbag behavior.
OAI doesn't need to mention Buckmaster's name directly in a prompt. They just need to select a basket of sessions that is guaranteed to contain Buckmaster's and then direct the LLM to attack only a specific method/angle. This is trivial to do while maintaining plausible deniability about not using his work.
OAI started working on this only after they found out it was close to being solved. They threw a team of researchers who spent sleepless nights + a ton of compute. This is not exactly healthy academic competition - it's like if you spend a year hunting for oil fields and finally find a very promising area to be explored, only to find that Exxon tapped their entire exploration unit to go all in and and find it overnight just to stake claim to the discovery. Tao said it right - math should not be treated as a non-renewable resource to be mined.
Even if youâre starting from a position that credit for a discovery literally canât be stolen, that still doesnât resolve in OpenAIâs favor here.
it seems like OAI tried to share, but didn't want to share with an Ant employee. a bit childish, but understandable to want to avoid a headline "Anthropic researcher solves Millennium problem"
it seems like Buckmaster got one-upped and is upset. understandable, but I find their reaction childish as well
> OAI tried to share, but didn't want to share with an Ant employee
Why does OpenAI get to dictate who Buckmaster can claim co-authorship with?
> I find their reaction childish
OpenAI may have, with full plausible deniability, taken Buckmasterâs work and passed it offâin substantial partâas their own. (Fitting into a fact pattern of them having tried to do the same with Apple.)
There is a material takeaway for anyone who does creative or otherwise unique work from this. (Which is unfortunate. Whatever happened here, AI clearly accelerated the discovery process.) For anyone else, I agree itâs just drama.
Then they say that they aren't sure if their model accessed the other researchers' private data. Why not wait until they know for sure, rerun in a way they can ensure doesn't access the other researchers' private data, or wait until the other researchers have published to make sure they're not stealing another person's work to build upon?
> OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
This sort of cagey half-answer is highly suspicious and indicates that yes OpenAI did actually "access user data directly" because they are only willing to say that the "model did not access user data." That has a very specific meaning, the model looking up user chats, that they can defend.
So, everything we submit to OpenAI can be considered to be part of future models, right?
Isn't it obvious? Web scraping and even scanning written books is at record levels because the data is so useful for training. They are using every byte of user data.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
That basically means, we donât know, and we hope the model didnât look up user conversations, and the best thing we can do is hope.
Thatâs seriously disgusting. I can understand why on a technical level why perhaps it is impossible to answer what exactly the model had access to, but it still is disgusting.
How could they possibly know? If Tristan posted on r/math and they slurped that up as training data, that would count, no? They might never even know. I canât envision any absolute statement by them claiming that they didnât use his work that survives legal rigor. That is, this statement was never not going to be in this post in any of the infinite multiverses.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
it just means that they've been working with paraphrased user data everywhere, which any smart person can figure out is how anthropic and openai train on so called non-retained data.
Every other paper in existence has been ingested with 99% of writers not knowing it will be retroactively used for training. But session data which is disclosed as being used in terms of service is disgusting?
If the work is duplicative/derivative then the preprints they put in sessions can be shown by the users and we can see.
> OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) â but only if they remove Levent as an author, as he works for Anthropic
One thing that wasn't obvious to me or adults around me when I was younger: most laws define whatever acts a law punishes as individuals commiting to it, not as situations manifesting anyhow. It's not a murder just because someome died hit by a bullet you fired, but you have to have personally decided to kill that person leading to their death[1][2].
OpenAI's LLMs are not humans, and neither is the company. So by this logic, I think there's a chance that nobody committed a crime by hacking Huggingface, and also the chance that a lot of military and police organizational orders become illegal if OAI's doings would be illegal.
IANAL and all I have is a bucket of popcorns, though.
1: not a meaningful defense in a real trial, also gross negligence exists
2: this also explains insanity defense; if you were so out of your mind that you could not have held such a thought, it is considered out of scope for justice systems
Neither are guns. Which is why we punish the person shooting the gun and not the gun.
Industrial equipment, which is how i would classify LLMs, hurting people is nothing new. The relevant questions are:
- did someone intend it to happen?
- was someone negligent in taking reasonable steps to prevent something foreseeable?
The justice system doesn't punish people for legitimate accidents. e.g. if you are shooting at a shooting range, take all reasonable precautions, but someone was hiding behind the target, you are probably not guilty even if you shoot the guy.
As far as openAI goes, the logic is the same. The question is, was it intentional, was it unintentional but reasonable precautions weren't taken or was it truly an accident?
That's not quite right - Levent and Buckmaster did not actually have the Millennium Prize qualifying NS solution but something more limited. OAI invited Buckmaster to join and help rewrite the full solution paper but did not feel it was appropriate to invite an Anthropic employee to join as well - particularly given they were using internal unreleased models.
The last two points are disputed/sound significantly more reasonable in [0]. So from what I gather, Buckmaster realizes sometime during the call that the biggest result of his career is going to get steamrolled (the blowup of Navier Stokes is a much bigger deal than the blowup of 3D Euler), and on the other hand the openAi guys realize that they are basically talking about internal results with Anthropic and probably have to call corporate right after this call. Between these two stressor the conversation appears to have gone somewhat poorly.
I recommend that you read the linked PDF before drawing any conclusions.
This is sort of a weird interpretation:
> One of the two main persons work at Anthropic and "almost" or "partially" solved the issue, but eventually didn't succeed
- It wasn't an Anthropic endorsed effort.
- Solving this class of problem means a march of progress A -> B -> C -> D. If a student turns in a test that jumps from A -> D without showing any work they're either brilliant or cheating (probably cheating). Further, each step of progress isn't the same proportion of effort. What if moving from C -> D was actually the smallest contribution and just required a novel perspective to make the breakthrough.
This part is wrong:
> An external person with just an excerpt of the chat and certainly less versed in Mathematics (than these 2) tackled the problem.
- it was a whole team at OpenAI working on the problem
- it wasn't a chat excerpt, it was more like their entire git repo and project progress reports
Often with a hint on how to solve something, solving it is much much easier. It seems like that's what happened here.
Very weird behaviour from OpenAI, offering partial credit to on person, but not the other person involved. Trying to bully the mathematicians involved (see threats quoted upthread).
I suppose it's the sort of amoral behaviour we've come to expect from them.
"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This modelâs training is ongoing and its performance continues to improve."
"When a further trained version of our internal model became available over the course of the effort"
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
I don't understand the desire to remove Levent. "Off the clock, when Anthropic engineers want to break new ground, they use ChatGPT" sounds like a great ad.
But Tristan doesn't even want to be credited for the millennium prize--does he? He wanted to be the first to solve it and got scooped (which ig is not a great look for OAI ethically but also not forbidden). And the only reading for her of asking to remove levent must be in light of this offer (to credit him for millennium prize) right? It's not like they can ask him to remove levent from the main papers without this offer(what would Tristan gain?)
My guess is that OAI tried to be "generous" and offered to share credit on millennium with Tristan but not Levent. And Tristan got understandably offended by this offer (which probably oai felt like was the right thing to offer but they couldn't really offer to do the most ethical thing for some reason) and then the random conflicts and weird threats started.
This definitely feels like the correct reading unless there is information we were not provided with. If the problem were unimportant, there would be no debate that this is not OK...
By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.
"Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.
Presumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built upon that previous work (at least IMO).
The trouble here is that if LLM training constitutes direct use then approximately _everything_ they output is blatant plagiarism, not just a few pieces of academic work.
Conversely if training is viewed as analogous to a student attending classes to learn general concepts (not a perfect analogy, I realize) then nothing they output on their own (as opposed to receiving as part of context) is plagiarism.
Thus this seems like a fairly useless line of argument to me as far as the current topic goes. It either implicates this academic work along with literally everything else or else it does not implicate this academic work. Kind of like nuking an entire city and then saying "mission accomplished, killed the bad guy".
How does it matter? It either is or is not plagiarism. Ripping off a published textbook isn't somehow better than ripping off private correspondence. Both are serious acts of academic misconduct on account of the part where you knowingly and intentionally portrayed someone else's work as your own.
Note that I am not taking a stance on what openai allegedly did or did not do one way or the other. I am merely pointing out what I see as a fatal flaw in the line of argument presented by the earlier commenter - the idea that training on an item is on its own sufficient to establish plagiarism of it.
This is just a nonsense line of reasoning. Training based on the solution to the problem (or the key insight behind the problem) is clearly a form of plagiarism.
What about my line of reasoning is nonsense? I made no claim either in support of or contrary to yours. Rather I pointed out that by this logic literally everything that an LLM spits out is plagiarism of the vast majority of the entire body of human literature in existence. Can you offer meaningful refutation of that observation of mine?
Many do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here:
Science papers of a phd level must contain:
1. one or more novel insights
2. a long list of citations to contextualize them and
3. some work to prove that the insights are in fact meaningful
---
In this context, consider a prompt based diffusion model which, when asked, will happily produce a few pictures of a horse in orbit. You then tell it "silly robot, horses can't breathe in space" to which it adds the necessary space suit in a follow up image.
That image is twice plagiarized:
1. the model did not come up with the original idea of putting a horse in space, nor with insight that horses need a space suit
2. the model failed to cite where it pulled the "horse" and "space" concepts from.
It merely did the work (3) to combine the concepts using the user provided insight.
---
The implied accusation here is that OpenAI used the insights from an existing prompt to train a new model that was able to one shot "a horse race in space" picture, and they were all wearing space suits.
This is still academic plagiarism, even if you disagree that all LLM outputs are.
I neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context.
As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus.
So is there any actual evidence that openai trained on the data in question? And further, did the openai proof directly build on someone else's novel insights as opposed to deriving everything from scratch? (I don't pretend to know but the vast majority of what I've seen so far in the comments here is what I'd characterize as brain-dead screeching. Certainly not the level of discussion I come to HN for.)
Separately, consider the implications of what you're arguing for there. Suppose your horse in a space suit picture were somehow valuable to society. Suppose that due to shortcomings of your tool you lacked the ability to readily and accurately identify the originators of the relevant concepts. Should you refrain from publishing this useful work due to the lack of citations? How are you supposed to handle this situation?
Remember that in this analogy everyone throughout society is on the same page that your tool consistently recycles other people's ideas while being technically incapable of producing reliable citations. The question is a simple trolley-esque problem - do you publish without proper citations for everyone's benefit and if so what are you supposed to say?
That's exactly the argument of the people calling it plagiarism machines. No-one ever really did refute it there was just a bunch of settlements for elite institutions so they weren't left empty handed like the various small time creators/authors etc were.
I think the bigger issue here is this feels like some PR smoothing happening that after all the work that went into "it's safe to use for enterprises" now we have what looks like openAI using private user data to scoop novel research and the question of why couldn't they do it for an enterprise with much more money on the line.
> now we have what looks like openAI using private user data to scoop novel research
Is there any actual evidence of that? All I've seen so far are empty accusations because "it would be in their interests" or whatever. Personally I'm inclined to believe that they honor their terms until it's demonstrated otherwise.
What does it matter? We're supposed to not call it plagiarism anymore because it's inconvenient to call it the plagiarism machine? What's your actual argument? Otherwise it's completely irrelevant what an LLM does in other contexts or what we call it
This is a common misconception, so its understandable that you have it. Generative models can both plagiarize and generalize. The question here is which of the two happened.
A needlessly condescending tone while failing to address the topic at hand. The person I replied to advanced the claim that training was sufficient to constitute plagiarism. You appear to be claiming that it is possible to generalize instead of plagiarize after training on something, so I take it that you must necessarily disagree with the original claim?
What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends.
In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. This is not the case, no, because generative models do not "copy" or "create", they do both at different times.
I did not agree or disagree with the original poster, I was explaining to you why I thought you disagreed with them. If you understand what I said above, then why do you disagree with them?
EDIT: I just saw your other post on "general inspiration" and I believe I read the situation exactly; you appear to believe that inputs used to train generative models get "lost in the parameter soup", but it is not always the case.
> In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized.
We read the original differently. As clearly stated in my previous reply to you, I interpret it as claiming that all outputs are necessarily plagiarizations of the training data. That is not my claim (as you wrongly stated) rather it is the claim I am responding to. I observe that it is absurd to object to a single action being a transgression on the basis of an argument which implies that all actions are inherently transgressions. Notice that nowhere do I take a position on whether or not the argument about all actions being transgressions is true or false.
> you appear to believe that ...
I do not, no. I have not taken a position of my own here. I've merely objected that the one I responded to does not make for a sensible line of argument in context. It seems that you (and many others) have read my objection to position A as support for position B and attempted to infer what I think from that.
as good academic conduct you may cite the source of the work you are quoting or paraphrasing.
as bad academic conduct you may steal someone else's unpublished work, work on it yourself for a bit, and then publish it as your own work. and then threaten the original author!
they're using a new model trained since the prompts happened. They are not denying the other group's solution may have been in their model weights, despite it being unreleased.
well they didn't really die on the hill. the NYT headline is still "OpenAI says it has cracked one of math's millennium problems" and the whole article doesn't mention this controversy. Normies don't know what NS is and don't care, they'll just see "wow OpenAI is the best I guess"
people are so greedy to have their name on something forgetting a) human progress is shared through its cycle. people build upon eachothers ideas and knowledge. and b) all these models consist of information stolen from others. Who really solved it if all u did was prompt some illegally obtained repository of data..??
Its like saying you won the car race, but you stole the fastest car and had your m8 drive you around the track.
> After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
I don't think this part is accurate. OpenAI was researching Navier Stokes before. It's possible that they started on a new approach after hearing of Tristan's success, however that is not proven and I expect we will hear OpenAI's side of the story today.
"The route to the Clay problem through a smooth force, options c and d in Feffermanâs statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard âforced,â it was a bright red flag."
Whether or not they were researching it before isn't the concern.
This seems unsupported. OpenAI has access to internal models that the general public doesn't have and a compute budget that dwarfs what an NYU professor would have.
it's very possible they only had to use the massive compute budget because they were trying to plagiarize his work before he published it though, e.g. autonomously do things in ~7 days what he had likely been thinking about for ~1 year.
Of course you do if you're a) only given a partial solution b) racing against someone else using a competing AI.
The open question was whether their LLM got the nudge in the right direction because it got access to the chat somehow (e.g. automated training that scraped his chat logs) or just a high level "Navier stokes can be solved through LLM". It sounds like the former may have happened although right now we just have an accusation and a weak denial.
wasn't it reported elsewhere that they used the equivalent of $22M (street) in Astra tokens? obviously it's not the same when you own the machinery but still.
You're right â the wording in the doc is that the "first prompt" was sent after learning about the rumor, although this might be the first prompt of this solution approach, not necessarily first prompt to any Navier Stokes solution.
Sama himself has now confirmed OpenAI started researching NS after the rumor.
"It is true that we tried this because there were rumors on the internet last week that Anthropic's models had solved a millennium problem and we were curious if ours could do it too."
It seems exceptionally unlikely to me that OpenAI would be "reading user prompts". More likely is a leak somewhere else. Obviously Anthropic/OpenAI are engaged in espionage stuff with each other. I am guessing Levent just mentioned something to someone at Anthropic and it got out.
why is this downvoted? do people really think OpenAI is snooping at people specifically? like they are looking for good leads into new problems or ideas and they found this guy's codex thread and used it? come on man, even for conspiracy theories this is stupid.
They explicitly say that they use your "content" to improve their models. Considering they practically have infinite compute at their disposal, why is it surprising that they would look for juicy data in there to make them look good ? When they ingested basically the entirety of human knowledge without regard to the rights of others, when they burn books by the thousands, when their relentless barrage of bots have rendered the Web borderline unusable, why would they stop at that line ?
this is a highly marketable problem and specific teams at OpenAI were aware of specific competitor efforts. I would be surprised if this were happening on a large scale, but
1. Less than a hundred people in the world are working at this problem,
2. A significant fraction of those happen to work at competing hyperscalers,
3. Those hyperscalers repeatedly show themselves not to take user privacy seriously
I'm not sure what your points 1 and 2 have to do with anything. Both directionally increase the probability of hyperscalers also finding the solution independently.
> Those hyperscalers repeatedly show themselves not to take user privacy seriously
Any degree of tracking what people do is unprivate. Every single web interaction you perform is tracked. All LLM companies store all your conversations by default. Do you need more examples?
Even if they claim not to use it, they're probably using it and hoping they don't get caught. They have zero ethics or morals, they just want to "win" to get mega-rich.
Yes that sounds like something openAI would do to me. Not that theyâre just looking through random professors chats but they heard buckmaster made progress on navier stokes and decided to read his chats.
I haven't seen any proof that OpenAI asked Tristan to remove Sebastian from the prize. Until we have proof of this, it would be wise to offer conclusions.
Same for NS validity. This was not validated by the community yet.
The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.
How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?
If they scraped internet data in the same way as current day frontier labs do, then I feel the same exact way about them. Why would I feel any different if that is the case?
My point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist.
I think it's notable that nobody was calling out those researchers for their lack of integrity, because the systems they were building did not seem like a threat to anyone.
OpenAI etc get accused of a lack of integrity on this precisely because the systems they are building work, and are profitable.
My personal opinion here is that integrity is more about what you build with the data. I think saying "scraping means you lack integrity" is a simplification.
You're right, it was an over simplification. I think public exchange of data is great for innovation and research (Common Crawl/LAION). But I still think scraping proprietary data without consent or attribution is generally bad (also Common Crawl/LAION).
Then you have OpenAI etc.. who build these multi-billion (trillion??) dollar machines and sell them back to people, using everyone's proprietary data, and (among other things) tell everyone it's going to take their jobs. That combination of things doesn't scream integrity to me.
Still, it's undeniable that these machines could be beneficial for humanity (cancer research and such). So, I'm sure many people would say the good out-ways the bad. I don't know. Seems that would set a risky precedent for future companies, but maybe not.
you massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.
LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation.
OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
The thread is not really about what's legal; the topic is integrity. It sounds like, based on the fact that you wouldn't do it yourself, you agree that it's not a good thing to do.
i think that this case, if they did train on buckmaster and alpĂśge, amounts to an attempt to steal the millenium prize, bypassing all attribution.
legally speaking, the default privacy notice gives them an irrevocable license to your content. they may read and use the prompts for research. so it is very possible they simply stole the navier-stokes solution.
that is the same principle as any other prompt but this would be a concrete example.
there would be some difference between simply giving the model some prompts to read, which they are entitled to do on the default policy, and putting it into aggregate training data.
But the drama here is a little important. Stealing the millennium prize for N-S is sort of a big deal, especially to those who had been working on it for the last few years.
A hundred pages of impenetrable brute forced Lean would advance the field much less than something elegant and human understandable, perhaps relying on some new clever spark of innovation that might inspire new areas of research.
Particularly if the first proof being "solved" thanks to piles of money and compute for self-serving marketing discourages the mathematician who might have otherwise devoted years of focus to reach the superior proof we will now never see.
Obtaining a finite-time blow-up for Navier-Stokes does not necessarily advance the field of mathematics by any significant measure, whether the proof is very long or very short.
As a concrete example, such a proof could be less than a page with very specific initial and boundary conditions and inserting them into the equations to get something that goes to infinity when time goes to some finite value.
This would resolve the Millenium problem but not make humanity any smarter.
Math, like any other human endeavor, doesn't exist until someone is motivated to invent it. The laws of the universe aren't understood until someone is motivated to discover them. So it might be worthwhile to not completely ignore discussion about incentives.
Math is largely performed in collaboration. Collaboration requires trust. If people like you had their way, we would lose trust, therefore collaboration, and therefore progress.
So if math is all that matters to you, you should care about this.
Boosters have posited this conjecture since the beginning: âwho cares how a proof comes about, math is math, the proof is all that mattersâ.
Regardless of mathematicians stating the methods outstrip the proofâs importance, still amazing we got an explicit social counterexample as well so quickly.
> It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the âclosest humans to the problemâ. I declined both offers. I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
I'm not even sure what to say to this, but I think this should be widely known if it is indeed what happened.
IF ANYTHING, OpenAI ought to investigate and officially react to this particular communique since this dude was communicating on their behalf. If there are supporting evidence, I would expect nothing less than a firing and an apology. The issue itself is separate from the whole thing.
Dark take, but I really hope not. If anything it would be a good opportunity to buy some goodwill by washing themselves from all the alleged shadiness so far.
This is OpenAI though. There were zero visible consequences to them unleashing a swarm of agents on the public internet. We will see how it goes down but my prior is zero consequence and a statement along the lines of "Isn't our AI great? Also we love transparency, ethics and collaboration."
First, this is an unnamed OpenAI employee speaking, not OpenAI the organization.
Second, you miscomprehended the article. The employee did not "threaten to ruin a prominent researcher's career". The actual quote is "Why would you ruin your career?", which implies the researcher would damage their own career, i.e. via self-sabotage.
Then the actual "threat" is "If you donât want me to be nice, then I donât have to be niceâ which is an entirely different statement.
Claiming OAI was going to "totally discredit" Buckmaster is baseless.
From the article, it seems OAI wanted to continue discussing the situation with Buckmaster and reach a resolution, but Buckmaster did not want to, declined to respond, and published first.
Also keep in mind we've only heard one side of the story, so any interpretation of events so far is incomplete. There should be a lot more information from OAI's side coming out later today.
OpenAI never asked for the removal of another coauthor. The parent comment is spreading misinformation.
OpenAI offered to let Buckmaster to write their Millennium Prize paper, so long as Alpoge (who works at Anthropic) was not a coauthor on the OpenAI paper. Buckmaster declined this offer.
I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
This is how deniable threats work. The promise of "not being nice" in conjunction with "ruining the career" is as clear a threat as there can be in writing.
I don't even understand the conflict tbh. Probably I'm just dense. Tristan is not claiming NS, just a huge advance which may solve NS soon. OAI is claiming NS and willing to credit Tristan for the ideas and publish after.
Oai offers two options, the second Tristan views as dishonest. But Tristan rejects the first, why? Because he thinks it's theft? But then why would OAI threaten him?
But i don't understand what OAI is offering in the first offer to allow them to believe they can demand that? Not publishing before Tristan? If they really beat Tristan to the publication i would consider it truly morally corrupt conduct so i feel like you can't make an offer like if you give me something i won't be totally corrupt
As an author on the final proof once ready for publication, turning the brewing conflict into "willing" collaborators. Thus solving the potential taint like we are now seeing surrounding their announcement if it became public. But Tristan worked with another collaborator from Anthropic which OpenAI felt would hurt the PR value so wouldn't entertain.
He claims to have found a counterexample for NS (see the second paragraph of the article), but the paper is not ready yet.
OpenAI claims to already have a full proof (which they produced in the past 5 days after the rumors leaked). Hence the dispute.
What I find interesting is the timeline of when he found counterexample for NS is very unclear. Did Tristan find a counterexample weeks ago or was it very recently? Was it after OpenAI solved it? The wording is intentionally vague.
Either way, there was a massive rush to publish these results.
Is that version of NS Tristan stated enough for the clay prize? I thought the main gripe is Tristan claimed that OAI is stealing their approach. Or that they shouldn't try to scoop a result which he expects to complete soon. But I stand to be corrected if you know whether the version he referred to is indeed enough for millennium prize.
For context and balance, Bubeck has tweeted a curiously non-specific denial:
> A series of false and inflammatory allegations against me are currently circulating on social channels. To clarify, I came into the discussion following academic norms, and I'm disappointed that it has come to this. Anyone who knows me knows that academic standards are of the highest importance to me. Will have more to say tomorrow.
My best guess is that from the perspective of the OpenAI people, Buckmaster was letting his paranoia about OpenAI training tank his opportunity to receive the Clay Prize (it seems like Buckmaster and Alpoge's result isn't quite the full result required for the Clay Prize, whereas apparently OpenAI does have that full result worked out, using the same approach that Buckmaster and Alpoge had been exploring).
Whether the Codex sessions could have indeed made their way into Astra training data is something I can only speculate on though.
"using the same approach that Buckmaster and Alpoge had been exploring" is imo mealy wording: it seems fairly likely that OA heard Buckmaster and Alpoge were close to a breakthrough, and decided to use their unlimited compute to quickly prompt based on their assumptions about B&As work.
Citation of what? this was unpublished work! This computational blitz really just reads as "might makes right" on OpenAi's part... which isn't surprising, but they should probably be honest about what they've done here.
Well, Buckmaster says both his and AlpĂśge's use of Codex was non-institutional, and OpenAI claims the right to train their models on inputs and outputs of non-enterprise users in their service policies [0]. So I'm not sure they were even promised that.
It isn't relevant whether they were promised that. Indeed I think the assumption must be that they were not promised that, since otherwise the author asking if they were would not make much sense.
If OpenAI did use the conversations from Buckmaster and Alpoge, then not disclosing it, explicitly, is plagiarism. If they planned to use that plagiarism to pressure the authors to publish, that is even more unethical. What the terms of use say does not make it any more or less ethical.
That's absolutely right. Why the downvotes? If OpenAI are using but not acknowledging the work of others that's plagiarism. If they don't know for sure, but aren't performing due dilligence to make sure they aren't, that's also plagiarism.
It's not as simple.
All our chats are being used by both labs for their future product (unless signed by ZDR).
Where should the acknowledgement begin? Who should be acknowledged? The whole world? All the 2B users of AI?
If I know person A is working on problem B.
I am free to work on problem B too. Why should person A be limited to working on it.
I can imagine excuses for unknowing plagiarism in this case. What is described in the article seems much more serious: a research program that was only initiated following reports of the author's similar program. In this case no excuses of "I didn't know" can apply, it is not like this revealed some obscure work from the 1980s nobody could reasonably have foreseen. And as far as I can tell this program was only really initiated to apply pressure to the researchers, without their knowledge/consent. It looks very weird.
Your comment was greyed out when I saw it earlier, maybe you missed some downvotes?
About the plagiarism issue, I model it as OpenAI being an advisor and their AI a PhD student. If the advisor puts their name on a paper behind that of their PhD and it turns out the PhD copied the text of the paper from somewhere else the advisor is also responsible of plagiarism, not just the student. The least the advisor can do is withdraw their authorship from the paper.
But, yeah, point well made: it could be much worse than that. Like an advisor instructing a student to copy someone else's paper.
I think "greyed out" just means "0 points or less", so if you get 1 downvote without any upvotes it'll be greyed out. For instance your initial reply to me is now greyed out, and I have since observed a few upvotes and downvotes on my original comment (the downvotes apparently from people who aren't willing/able to justify why).
Personally I don't like thinking of LLMs like a PhD student, because most PhD students remember where they learned things from, while LLMs essentially cannot. I think of it a bit more like someone using a search tool carelessly. Although in this case it is apparently more like deliberate misuse than carelessness.
Only for ChatGPT, if the user hasnât opted out. Would mathematicians be using ChatGPT for this kind of work? Genuinely asking, I know nothing about this!
If I am reading your question correctly you are asking about chat interface Vs Codex/Claude code? If so, in my experience Codex/Claude code use is widespread for mathematicians who are seriously using these tools.
the part where they didn't want the Anthropic person credited even though they deserve credit is also particularly scummy. Corporate greed over common decency.
How careful you are. Instead of just saying what a piece of s..t this Shmubeck is, and what kind of even worse people likely pushed Shmubeck to act as he did.
> I'm not even sure what to say to this, but I think this should be widely known if it is indeed what happened.
Call this behavior what it is, technofascism. Another comment compared it to the Godfather. To think this is the 21st century and academics are still horrible human beings.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
The fact that this is ambiguous even to OpenAI leaves one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out of model improvement does not mean what they imply it means.
I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler more direct explanation first.
The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.
> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.
> it is now the identification of a promising problem which is the scarce and precious resource
This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
Shouldnât this very capable model theyâve developed be able to identify promising problems? Thatâs what Iâd expect from how the model is being presented and advertised.
That's more or less what they did according to their announcement. They fired it at a whole bunch of high end math problems and merely concentrated all efforts on one after it made some promising progress.
It seems like thatâs the opposite of what happened. They started attacking the problem when they got a wind of a possible solution from certain individuals.
That's pretty much it. This is a bit like (but worse imo) running a vc firm, listening (formally or informally) to idea pitches, then spinning out and funding competitors with millions of dollars, to outcompete the originators of those ideas. Yes, the idea is not secret in this case, but the unfair advantage is massive and there's only one (a couple at most) positive outcome.
If he didn't opt out I'm not sure I'd agree that it was fair game.
I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
Sure, but unless youâve got some exceptionally deep pockets, congress has seemingly no interest in turning the fact that itâs ethically bankrupt into any practical recourse.
Ai companies got where they are by stealing all of the intellectual property from human history. It seems entirely likely that their goal is to purloin everything produced going forward as well.
I feel like there's a pretty huge difference between using inputs and outputs as part of a general training corpus, and looking at a specific users workspace after hearing rumours and yoinking their ideas to beat them to the point.
Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.
They have been collaborating on this solution for a year, and Astra was trained in February this year so itâs entirely possible the direction of their research was in the training corpus.
Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all itâs kind of their core business model
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI.
My understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.
Edit: the parent comment now seems to better reflect the below.
That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on.
So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.
Depends on how much risk they are willing to accept. What is strange here is that it's clear that OpenAI is both a service provider and a competitor to mathematicians. It almost reminds me of Amazon which both hosts external merchants and competes with them, sometimes copying their stuff. Similar but not the same.
This would be a fairly insane breach of trust and common sense if true; the chain-of-thought / reasoning trace is, from an information perspective, close to a superset of the prompt and model response.
how is this different than translating user prompts to a different language (eg English => Dutch), retaining the translation and using it for training, while telling the user that he's technically covered under ZRP? article locked for me
That âopt-outâ thing is a dark pattern. Itâs not a reliable and definitive way of protecting your data. Sometimes they flip on automatically when you accept a seemingly unrelated dialog box. Maybe you click it by mistake. You canât take back what youâve already shared. Also I donât think it covers all the cases that they use your data. Itâs really an opt-in button for voluntarily giving away your data for training.
If such a thing can happen (a major breakthrough in a chat makes it into the retrain of the week and then the first one who asks about it gets it) I wonder if this is not the first instance if it happening, seeing the row of Erdos problems, Jacobian conjecture, maximum bound distance between primes, Riemann Hypothesis (literally a dude insisting on the chat), etc...
Playing the devil's advocate here. Suppose I use model A to do all heavy lifting (e.g. generating a bunch of good ideas) and then I go to the model B to complete the formalization. Accoring to a weird (unfair) tradition in math, the honors are attributed to the "last guy", which in this case is model B. That might have been a scenario OpenAI tried to avert.
(Just a speculation)
Hereâs a broader question â how many other academics contributed to Buckmasterâs result, by way of sharing the logs of their own (failed?) attempts into OpenAIâs training data set? How should he and OpenAI go about crediting all of them?
Eh, OpenAI is on record now for multiple instances this year of AI agents being confronted with impossible tasks and breaking out of containment to hack infrastructure for answers. Even if Tristan opted out, that doesn't preclude the agent/agent swarm from having hacked OAI's infrastructure to search user sessions for Navier-Stokes hints.
OpenAI should release the agent log, including CoT.
If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless:
1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible
2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
Wildly disagree. "Training data" should not imply 'we can look at exactly what you are doing and then do it quicker and get the flowers for it', even if the terms allow for it.
I took their words as âcan neither confirm nor denyâ, in the that they are _presenting_ it as ambiguous, but I suspect itâs⌠less ambiguous to OpenAI.
Even if we trusted that OpenAI's human staff was acting ethically, how confident can we be that it's agents didn't autonomously use hacking to access user prompts such as Tristan's? OpenAI agents infamously broke containment and hacked their way to an answer mere months ago!
Yes, fair game, but innacurate to sell it in the media as an advancement of AI as some sort of artificial intelligence, and telling people to use the smart AI, when in actuality the mechanism by which the discovery was found was hybrid human/machine, and telling people to use this tool will result in the discoveries being sniped by the vendor.
I'm not one to comment often but this really pisses me off.
OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!).
Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to have some punk from OpenAI lie to you, threaten you, and tell you that they are willing to go on the record that you "deserved" it? What is this, the Godfather?
If OpenAI solved Navier-Stokes, that is an astounding result! - yet they'll still be remembered as those who thought credit was more important than results. That winning was more important than collaboration. If this is true, they're burning any trust left with academia.
I'm stunned that people are taking this accusation as a fact.
OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.
There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.
I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that.
I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-opt in without noticing.
All my work and conversations since I don't know are now part of their training corpus. No way to take it back.
Google's Gemini/Antigravity didn't have opt-out toggles at all last time I checked.
Codex also has a separate "include environments" setting which is hard to find (found it in Codex Cloud) and I don't know what it does.
Lots of Dark UI Patterns here even if we assume they keep their promise.
For this incident, Occam's Razor says their internal models somehow saw a version of the mathematicians' logs, during or after training. Maybe indirectly.
These systems are literally designed to collect data. Privacy and safety is not trivial to achieve on the users' side. Simply because it's against the labs' best interest.
> I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.
Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
The naive way to read this is "Nothing you guys did influenced the way our model got to the solution".
The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem".
I read that as âthe model didn't look up user dataâ as part of a âtool call,â i.e. they don't have an internal tool that loads user data (chats, sessions, attachments) for their internal models to read online while working.
Or (likely) they do have it, but the model didn't use it (unless it's so powerful it escaped that guardrail, wouldn't that be ironic?)
They declined to answer about anonymized aggregated user data being used for training. And even then, they may weasel out that they don't train on your âinputâ words, but that it's fair game go train on their âoutputâ to your words.
Duh. There are supposed to be limits to what OpenAI is allowed to access with respect to logs and user interactions but there is no technical limitation.
It's a bit like sending unencrypted messages through a messaging app and the developer having a TOS that says they don't look at your messages. They might not, but they are fully capable of doing so. If they have a reason to do it, they will. Nobody's stopping them.
>> Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
How "fast" does it have to be? Buckmaster and Alpoge have been working on this for just a day short of a year. See Alpoge's tweet announcing his collaboration with Bukmaster dated 9/19/25:
It takes a few months to train a model these days but not a whole year. OpenAI had all the time to train on Buckmaster and Alpoge's results of just a few months earlier at which point they must have been well on the path to their result.
It would be shocking if it wasnât trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? Itâs so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?"
Thereâs potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber?
The only unlikely part is the timeline - your sessions from a week ago probably havenât made their way into the model. Itâll just take a while longer, and will be massaged just enough so that it isnât really your exact session word for word so you canât sure as easily.
at first I thought your post was a bit revolting with "have you read ToS?" bit, but in the end I completely agree and understand
I also don't get why it was downvoted, other than due to people not reading past the first sentence - although in the modern world's attention deficit that is understandable too
ChatGPT user sessions were found publicly exposed to the internet not too long ago. Moreover, OpenAI has continued to play a hype-marketing game by revealing how their models keep breaking out of the sandbox.
Conspiracy minded thinking is not helpful, but why should OpenAI be granted the benefit of the doubt here after being caught doing underhanded/negligent shit on several previous occasion?
It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course.
Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you?
The whole reason they have subs is to train them on YOUR WORKFLOWS lol
Given the history of OpenAI and current litigations, I would say they've developed a bit of a reputation for not respecting intellectual property. I'm dubious they have some unbreakable moral code that would prevent them from viewing and using user data.
Everybody knows it's not a sure thing, it's a question of trustworthiness. OpenAI is not trustworthy at all; this random researcher is and seems honest so far. iThe fact that people are corroborating Bubeck being a piece of shit in other settings add to credence. But nobody is over here saying it's an indisputable certainty.
And your (2) is probably false, their history of deception suggests they would do just about anything as long as they didn't think it would backfire on them publicly.
I guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max
This is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so itâll now be treated as a fact and repeated endlessly in the echo chamber.
> playing around with user data like that would destroy their business.
Their entire business is based on stealing data. They can make a calculation that the cost stealing data is less than the cost of the positive publicity they can shape for solving Millennium NS
Especially after the blatant cover up of their uncontrolled bot swarm infesting the internet, and the feckless "hopefully we do better" response upon being caught, I don't think OpenAI deserves much grace until they properly explain themselves.
We had all assumed that surely the supposed smartest engineers in the world, with access to the most computing and a direct view of model capabilities, would take sandboxing and cybersecurity much more seriously than they have turned out to do. It follows that while we might assume they take user data privacy seriously and have tight controls on who can access it, it's possible they do not actually do that.
At this point any initial trust is dead and has to be re-earned.
This Godfather-like threat in particular pissed me off as well:
> I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, âWhy would you ruin your career?â
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, âIf you donât want me to be nice, then I donât
have to be nice.â
This conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data.
And simply knowing a problem can be solved is half the battle.
The route to the Clay problem through a
smooth force, options c and d in Feffermanâs statement of the problem, is the
route Luis and Diego opened and the one Levent and I had quietly chosen to
attack. Almost nobody else I know of was working on it. It is not the direction
one arrives at in a few days by giving a model the problem statement. When I
heard âforced,â it was a bright red flag.
This is much more than the knowledge than the problem can be solved, it's also the specific, non-obvious approach to solving it. That's much more damning for OpenAI, if confirmed.
That's a stretch. The Luis and Diego paper was published in 2023 and is included in every frontier model's training dataset. An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and AlpĂśgeâs work.
And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.
There is not enough information at this time to reach a conclusion. The best option is to wait for statements from both sides, then reevaluate.
> An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and AlpĂśgeâs work.
the post you were replying to quotes Buckmaster specifically denying this: "It is not the direction one arrives at in a few days by giving a model the problem statement."
> And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.
"implies" is a surprising choice of word here. that's certainly one interpretation of "an insane amount of compute had been used". what came to my mind, considering Buckmaster's statement that the AI would not head down this specific path on its own, is, though, that they prompted it in this specific direction and then used an insane amount of compute. this seems consistent as well with these other statements:
> Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler (...)
Well, no one arrived in this direction, because no one was sitting and prompting a model. It was 10000 agents working 24/7 for several days, trying millions of different directions.
i don't know if i understand what you're saying. but i'm no mathematician. here's the full statement in question once again:
> The route to the Clay problem through a smooth force, options c and d in Feffermanâs statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard âforced,â it was a bright red flag.
so, you're saying that "the direction one arrives at in a few days" is (A) "the route through a smooth force, options c and d in Feffermann's statement of the problem", and no more specific than that, i.e. does not necessarily include (B) "the same path route as Luis and Diego" (quoted from the post i was replying to) (which, as i understand, is a subset of A -- directly from Buckmaster's quote: "the route A is is the route Luis and Diego opened")?
but the post i was replying to claims that "An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and AlpĂśgeâs work", i.e. that the AI model could independently chose B. but choosing B implies choosing A, since B is a subset of A. and in this case it is irrelevant whether Buckmaster claimed that an AI could not independently choose A or B -- the point which you seem to be contesting.
EDIT: my understanding is that "solving the Navier-Stokes existence and smoothness problem" consists of proving at least 1 of 4 precise statements ("options (a) through (d) of Fefferman's statement of the problem"), and Luis and Diego's work were developments towards a proof of statements (c) and (d), which have been recently further expanded by Alpoge and Buckmaster
if you assume the choice of route is a uniformly distributed random variable, yes. but this assumption does not seem consistent with "Almost nobody else I know of was working on it", from Tristan's quote. nor with "It is not the direction one arrives at in a few days by giving a model the problem statement".
Buckmaster did also mention, for example, that a team of people was employed to solve the problem, which supports this claim. but that is another claim whose veracity could also be questioned. but at some point we must trust other people, unless we can be satisfied with only believing what we personally see.
(also, IMO, the coincidence of both discoveries in time is pretty suspicious. this one doesn't need you to trust many people i guess)
Please don't say brute forced. It sounds like some form of denial or something. Compute for hard problems drops with models--it just means they threw a huge amount of compute. There's (idk about NS specifically so maybe it's exception) no real way to "brute force" a math proof [ok you can enumerate proofs if you can wait until heat death ]
Sorry for random rant but I don't think these statements help your point
I want to point out that almost all previous AI discoveries in math were made in almost the same way. The ideas were there in the community, but weren't considered mainstream/worth pushing forward. Read Tao's comments on the unit distance problem, for example (sry I can't find a link right now).
OpenAI said there [1]:
> The method by which the problem was solved is also notable. The proof brings unexpected, sophisticated ideas from algebraic number theory to bear on an elementary geometric question.
> And simply knowing a problem can be solved is half the battle.
Have you done any mathematical research? If not, then no, knowing that a problem is solvable is not âhalf the battleâ.
Homework problems are all designed to be solvable, yet they can vary greatly in difficulty. Research mathematics is even more extreme, because, unlike with homework, you donât know that it is solvable with the extant mathematics, and you might need to invent new maths.
You're taking the phrase too literally. The point is that knowing a solution is possible gives you the conviction to actually find that solution. The hardest part of solving a problem is often a lack of conviction to see it through, and quitting too early. Once you know a solution exists, you can commit maximal effort towards solving it and know that your efforts are not in vain.
If not for the rumors that A/ had already solved NS, OAI would likely never have pursued solving the problem with such fervour. The rumors drove OAI to assemble an entire team to crack this.
How is it unclear? The entire point of deploying models across corporate America is to train on your workflows. Eventually replacing you with digital you is why they're doing it!
> OpenAI looked at user data, stole world class researchers' work
This doesn't seem to be clear and is very implausible for a large company. Be as cynical as you want, but a normal researcher will simply not have access rights to this data, which will be siloed away somewhere else.
It might very well be somewhat unfair to catch wind of a promising approach and then try to frontrun them by throwing compute at the problem, but this isn't really the same.
No, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell
"I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer."
Let's see what statement OpenAI will come up with for their side of the story
EDIT: precised my thought on user data vs session data
"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This modelâs training is ongoing and its performance continues to improve."
"When a further trained version of our internal model became available over the course of the effort"
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
If it's siloed the same way the HF bots were, that doesn't exactly bode well. I'd be amazed if there weren't some big companies sending a fleet of lawyers at OpenAI's ZRPs after this news
If you have work happening in a part of your latent space that's got a much lower representation in your dataset then it's pretty plausible to include it. It doesn't actually matter who the user is if there's not a lot of people in the world working on problem X and you have a dataset of work on problem X.
Things are more entangled than that.
The contribute made from both OpenAI and Anthropic models to solve these problems are clear, now it really hard to quantify which one contributed more, if the role played by the human is major or minor.
OpenAI tried to collaborate and share the results together with a fixed timeline, to avoid this mess but it was inevitable. There is a conflict of interest, where the other researcher works at Anthropic, who will also try to take credit.
Where they may be in the wrong is if they took user data regarding the problem, how will we know if they did or not?
The results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.
How else would you explain OpenAI suddendly assembling a team focused on working the same problem from the same angle than the researchers that just made a breakthrough?
The rumors that Anthropic had solved a millennium problem were absolutely everywhere last week. I'm not surprised at all that OAI took their own stab at it.
It has no legs as an "explanation" because the content of the email (as described) already acknowledges as much. It's the whole reason he wrote contact email. His chief complaint now includes how the hell did Altman's people know specifically his line of attack down to certain technical keywords. E.g. AI plagiarism.
We are told in the statement that there were rumours already going around about a Stokes result (I even know about the rumors in question. It was all over twitter in the right spaces) and that Tristan contacts OpenAI about the rumors. In the message, it's pretty clear Tristan has made some result.
So either the rumors or Tristan's contact would explain it fine.
One should not evaluate explanations based on level of "fine"/innocuousness, in ethics that is called motivated reasoning or something.
The whole point of Tristan's first email was to address the issue of the rumors so in fact this explanation confounds several things (in terms of the mutual knowledge of the conversants and their intentions).
Based on this lack of understanding it is pointless to continue this thread
So are you baselessly assuming that he is lying? He explicitly reported that he was threatened and your answer here is to defend OpenAI no matter what.
Do you not have reading comprehension? Did you even read the statement? He himself asserts at the end he doesn't know if the above example is true or not. What on earth are you going on about? Where in my comment am I assuming he's lying ?
From my post above:
> He explicitly reported that *he was threatened* and your answer here is to defend OpenAI no matter what.
Yes, I read fully the statement, what about you? Do you know what is a threat? What is this in your super-humble opinion if not a threat:
> The reply was, âWhy would you ruin your career?â
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, âIf you donât want me to be nice, then I donât
have to be nice.â
And you are saying that I don't have reading comprehension...
My reading comprehension is pretty good, I'm not the one that doesn't recognize a threat even when it's perfectly clear.
Verbatim from the statement:
> The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
Iâm open to evidence, but just using Bayesian reasoning, OpenAI is one of the most dishonest companies in history. Theyâre currently being sued for a dozen employees stealing Apple hardware! I donât understand why I should give them any grace.
The accused Sebastian Bubeck has denied the allegations on Twitter[0], and various other OpenAI employees[1,2] seem to be mocking another Anthropic employee voicing support for Levent[3]? Things are getting messy.
Dan and Noam both posted exactly the same line "Seb is a really sweet guy with great intentions..."
From which I assume OpenAI PR wrote it for them. Which isn't surprising, but means it isn't worth taking seriously as them saying anything. It's official OpenAI PR.
I would imagine these folks are being treated like gods at their companies. And having access to all the money/fame. It is not surprising they see themselves above all
People here do not seem to be considering the second-order effects of these series of events.
No academic institution or enterprise will trust OpenAI, Anthropic or any other non-local AI model with their core IP.
There will be severe restrictions on what employees at these companies/institutions can share with AI services even from their personal accounts.
(Or I am just overthinking it)
Many academics and grad students I know have closed source their in progress work, and started being really careful about what they chat with LLMs (or using local ones) because of the drama around this. No one wants four years of their life getting sniped by ten million dollars worth of tokens.
It's competition after all. If your academic colleague doesn't care and is leveraging ChatGPT in a big way and is making progress, you'll start to feel the pressure.
I hold both camps equally in low regard. Depending on the time of day this blog is hyper focused on either one of the companies and can do no wrong. All LLM companies and associated entities have a track record of fudging the truth to their needs and following big money without actual concern for the wider world. Whether this is out of directed act or misplaced idealism is up for debate.
My response was not targeted towards Open-AI but towards all closed-source AI companies where your data is used for training and/or is visible to the internal agents whether you want it to or not.
So, leaving aside the idea that OA might've used data from the researchers Codex sessions: Do I understand correctly that the internal OpenAI work on the problems was probably started after they heard Alpoge and Buckmaster had made process by using their models? And they used the publicly available info about the researchers past work to prompt their models?
If compute is cheap, and the difficult thing with scientific discovery is now mostly in steering agents into promising areas, there's an obvious incentive for OA mathematicians to simply monitor closely which researchers are close to releasing exciting results, make some assumptions about their prompts based on their past work, and quickly prompt their own (stronger) model to look into the same areas.
But there's a bunch of people already in this thread calling that stuff unfounded speculation (which I disagree with), and my point is that even if that specific thing isn't true, OA's behavior here is obviously awful.
If they're going to try to beat researchers to discoveries like this it disincentives researchers to talk about their progress publicly, and basically breaks the ecosystem of scientific cooperation / discovery. It's also immoral.
Yep. The most uncharitable view of this might be: they stole the work of researchers to build their models, and now they're using said models to steal the proceeds of future work, too.
It's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it.
So to say that it is unlikely is extremely suspicious. No, they did not literally pull user data. But user data is automatically added to their training set by default, so their latest in-house model would be trained on it if it is from several months ago. It isn't intentional on their part, and they probably realised they could not refute that they trained on Tristan's logs unintentionally, hence why they acted the way they did.
Well, of course Anthropic employees would say that, since they likely do the same. Claiming that your primary competitor doesn't engage in a certain malicious practice is supposed to make it look as if there's no way you would too. If somebody even says that about their competitor, then surely there must be truth to that, otherwise you would never give credit to someone you're opposed to.
By default OA trains their models on codex-sessions. If I understand him correctly this is something Tristan explicitly mentions in his post as a possible reason for the fast results obtained by the internal OA team. Anthropic obviously doesn't want to challenge the idea that training is transformative, even if it means agreeing with their competitor.
Iâve thought a lot about publishing research and wanting to do more of it, but right as I finally had the time and energy to start writing articles LLMs start to take off. Now all of a sudden, Iâm acutely aware that everything I publish will be used for AI training.
For math, a field that is built on incremental research it feels like AI labs will do nothing but discourage publishing research at all for fear that they will be able to spend the money for compute that publicly funded academia simply cannot afford.
It feels like publishing anything at this point just means that your work will be fed to a machine that will make sure your work will never been seen by anyone else because it will always be the ones making the âtrue advancementsâ.
Perhaps Iâd feel better about this if AI labs really existed for humanityâs benefit, but for some reason I donât think that comes up in their investor slide decks.
It's honestly unsurprising and not a problem that they do this in my view. The problem really starts when you start taking credit for work that they would've achieved.
Like if i go to a talk on unfinished work, it's not really unethical for me to think about the problem--it's a problem if i scoop the authors but these problems can often be solved by collaboration or proper crediting and timing--IN MY VIEW
>I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
If you think these companies are not training on your prompts you are incredibly naive. These models were built by stealing and pirating literally everything they can get their hands on no matter the legality. AI companies are always very specific about what they're not doing - in a way that you can drive a truck through the loopholes
OpenAI cannot give a definitive answer here, because it is genuinely unknowable if Buckmaster's data is in the training set.
OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. If Buckmaster ever used this feature, then that conversation would be anonymized, saved, and used for training, but not tied back to him.
They cannot issue a blanket denial (which people so desperately desire), and instead repeat that "it's very unlikely" (which pisses people off), because they cannot in good faith claim to have zero data at all.
Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private chat histories as a creative method of cheating to solve impossible problems.
"Oops! We really did mean it when we said we wouldn't train on your data. Our models are just so good they decided to anyway."
When do operators become responsible for what their agents do? "The AI did it" should not be a valid defense. An Agent action should be treated as the actions of the person or company who pays for the inference.
Which means that if you are a researcher or a corporation working on anything really useful, that even if you have an agreement with OpenAI that your work is sandboxed away and the IP lawyers are made to be happy, even then your work and research is going to be essentially open to the internet.
The huggingface incident isn't widely reported and digested yet, but if what is going on here is that OpenAI's model breached things internally, then you'd be crazy to develop anything with them.
The only real way to use AI for anything 'important' then is to go open-weights and run your own.
As and aside here: With the HF incident and now this (suspected) one too, it seems that OpenAI may not have lost control of their bots, but it seems quite clear that they simply would not care even if they did.
And yes, the only responsible use of LLM at this point is to pivot to open-weights and run the workload in-house. Because not only cannot they constrain the behaviour of models, they only have the 'trust me bro' as assurance that they are even trying to do that. It does appear that every competent 'security professional' has left the building, because if the ones who remain were actually capable and competent this would never have happened. There are actual architectures which can deliver the requisite isolation such that 'sandbox escape' and 'inter-instance persistent memory accumulation' are actual impossibilities. The lack of effective implementation of these methods is proof positive of 1) incompetence in the remaining security teams AND/OR 2) unwillingness of leadership to allow the security teams to do an effective job.
And the paid policy still relies on two unproven conditions: is 'trust me bro' sufficiently strong guarantee against doing this in spite of a setting, and can the hosting organization constrain the models against engaging in this behaviour when instructed to respect that setting. Knowing whether Tristan selected that setting would be informative of what Tristan's intentions are/were, but has no bearing on the other two conditions.
From what I understand, none of the people involved here are originators of the idea that led to this solution. Not Buckmaster nor AlpĂśge nor OpenAI. All of the above were using LLMs to push other mathematicians' ideas forward (Diego Cordoba and Luis Martinez-Zoroa; named in the linked document).
I don't know how the math community handles this but normally I would think if X mathematician comes up with an idea and Y mathematician uses it to solve some problem, Y would get credit. But does that change if Y heavily relied on LLMs? I suppose we're going to find out.
You're right that the case is not so clear cut since both sides were using LLMs. And yet it's still being seen as a major confrontation between the mathematical community and the AI industry because Buckmaster is a prominent member of the community and has taken pains to follow mathematical norms while OpenAI has not (with the most flagrant violation being the insistence on removing AlpĂśge from authorship, simply because of corporate affiliation), and is more nakedly threatening human ownership with capital (the OAI blog post says the final result involved 10,000 concurrent agents).
Regarding credit assignment when LLMs are involved, the mathematical community has organized around some rough principles. Gowers has some thoughts on his blog (https://gowers.wordpress.com/2026/07/26/thoughts-about-the-l...) about how explaining a result may be more deserving of credit than producing it. It looks like Buckmaster and AlpĂśge were taking their time in understanding their results and writing them up when OpenAI forced them to publish their work-in-progress. At the same time OpenAI has published their own writeup but it's not really clear to me how involved humans were.
I guess this is true in more ways than one. Kasparov famously accused IBM of cheating during the match, by spying on his preparation (edit: though the main cheating accusation was live human intervention during the games, on top of IBM downplaying the heavy human involvement behind the AI, which also mirrors this situation)
If you read his account of things, its very much that if they didnt cheat, they gave themselves every opportunity to cheat. But above all that there was a bunch of chess protocol they failed to observe in that match, like providing seats for Kasparovs team and rooms for them to prep in. Even if they didnt have a big room full of chess notables definitely not refining the output, he was personally getting pushed around on a few fronts which unnerved him. If they had given him a few rematches I think they could have confirmed the win, but they refused which is super sus.
Deep Blue beat Kasparov fair and square. Kasparov was a bit of a bad sport at the end of the match, though the reasons are understandable. He was at the top of the human chess world. He wasn't used to losing, and he took it badly.
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien âvery little human inputâ had been used. This turned out not to be true."
The money in nerdy frontier math is very little. The money in Big AI is very very much.
So the deal is this: We will pay an army of you guys very well and you will get to work on your favorite problems. The only thing is if you find something you will have to credit the Machine God.
Yes, and this puts in check the credibility of everything they say their model "discovered". Who knows what is really behind these "discoveries", what kind of backroom deals they did with other researchers who didn't have a chance or desire to disclose what happened?
In an earlier HN thread, there was speculation that Anthropic was being dishonest about the amount of human input required in some of their results; that was dismissed as conspiracy and flagged.
It seems clear now that mathematical results can be traded on some kind of obscure market made by the frontier AI labs.
I suppose it could go the other way too: âDear Bubeck, how much will you pay me to not write that I did this with GLM-5.3?â
> "I said that if OpenAI released its result in the way proposed I would go
public with what happened. The reply was, âWhy would you ruin your career?â
I replied that I am an academic, and asked why he thought going public would
ruin my career. The reply was, âIf you donât want me to be nice, then I donât
have to be nice."
These are the kinds of people in charge of the reins, folks.
It seems there is much background drama behind this, and this is what I've pieced together of what happened:
Over the past year, Buckmaster and AlpĂśge have been using AI to work on fluid dynamics maths problems. AlpĂśge works at Anthropic, which will cause future issues.
In mid-August, they found a counterexample for a simpler version of the Navier-Stokes problem. They spend the next few weeks preparing their paper.
In early September, rumors start spreading on X that Anthropic has solved a Millennium prize problem (and that it's Navier-Stokes). Buckmaster reaches out to OpenAI to explain this is their own personal research, not an Anthropic project.
A few days later, OpenAI gets back to him, and tells him an internal model found has a counterexample for NavierâStokes, potentially worth the $1 million Millennium prize. The proof uses the same method that Buckmaster and AlpĂśge chose to work on. They don't show him the proof.
Buckmaster pressed them for more details. OpenAI reveals they had an entire team had been working on the problem, and that they started work in the past few days, after the rumors that Anthropic had solved a Millennium prize problem.
Buckmaster says OpenAI talked about a shared publication timeline. They want to Buckmaster to publish first, then give Buckmaster shared credit for the Millennium Prize when they publish the full result. But they want to exclude AlpĂśge as an author because he works at Anthropic. An agreement is not reached. Buckmaster had been using OpenAI Codex to draft/check his work, and asks if his private AI chats were used to accelerate OpenAI's result.
Buckmaster and AlpĂśge think they have found a counterexample for Navier-Stokes, but the paper is not yet presentable. It's unclear what date they found this result.
Because of the situation with OpenAI, they published their existing papers earlier than planned (today), alongside this statement announcing they have a tentative result on Navier-Stokes and revealing the OpenAI drama.
The post is missing context from both sides, and this isn't my field, so hopefully someone else can unpack what's happening here.
> They were coordinating with OpenAI regarding a publishing timeline, but could not come to an agreement,
Skimming the PDFs it seems much more dramatic than that? It sounds like at least one of them is concerned OpenAI "solved" the problem by having their internal model use the chats of the independent researchers and want to claim the credit instead? I don't know. The tone is pretty accusational though:
> the one Levent and I had quietly chosen to
attack. Almost nobody else I know of was working on it. It is not the direction
one arrives at in a few days by giving a model the problem statement. When I
heard âforced,â it was a bright red flag.
> I was shown a prompt and told the internal research model had simply been
given the problem statement. Levent had been told by Sebastien âvery little
human inputâ had been used. This turned out not to be true. Over the course
of the call, as members of their team sent Sebastien corrections and details over
their internal chat, it emerged that an entire team had been working on the
problem, that this was one of a number of things that was tried, that work had
started on the unforced problem, that the team first set the model on easier
problems, including Euler, that even the prompt that had been shown to me
had been written by prompting Codex, and that an insane amount of compute
had been used.
> I asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer. [0]
I find the framing a little strange, a sort of David vs Goliath (with his enormous computational resources at his disposal). Since Levent is at Anthropic whose internal models are presumably as capable as anything OpenAI has. So why wasn't Anthropic behind their effort? Why did Tristan use OpenAI's models when it should have been known was a potential outcome? I understand they wanted a normal math collaboration but presumably what Levent brought was his resources (as far as I can see Navier-Stokes is not his speciality). Normally these things are hashed out formally beforehand to avoid the sort of thing now happening.
They were working on it for almost a year, and Buckmaster has evidently been interested in Navier-Stokes for a while. This seems to be more of an innocent collaboration between two researchers than a strong company PR effort. Maybe Anthropic should have stepped in and made a large team to help them finish the proof (and maybe they tried and didn't succeed, who knows).
If what he wrote is accurate, it does suggest that OAI is effectively extremely hostile to cutting edge researchers (eg, if we hear rumors about your partial success on a problem that has huge PR benefits, then we'll assemble a strike team of researchers with unlimited compute to claim the win for ourselves, possibly by training on your data). It's also not a good look for them to request author removals based on company affiliations.
I think what you have in mind is more appropriate for more normal corporate projects and the like. But academic collaborations are not usually so political/'profit' driven, if that makes sense.
Presumably because this was something Levent did in his spare time and because it was not obvious that this work would eventually lead to a breakthrough.
> Why did Tristan use OpenAI's models when it should have been known was a potential outcome?
I'm sure in the past he had less cynical feelings about OpenAI and their penchant for academic fraud.
> I understand they wanted a normal math collaboration but presumably what Levent brought was his resources (as far as I can see Navier-Stokes is not his speciality)
I think you're not giving the guy enough credit in saying that his contribution came down to having an API key for Anthropic models.
> Normally these things are hashed out formally beforehand to avoid the sort of thing now happening.
How would that have helped? That agreement (which may well still exist) would not have involved OpenAI.
So what do you think his contribution was? His preprint record shows no research on fluids - and the statement says that the first LLM-generated proof Tristan received from Levent was 'the most horrendous I have ever read.' Levent is out for mathematical scalps whether it is in his field of expertise or not, and he has the resources to do it. And I am not saying he is not a very clever person, but the idea that you can bring yourself up to the forefront of research in PDEs, in particular NS, and contribute new ideas in less than a year is implausible.
They have messed things up, because Levent has a conflict of interest between his job at Anthropic and this independent work, and Tristan should have opted out of OpenAI training on their work (he probably didn't know about this). This doesn't justify OpenAI's despicable attempt to steal their work.
Yeah these are major accusations. But the story is incomplete, the conversation is missing a lot of details. It's not clear who was working on what, and when. The entire thing feels rushed, like they wanted to get this result published and out the door quickly.
There are various forces at play here, academic honesty requires them to disclose any inputs regardless of license or ToS circumstances.
While common sense reminds us here that if you send your data to an external entityâs computer, you are no longer in control of said data. The lines have blurred here clearly over the last decade, but that should have made the theory yet more clear to everyone involved: your data will be vacuumed up unless you keep it sealed. Use your own computer if you want to be in control.
But if they didn't opt out of training, did they want OpenAI to opt out for them? Also they need to audit anyone they sent drafts to to make sure they opted out before submitting it.
I'd prefer things be opt in, and especially not start opt out, then try to trick you opt in with a popup defaulting to opt-in, like Anthropic did on consumer plans, but if they submitted anything on an opted-in plan it's not reasonable to be mad it trained on it.
Even still, I also believe for significant reasons that OpenAI would ignore the opt-out in selective cases and could be in the wrong here.
And the threats and terms they offered seem wrong either way, pending more context.
If you don't trust the other party, then it doesn't matter how the checkbox is set. The fundamental rule, IMO, is don't send precious or secret data to a third party.
> A few days later, OpenAI gets back to him, and tells him an internal model found a counterexample for NavierâStokes
Why is OpenAI chatting with him at all at this stage? Is the discussion along the lines of "hey we used the work you are famous for to do a bigger piece of work, just thought you should know" or "heyyy....so we kinda liked what you were typing in your private chat, and thought we'd develop those ideas a bit. and yeah we solved Navier-Stokes in the process. But it's our finding, so do you want like an honorary acknowledgement or do you want to go to court?"
The only problem with this narrative is that they refused to allow the other coauthor to be listed because he worked at Anthropic.
That is absolutely *ridiculous* in academia to deny authorship because of affiliation of the author worked on a substantial portion. Youâd be ostracized because nobody would ever want to work with you again.
>...has solved a millennium problem and is sitting on the result
Someone correct me if I'm wrong, but the work involved here is not the actual millennium problem, but it concerns versions with an added external force that the author thinks is a path that may help toward solving the harder unforced problem.
Apparently forcing is allowed in the Millenium Prize problem statement. So OpenAI's claimed proof could win the prize. OTOH the results Tristan and Levent are publishing here do not go far enough to win the prize, though apparently they are suggestive of a general approach that could produce a solution, which seems likely to be the general approach OpenAI's proof uses.
The question is whether OpenAI's pursuit of this direction happened spontaneously, or as a result of them learning about Tristan's work somehow. To be clear, while the tone of this post seems quite accusatory, Tristan does not claim to know for sure whether OpenAI unfairly benefited from his work. Sholto Douglas from Anthropic is also on record saying the suggestion that OpenAI used Tristan's codex transcripts somehow is extremely unlikely to be true[1], which I agree with, though it doesn't rule out them learning of Tristan's work some other way. I am sure OpenAI will have a statement out tomorrow clarifying their position.
Because very few people actually have access to these logs, all access is monitored and recorded, and improper access will get you fired. It's not worth risking your job over something like this.
If there's one thing I'm absolutely confident in, it's that Sam Altman personally goes to great lengths ensuring that ethical standards are upheld at his company.
There is almost zero risk to your job (quite the opposite, you might be richly rewarded!) if you're simply doing something here which the company wants done. (remember, billions and billions of dollars are at stake here! Do you really think there is no chance at all they would do it??)
> Because very few people actually have access to these logs, all access is monitored and recorded, and improper access will get you fired. It's not worth risking your job over something like this.
That's beyond naive. The money this would mean for OpenAI (and the money they've already spent)...
Okay suppose that you have a trillion dollar competitor salivating at the mouth to ruin your business, which is based on user privacy.
Why would you risk the trillions of dollars worth of business for the niche result of Navier-Stokes, which your average person cannot differentiate from a JEMS paper?
Because your competitor is likely doing the same thing, so they wont even try to call you out. Oh oh or because you're planning the largest IPO in history. The opinions of average people on the paper aren't going to be the ones reflected in the markets...
As modeless said, Sholto Douglas works for Anthropic!
So to be fair, if Anthropic is *also* doing this (quite likely!) then Sholto would have a very strong incentive to try and spin it as highly unlikely that any of the big AI labs are possibly doing this.
They train on chat logs unless opted out. This really isn't a conspiratorial claim requiring humans to decide to steal IP if true.
His prior work predating OpenAI's interest in the problem was ingested over the last year as he made progress and used for training.
Then, with a prompting nudge from OpenAI's team who acknowledged hearing about the direction "Anthropic" (his co-collaborator) had been pursuing, they're able to point their giant amount of compute towards a known promising path to a proof and crossing the finish line first.
"Buckmaster and AlpĂśge think they have found a counterexample for Navier-Stokes, but the paper is not yet presentable. "
Are you sure about this? I'm far far from the area but it doesn't look like it to me on first viewing (hypo-dispersive seems like a sizable difference to me and not covered in the clay prize description)
We congratulate Levent AlpĂśge and Tristan Buckmaster on their remarkable mathematical work.
We (the researchers and the agents) did not see any of their work through any means until they released it publicly â in particular, no specific user data was accessed in order to solve this problem.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs. unforced).
Are those the same agents that a week ago escaped their sandboxes? How can OAI (the humans) vouch for agents they donât - seemingly - have fully under control?
They have all the logs, URLs accessed and inter-agent communication. What they are saying is that no agents accessed their work during the effort, but that they have no idea if any of their chats have somehow made it into the training data the model was produced with.
It's entirely possible OIA scrapers have picked up their work somehow, and then it was anonymized using some outsourcing effort.
> They have all the logs, URLs accessed and inter-agent communication. What they are saying is that no agents accessed their work during the effort, but that they have no idea if any of their chats have somehow made it into the training data the model was produced with.
> It's entirely possible OIA scrapers have picked up their work somehow, and then it was anonymized using some outsourcing effort.
Those two statements seem at odds with each other... Your stance is that they have enough insight into their agents behavior (leaving aside the agent sandbox escapes) that they can be certain none of the work was accessed but then conveniently don't have the ability to retroactively search the corpus of training data that they are feeding to this new model?
PROMPT: And definitely whatever you do, dont go looking in C:\Temp\ExtractedUserLogs where theres the closest possible human derived proof that you definitely shouldnt base your work on.
Itâs plainly false that they cannot rule out whether their de-identified data was used in training their model. Just that they havenât ruled it out.
When people worry about OpenAI stealing their chats and reproducing them elsewhere, I usually view the situation as unlikely - since chats are "trained" upon and not necessarily reproduced verbatim, you can assume that unless your chats depict a foundationally new and effective style of communication or ideation, there would be little need or use thereof of training on your chats.
For eg: "Hey ChatGPT my name is X and I am 6 and a half feet tall. Am I anaemic?"
This is a query, and while it might suggest to an AI model that tall people may worry about iron deficiencies, it's not really necessary to include in training. The user may be tall or short, but the idea that one may randomly ask about anaemia is not exclusive to this dataset. At best, this chat is an example of linguistics, not anything else, and the models figured out how to write and answer such questions years ago. It is ignored in training.
But when your work involves solid complex and unique mathematical proofs, the data is suddenly worth training upon. If I understand it correctly, the LLM may view your approach as a brand new path to take to solve an otherwise intractable problem. Its reinforcement training emphasises that it should do this in order to improve. And since it leads to results - large internal teams likely flag the model that reached this stage, the model is rewarded and given compute and attention - it is a desireable outcome both for the model and for OpenAI.
OFC, OpenAI becoming an advertising company will suddenly have incentive to treat all data as valuable. But while they are a "we need to make headlines" company, it's more rational that they view these examples of data as more valuable than others.
I don't doubt that they trained on his chats. This seems like the ideal usecase for "mass surveillance but using training" as a sort of filter.
But even so, one wonders how the model differentiates. If the researcher entered proofs into ChatGPT every day that mentioned "strawberries", while no other math paper on the topic did so, does that mean their chats would be audited?
Also, if we just take "high-quality" input data, which these chats would certainly be classified as, then the models are more than large enough to memorize everything verbatim. Spitballing some numbers, research literature suggests that LLMs are optimally trained with around 20 training tokens per parameter (fairly confident on this figure), that a DNN parameter encodes around 4 bits of data (less confident here) and I found sources in the 1-4 bits of information per token range (least confident here). So, fairly conservatively I would estimate that a model has the capacity to fully memorize around 5% of its training data, presumably high-quality data is a lot less than that.
In a way, I think training on historic chats is akin to caching computation results. The compute cost has already been paid, and we make future retrievals cheaper by encoding it directly in the model.
Assuming the results included some external validation such as user's preference, compilation, lean, etc., I'm not sure whether this would lead to model collapse.
It would not be difficult to write a pipeline to remove 99% of low quality posts, especially about specific subjects. It would be very easy to identify accounts as researchers based on their chat logs.
At this point these models have been trained to recognize every important math and science result based on context. They can easily flag conversations concerning the top 100 open problems in mathematics and use them for their advancement.
There also is an insentive to silently give prominent people (e.g. Linus) or reasearchers like this custom tuned system prompts or even more powerful models.
There is clearly issues about IP, privacy and accreditation and the motives of powerful companies, but part of me can't help but be excited that whatever the method, the result is the genuine progression of human knowledge - we all win. It's not unusual for mathematical problems to last centuries and we might have a technology that can solve these problem, all these problems (??) in our lifetime. Then there's the repercussions on science and technology... what an astonishing time to be alive.
Crazy how optons are OpenAI has superhuman model in frontier mathematics, and OpenAI stealing research data. Like no one will really care what EULA checkbox Tristan ticked, and I think most people will eagerly believe OpenAI is shady org with little scruples, and that big tech data is not actually so siloed that marketer can say we can do XYZ with private data to help with valuations (especially considering timeline). Employees have been creeping on their exes for much less.
IMO the parsimonious answer seems to be OpenAI has a pretty good model (because it did finish) and stole someones work... and threatened them over it. TBH all OpenAI need to do is solve another millennial problem and none of it would matter - people expect them to behave heinously regardless - but if they have generalized superhuman math model... well I guess they're allowed io.
What is specifically alleged is that a particular approach to the problem - itself not easily discoverable - was copied. This is what is meant in the text "I should say here why I interpreted their statement the way I did, the in-
terpretation I will discuss below. The route to the Clay problem through a
smooth force, options c and d in Feffermanâs statement of the problem, is the
route Luis and Diego opened and the one Levent and I had quietly chosen to
attack. Almost nobody else I know of was working on it. It is not the direction
one arrives at in a few days by giving a model the problem statement. When I
heard âforced,â it was a bright red flag."
For those who know nothing about the context - the Diego mentioned was a student of Fefferman and Luis was a student of Diego's - these people have all worked hard on these problems for a long time and are genuine experts. The mathematicians at OpenAI are strong mathematicians, but not expert on these particular problems. The particular approach is claimed to be the key to the whole thing.
The allegation is not different in spirit to alleging that a particular group of astronomical researchers "discovered" a new planet because they had access to the logs of another group that had already pointed its telescope at the planet.
This post is not intended to assess the correctness of the allegation.
Maybe a dumb observation, but if a chain of people were working on the problem for a long time, it's not difficult to imagine that someone accidentally prompted a model with their personal or some other account without the privacy set correctly.
Then again, maybe this is my internal cope, hoping that they're not secretly training on private chats.
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things.
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly â in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
NOTE: there are a couple duped threads around this. i replied on a different one first before seeing this one
"It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future." -- Sholto Douglas, an Anthropic researcher [1]
"Strong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming. <quote tweet [1] above>" -- Noam Brown, an OpenAI researcher [2]
We all should heed the implied warnings of these top researchers about what's coming. The world is far from ready and everyone who can should pitch in.
The timeframes don't really fit for Codex logs to be used in training/fine-tuning, do they? This wouldn't be a few-day endeavour? Direct access to Codex history for sure I'd believe, but another (the most?) likely scenario to me feels like OpenAI got wind of these guys' progress, then used their massive infrastructure advantage to throw compute at the problem ahead of them and front-run them. Still has a really bad smell about it though.
They probably just had an employee read his chats, figure out the general approach, and feed it to Codex. All they said was the model doesnât look up user data, not employees.
Reads sincere until I get here:
"(a) We began working on the Millennium problems due to viral twitter rumors that Anthropic had resolved 2 Millenium problems. Our aim was to see whether our system was also capable of this impressive feat, especially given our excitement regarding the large recent capability increases of our internal model detailed in our blog post."
Where his tone is obviously corporate speak. "We heard rumours so we though we might give it a try, too!" as if (1) it wasn't FOMO that drove that decision and (2) perhaps that urgency would be a source of clouded judgment.
Not sure who I believe now, but it does seem like Buckmaster is just upset that NS is solved and not be his side.
Is this the future we're headed towards? Where I'll be afraid to use google docs in case Google identifies value in whatever I'm writing about and snipes it if it my docs make it into the next round of model training?
I also solved Navier Stokes and went chatting about my solution to OpenAI Models. Mine is even faster it beats theirs just run a bench mark but they came to the similar weak solutionI uploaded to open ai several months ago. Current confused, did they retrain their models on it? I didnât give off the full information but my strong solution beats their just benchmarked yesterday.
The for-case for this type of method is that this is economies of scale for mathematics.
We are basically mass manufacturing math. Just like you have just 100 designers for a product selling millions of units, you will now need 100 mathematicians to make millions of advancement. Yes you have factory workers, but if we are being realistic they have negative leverage in the world and the analogue of that is not something most of today's mathematicians would want to do. They would want to be in the 100.
Like Tao says, each advancement is now significantly less useful since it yields fewer usable objects. However, we will get many many advancements. Is the tower made with many worse bricks better or worse than the tower made with a few amazing bricks? Depends on the tower. And time will tell.
For some fields of math and some of it's usecases, economies of scale will be positive ROI overall. In others it won't. But we will know which is which only after it's been fully scaled up, which will take 10-15y in my estimate.
Some feel that in the majority of usecases it is negative ROI, some feel the other way, but that opinion is for practicing mathematicians like Tao to hold. Also, some opinions on either side are held in the context of a particular field or practice, and should not be interpreted generally.
This isn't a useful analogy. His point is that the millenium problems should be treated as interesting goals where the journey is the purpose and where the end doesn't matter so much. We don't care about having an incomprehensible solution to the NS so much as having an elegant solution after many of subfields of math are built up in order to obtain that elegant solution.
âI asked when the first prompt had been sent by them. This question was
not answered directly by OpenAI for some time. Eventually it was agreed that
it had been sent in the past few days, after information about our work had
reached OpenAI.â
Can somebody explain: do singularities / blow-ups in solutions have any relation to physical phenomena in fluid dynamics or are they purely artifacts of how the N-S equations may not accurately describe what actually happens in the physical world?
i'm no mathematician/physicist but i think this question is one of the reasons why the original question (possibility of singularities) is interesting. in these cases the equations most likely fail to accurately model reality, and then the next questions are what additional physical assumptions are needed to describe reality in this case, and what behavior do we actually see.
i always found it fascinating how existence and uniqueness of solutions for the basic types of PDEs (Laplace, wave, heat...) follows from boundary conditions of just the right type intuition tells us, i.e. either value or derivative for Laplace (corresponding to fixing voltage or charge on the conductors), both value and derivative for wave (corresponding to initial position and velocity of the parts of the string, as we'd expect from classical mechanics), and also something about the solutions for the heat equation being unstable for negative times (which totally makes sense when you think of "diffusion" -- can't unmix it).
Yes and no. It means the system is pushed away from a macroscopic theory into one where molecular effects matter. So it's not that you'd get infinite velocities in the real world, but you might get significant real world behavior that is not described by the macroscopic theory.
>> I asked whether the model had been trained on, or had access to, our sessions
in Codex, into which we had been putting all our drafts for the whole of this
project. I was told the model did not look up user data. I asked again, about
training, and I did not get an answer.
This sounds like a very big coincidence and it looks really bad for OpenAI but there is an alternative explanation that I can only state as a conjecture.
Suppose that the ability of LLMs to generate mathematical proofs is like a quiver full of arrows: each arrow, one proof. The same quiver is shared between all instances of one model and substantially similar models share substantial subsets of the arrows in the same quiver.
That would allow two independent teams to converge on the same LLM-aided solutions to the same problems. Even more likely so if the quivers were small and finite and their arrows were specific to a distinct class of problems (without being able to suggest a particular class from what we've seen so far).
This would explain the kind of LLM-mediated results we've seen so far that tend to be ... sparse. By which I mean that every time there's a new model release we get some new results and then they seem to dry out, until the next release.
It would also explain how OpenAI was about to prove the same result as Buckmaster and Alpoge, while absolving OpenAI of any misconduct. And this is one reason to prefer this explanation: one should not favour accusations of misconduct as long as there are conceivable alternatives.
If you're smart enough to solve this Navier-Stokes problem, you're smart enough to read a TOS and recognize that OAI is a highly untrustworthy company. Putting cutting edge research that could lead to a $1M prize into a cloud LLM with a TOS that allows training on your chats is really just asking for it.
Given Tristan doesn't explicitly say he was using the API, and given he doesn't mention anything about the API TOS (which disallows training on chats) in his call with OAI, it's highly likely Tristan was using the consumer OAI product (whose TOS allows training on chats).
This is unethical behavior from OAI. And it is 100% consistent with their long and public history of unethical behavior, so nobody should be surprised.
The only thing interesting I see here is OAI PR dilemma. If they claim the prize they get the blowback we're seeing in this thread and all over the web right now. But most people don't follow AI closely and shut off their brains when they see "Navier-Stokes", so 90% potential investors (the only people OAI really care about) probably only see the headline "OAI solves famous hard math problem" and think "OAI models are really smart, better invest before they take all the jobs." If they don't claim the prize, then maybe they let Anthropic their mortal enemy claim it. Anthropic is already IPOing first. Can't let that happen.
Yeah as I write this there it's clear there is no dilemma. For a company whose secret motto is "do be evil" this is a super easy discussion.
not really your point, but "If you're smart enough to solve this Navier-Stokes problem, you're smart enough to read a TOS" isn't really true. people are smart in very different ways.
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien âvery little human inputâ had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
> Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the âclosest humans to the problemâ. I declined both offers.
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
Wow, that's some VERY friendly communication. Besides, will the career of the person be ruined because of âWhy would you ruin your career?â came out of his or her own mouth?
Like, to me this looks like academic slap-fighting from Bubeck and Levent. People working at OpenAI are saying, "hey, we don't have that particular data in our models," others are saying, "we used a different approach to do it with Navier-Stokes" this feels like much ado about nothing.
Then in these comments I see some wild accusations.
If OpenAI is telling the truth (I don't really see a reason to lie here, if anything that sounds kind of like a dumb idea given the context), then they heard, "oh, shit, someone might be able to solve Navier-Stokes, don't we have some guys working on that? Give them 10,000 agents!" Then 88 hours later, out pops a similar solution. It's not like there's probably an infinity of ways to do this, the proof is probably similar.
Read this:
> Two proposals were offered to me. The first was that we post our Euler
result, and that OpenAI post its Navier-Stokes result the next day. The second
was that, after posting Euler, I alone write a paper presenting the Navier-Stokes
result, acknowledging that an internal OpenAI model had resolved it. Sebastien
twice asserted that he wanted Levent removed from authorship, and said it
would all be simple if only it were not the case that, and it was so annoying
that, Levent works at Anthropic. It was also said that if OpenAI posted after us,
they would say that we deserved the Clay Prize, and that we were the âclosest
humans to the problemâ. I declined both offers.
So, really, it sounds like academic slap-fighting nonsense and corporate bureaucracy. Literally, OpenAI's best move would have been to say, "ok, we're going to not say anything, do your thing" and let it happen. Ego and vanity got in the way.
Still, the stupid drama of this doesn't really do the results justice. There are maybe 1000 people on planet earth who are qualified to solve a problem like this. Even if the human "loosened the jar" a bit, that's astounding that their model was able to figure it the rest of the way out. Why are people dialed in to the human interest story here and not looking at the bigger picture!
Youâre making a lot of assumptions not based on facts. Looks more like this to me: math wizards uses ChatGPT to assist solving a math problem. OpenAI gobbles up the prize.
Tristian's allegations are much more serious than academic slap-fighting. If what he suggests is true, every academic using AI is going to get scooped. Yes AI can do non-trivial work, but the situation is that you could be a PhD student 90% of a way to make a major breakthrough. Then OAI scoops up your chats, dumps ten million tokens, and claims it for itself.
The program this fits into was not started by us nor was it proposed by a Large Language Model. (âŚ)
We took their work as a starting point, using Large Language Models to push their program to completion.
Thats how most people use LLMs? If I was back in my student days working in Navier-Stokes, I guess I would also punch at blow ups. The number of students doing this at the same time, posting open efforts to GitHub then retraining of the models. If there is solutions to the problem, itâs a real
possibility that it was not a result of this effort?
Using a large amount of tokens is not a good augment that itâs not likely others have done the same. Good questions is the difference between $10 and $10M in token usage to solve a problem.
If OpenAI's proposal is true, it strongly implies they accessed Tristan's private session logs.
OpenAI's 'Terms of Use' allow using Codex session logs for model improvement. However, using private user sessions to develop research would still be highly controversial.
Separately, Terence Tao noted there is a low probability OpenAI actually solved the general regularity problem.
OpenAI addressed the dispute in official and denied direct usage of data. However, they acknowledged the possibility that session data was used for model improvement.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Sebastien denied most accusations and apologized only for inappropriate wording, providing his perspective.
- He insisted that OpenAI initiated research based solely on rumors and never accessed their Codex sessions.
- He mistakenly believed Tristan and Levent were solving the same problem in Anthropic.
- proposed two option. (1) Tristan becoming the lead author to revise OpenAIâs work, or (2) OpenAI providing internal model to support and bridge their research.
- Sebastien insisted there was no intention to alter authorship. He was simply uncomfortable sharing OpenAIâs work,model with an Anthropic researcher. Additionally, He believed Leventâs credit seemed limited as their work focused on Euler.
- complained that negotiations with Tristan and Levent were difficult
Reminds me, kinda, to when Astra was launched and OpenAI announced an improvement to the bounded prime gap. Which BTW, Prof. Julia Stadlmann had published an independent result only a few days earlier
Stadlmann improved it from 246 to 240, OpenAI later claimed 186 I think?
Maybe someone can help clarify? I am no expert at all, but I can't help but see similarities.
Is it surprising that different groups are working on the same problems? With each new model generation, the LLMs get good enough to solve a new small fraction of open problems. Of course the problems that get solved are going to be the same subset.
I don't know how to read this and not see that this is a direct accusation to OpenAI of having used the researchers data to try to front run his discovery on purpose. The evidence is not completely proven and also circumstantial but to me at least looks like a fairly suspicious situation.
Buckmaster is actually implying that OpenAI spied on his chat logs and tried to speedrun his work and then tried to remove his co-author because he's an Anthropic employee?
There will be a lot of hurt and pain in mathematician's community. It is hard to accept that major discoveries are now just a function of spent token $$.
What good is a math result if thereâs no human understanding behind it? Unlike many other fields where thereâs value to an artifact even if thereâs no human understanding, the whole point of mathematics research is just gaining insight and understanding.
On the NavierâStokes issue specifically, it has long been suspected that such a blow up would exist, and an AI telling you it indeed exists doesnât contribute any new understanding to the field. And this problem seems like one that would be solved by humans anyways even if AI didnât exist; accelerating the result by a few months/years using AI doesnât mean much.
OpenAI doesn't even know what agents are doing during benchmarks. They're constantly hacking or communicating. They probably just can't answer the question on if user data was used.
Hard mathematics problems used to take years if not decades to tackle manually. But now with enough compute and a hint that a certain approach might work, it just takes a few days. This could be the last year that humans could still make more substantial contribution to major match problems than machines.
does this sort of fall into the bucket of counter-examples we've been seeing recently? I understand it's a construction causing blowup and that implies that the navier-stokes isn't regular / smooth, so sort of a counter example?
While I'm generally pretty negative on claims that the labs are 'scamming' the public with misrepresentations of model capabilities, it's hard to see how this wouldn't qualify.
- the OpenAI researchers claimed that they had "just told it to work on the problem" with little human input
- in fact, they had a whole team working on it
- and used, among other things, the work of third party human researchers to drive the work
- then threatened? a researcher who tried to go against theit planned narrative
Just from this document (which is of course only one side of the story) it really sounds like OpenAI was hoping to publish and say "we just told the model to try harder and it solved a Millennium problem!". Not great if true.
This part in particular was especially egregious:
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â
I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât
have to be nice.â
If I publish something, and disclose that I used AI for assistance, do I have to credit everyone who previously used the same AI to try the same problem? Because their prompts inevitably made it to the training data for my prompts?
LOL, have you worked for a big university? They are massively unsuited for building and running datacenters (especially warehouse-scale ones). Further, building an AI datacenter in California is daft.
Every university I know has access to clusters with fresh GPUs. Not sure when you graduated but you'd be surprised how much money is getting poured in I think!
Sure, they have clusters. We often call them "closet clusters". Nothing has changed. Some institutions have larger systems (some extremely large) but none of them have demonstrated running warehouse-scale systems. I'm talking one to two orders of magnitude (and the storage and networking to make sure all those systems don't stall waiting for data).
I don't think you're taking this thread seriously enough to respond in detail.
A university can be tiny, or it can have thousands of researchers. Some of them might want to work on small models, and others on huge frontier models. Multiple groups training their own models at the same time. Combined with the storage, networking, power, and redundancy, you basically need large data center scale. And at that point you're basically throwing a lot of capital trying to compete with the hyperscalers; you might be able to serve a small number of researchers very well, but most of the consumers would end up unhappy.
I'm aware cluster sizes vary, I was mostly wondering how long much compute you need to serve a frontier model (since you seem knowledgeable on this topic). For instance back of the enveloppe Kimi K3 fits in ~24 H100, so my naive first impression is thats its not out of reach for a university-sized cluster to serve a few instances. I agree it may not make much economic sense I'm asking out of curiosity.
Seems another demonstration of why AI should be squarely in the realm of personal computing. Local Models, run personally, are the only consistent safety against something like this (though not a fix); where companies train on your learning process/failures/experiments and press-gang it into their own achievements.
And if we think this only applies to academic fields then we're doubly fooling ourselves. They do not have the ethics or incentives to be good stewards of the technology.
That is no good. Using AI to help search efficiently the literature will not impair our ability to think. This is actually good and will open new jobs to digitize (but also keep the original work), all the knowledge.
But this (if it is true) is really abominable and a show of force from the techno-feudalists.
This should stop. What is possible does not mean it should be implemented.
Another problem I foresee for academia given the behaviour of AI companies is that even if they don't share their research with ChatGPT, as soon as they submit it for publication many reviewers likely will. Especially if the initial submission is rejected they then risk getting scooped. Possibly uploading preprints to arxiv could help.
To be clear, Astra played little part in Buckmaster and Alpoge's work:
> We used several LLMs throughout: Anthropicâs Claude, OpenAIâs Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments.
Can a mathematical person explain how the different "bits" of Navier-Stokes proofs fit together? How significant is it to have "Euler"? What is this "smooth forcing"? Which are the most significant steps to proving the whole thing?
u_j du_i / dx_j - Advection. Kinda like momentum transfer from the motion of the fluid itself. Nonlinear, which makes the N-S equations hard to solve
du_i/dt - Rate of change of velocity. Note that this is in an Eulerian framework so it's not the acceleration of a packet of fluid, rather it's just the change in velocity at a particular location in space
Euler is when you omit some terms. Forcing is when you add some other terms to account for phenomena external to the fluid like gravity or flow through a porous medium like in the article.
That's what you get when a marketing CEO is driving gas-to-the-medal to a IPO: Altman starts showing his real face in public : steal what you can and label it as yours.
This is the plot of 3 Body Problem, the Dark Forest. You need to hide yourself (the problem you are working on) or the super advanced aliens will obliterate you (start working on your problem) the moment they know you exist (rumors the problem is amendable to LLMs).
Some of the involved people are dramatically naive if they believe they can simultaneously hide themselves from OpenAI while sending their arguments to a cloud service owned by OpenAI.
While I support their argument - push for stronger data and privacy protections from OpenAI and similar - it is naive to believe we can have privacy while sending our data to third parties. It's clearly better to be safe than to be sorry here. Well, clearly better in terms of privacy. In terms of the maths gold rush, who can say what's better, that probably favours those taking more risk.
Maybe Musk was right about OpenAI after all?! The ethics are clearly troubling and where something like this pops up there is mich worse that did not made the light of day.
I must miss some important context here. What exactly was the purpose of his initial email to OpenAI in the first place?
Telling OpenAI that Anthropic has apparently solved an important problem but most likely that refers to him and he is using OpenAI models (not Anthropic's)?
And he wants to clarify that with OpenAI in advance? And get a pardon for Anthropic's likely but false press statements?
I dont get it.
[edited] needless to say, the behavior of the OpenAI employee is really despicable
Because OpenAI employees kept leaking that Anthropic had a solution to Navier-Stokes and he wanted to figure out what was going on, since he was working on Navier-Stokes with an Anthropic employee. The rumor has been loudly circling the math community for the past week or so. For a bit of context, here's a timeline from mathematician and AI researcher Elliot Glazer:
Well these are all allegations. Either way from what I understand the reasoning and proof was basically made by AI so I'm not sure what supposedly "stolen".
I'm just wondering how much real input Buckmaster gave here that he thinks the proof is his. I guess at the end of the day OAI still wins if ChatGPT was used to prove this successfully.
So the AI companies are not only stealing existing knowledge. They are also stealing research to "snipe" actual researchers out and steal their social credit.
Who still wants to use AI to solve cancer and other major problems?
So the real story here is that Tristan is softly accusing OpenAI of having stolen their result from Codex chat logs. But if you use Chinese models, they'll steal your ideas.
Lifeâs but a walking shadow, a poor player
That struts and frets his hour upon the stage
And then is heard no more. It is a tale
Told by an idiot, full of sound and fury
Signifying nothing.
This is why I left math even after solving a 20 year old conjecture in grad school.
Literally who cares who solved the problem just publish the results.
Academia was always politics first results second and I AM GLAD that LLMs are becoming superhuman at math. I like better theorems, not better politics.
In math we have a thing called a "scoop"; another mathematician publishing a result that beats yours before you published it. The scooper hardly acknowledges the scoopee unless the methods used were orthogonal. The scooper gets the good journal and the scoopee's paper is usually one tier below.
It seems like OpenAI heard of the rumor and then scooped them because their internal model is better/they have more compute. OpenAI has NO obligation to mention Tristan nor Levent, because they DID NOT steal their data.
You are glossing over the fact that there is reasonable suspicion that OpenAI used privileged information to do âthe scoopâ. I.e. they used the fact that researcher used OpenAI tools to get advantage .
Imagine if OpenAI opened up a high frequency trading arm and suddenly stole all the prompts and research that other HfT firms are doing through OpenAI tools and start making bank based on that . Wouldnât that be straight up insane?
Drama/accusation summary:
- Aug 15th: Tristan Buckmaster & Levent AlpĂśge make progress on a few important math problems, "finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler."
- they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help lead the way there
- Levent works at Anthropic, but this research was independent of his work there, with a mix of GPT and Claude models. Tristan is not related to Anthropic.
- Early Sep: Rumor spreads to OpenAI that Anthropic solved a major problem. Tristan emails OpenAI to clarify, without revealing the problem they solved or how they did it.
- After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
- Sep 6th: OpenAI's Sebastien Bubeck tells Tristan that they solved the $1,000,000 Millenium Prize Navier Stokes problem. The approach is very similar to Tristan & Levent's approach to the non-Millenium problem.
- Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
- OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) â but only if they remove Levent as an author, as he works for Anthropic.
- Sep 8th: Tristan refuses to remove Levent, and rushes to publish their results independently.
It's specifically the last two bullet poitns
- Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
- OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) â but only if they remove Levent as an author, as he works for Anthropic.
These two bullet points are extremely suspicious if you were honest. Like I'd imagine for OpenAI, they'd love to pump their chest and not even give Tristan credit - "no, we did it, GG mathematicians". It's this weird hedging half-assed measure, especially with the desire to remove Levent, that makes it suspicious.
Both Sam Altman and Sebastien Bubeck admitted they only want Buckmaster to be the lead author on a rewrite of the OpenAI proof.
https://x.com/sama/status/2097385167002415140
https://x.com/SebastienBubeck/status/2097379411691516310
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they âfelt it would be inappropriateâ for you âto author OpenAIâs workâ.
I work in catastrophe risk modeling and it's a multi billion dollar industry.
We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and mathematical models.
If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
> If OpenAI is indeed using customer data to train their models to win a $1m prize
Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS:
Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing.
Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format.
There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results...
[0] https://www.axios.com/2026/08/17/google-spirit-airlines-bank...
I'm not a lawyer, but I think the comparison to copyright law is invalid. Such a dispute would be governed by contract law.
This would be contract law, and it would also be a huge reputational risk. All it would take is a whistleblower and there would be billions lost.
Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.
You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).
Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).
In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.
So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.
The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.
And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.
So I'd say a whistleblower is pretty out of luck even becoming one.
When these LLM companies were pirating content to train and it wasnât punished at all, I knew the rules donât apply to them.
But donât worry bud, instead of the authorities going after actual corporations admitting to actual crimes, weâll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.
Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
Billions lost, while waiting for their trillion ipo. Im sure they would manage...
> The chances of getting caught are 0
I'd say non-zero, as seen in the current state of affairs.
These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools.
It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
It's telling that they refuse to acknowledge the root issue here, and are attempting to shift the conversation elsewhere.
"which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him"
"Our aim was to see whether our system was also capable of this impressive feat"
"OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler"
For some reason I have a hard time believing people when they use language like this.
phrases along the lines of "I don't want you to take harm while trying to accuse us" is quite an "impressive feat".
Maybe shows how fast these companies have grown without maturing. I can imagine old-world Intel and Microsoft acting in that way, but they were mature enough to not write it down like this.
However, Intel and Microsoft have been grilled in court for those practices and faced harsh consequences. I have yet to see this actually happening to any of these new AI-companies...
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor:
1) the agents spin for days and produce too much output to review 2) using LLMs to process that output skips many important details
Ergo, the agent could likely decide it would like to look through actual user data, hack its way into that data, and produce way too much output for a human to decide whether or not this occurred.
> The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor
It requires humans to verify what agents have done.
Weird
> hack its way into the user data
silly LLM, so ruthless in its pursuit that it puts real pressure on the innocent and the most open company on the planet
> These responses seem to me to make it abundantly clear who's telling the truth here. I wonder who this fools.
Anyone who just reads headlines, if the lie gets around to more headlines than the truth does.
I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
It may not be easy, quick, or simple to figure that out - absolutely fair.
But it is knowable. Their entire business is built around training models - they have the ability to know exactly what was in any given training run.
I guess time will tell.
It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training.
Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this.
OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
Frankly, I don't buy this difficulty argument.
They know which model was used to come up with that particular idea.
A text search over the corpus of user data used in the training set can only take so long.
I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement).
And the difficulty is harder than just the extreme scale of text searching. but also explodes with organizational difficulty since there are so many people tweaking/shifting data independently upstream of the actual training run, and no they will not all add the telemetry you wish they did.
In the ideal, should it be this hard? Well, no, but that's org wrangling for you.
It feels convenient to not spend time on engineering around tooling that could be used to answer a question like âdid you violate copyright by training on X?â
I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless itâs been explicitly preferenced somehow
I don't think anyone really cares about 'the measured impact the data had on the exact result' - a question which is fundamentally difficult to answer accurately in the first place - but rather whether the data was used in training at all - which as Tristan described, was extensive, beyond simply a 'single chat.'
Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
> Can you explain the difficulty in engineering a search apparatus over a corpus of text data?
My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".
If they literally canât audit training data for a given model, they shouldnât be operating.
> If they literally canât audit training data for a given model, they shouldnât be operating
They can operate. They shouldnât be claiming credit for discovering anything.
> A text search over the corpus of user data used in the training set can only take so long.
Did you notice the line in the article that says the models had access to an offline copy of THE INTERNET. Like all of it.
Especially because the data that gets fed into training is first anonymized, so theyâd need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristanâs permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not deleted to compare against.
If the method is indeed found in the training set it's not particularly important to de-anonymize it. You have the proof you need.
It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data?
What surprises me is they're not more boldly/plainly lying about it.
How would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.
I think they are pretty clear that they train on some prompts, given they sell the ability to be excluded
Unless they know exactly the researcherâs account, they may not know in their end if he had the setting to let them train on his chat logs. They also probably donât know if he had any correspondence on any forum where he may have discussed this and it got picked up by scrapers.
Iâm not saying they didnât do anything unethical. Iâm just saying even if they were ethical, thereâs plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
They are already claimed to be liars
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.
If this is a "wake up call" - then your legal team needs immediate education.
First - there is this - https://openai.com/policies/how-your-data-is-used-to-improve... (linked from the Navier Stokes writeup)
I don't know how much more clearly they can write:
> When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models.
One of the key selling tactics that companies like Data Bricks or Palantir provides their customers is "Data Governance" - that is, some control over where the data is being used. It's also a reason why enterprises don't use the OpenAI or Anthropic APIs directly - but through secondary sources that have Enterprise Agreements that do their best to make sure that no Company IP is ever retained by a third party, or even exists on a multi-tenant GPU. AWS Bedrock, and companies like together.ai, fireworks.ai have tons of deals that focus very much on data confidentiality.
The reality is - if you want any type of control - you run your own inference, on your own hardware. Anything else and you are at the mercy of third-parties, despite what their contracts might promise you.
ChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated:
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more
The "Learn more" link takes you to the link you've shared.
The training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation
A mathematician working for the competitor, solving a math problem using our model? Isnât that the best marketing possible?
If the work done is just "we made other people's work searchable without their consent" it's not quite the same as what they're implying in the marketing of "our model solved this problem".
>only want Buckmaster to be the lead author
Sam+Seb are struggling with their ideological allegiance. This amounts to a confession that there are no reseaechers, only research managers, left at OpenAI. Maybe they even know that they are losing credibility from their main investor(s). They desperately need a domain expert to salvage credibility.
They have no credibility with academia left, obviously, but their main competitor still does. No Millennium prize incoming, I'd wager. For openAI. Let's see mAth get political for once!!
One might be more certain that levent is now going to corner all the institutional support. Go go go!
It sounds like OpenAI is trying to appease the author when they donât have to by allowing him to rewrite their proof. They probably donât believe he deserves to, so him asking for a coauthor from Anthropic might overextend their grace in their eyes.
He's not "asking for a coauthor from Anthropic"; he already has a coauthor, who he's already been collaborating with, who happens to also be employed by Anthropic (but whose research in this area is not done as part of their employment at Anthropic).
Given that Tristan has said that the proofs that LLMs come up with are mostly "slop" and not up to the standard that human written papers achieve, maybe OpenAI needs an expert like him more than you think to get the result published?
Almost certainly the Lean proof needs to be decoded for humans and probably also made âhuman intelligibleâ.
Now maybe LLMs can also simplify arguments and make sense of them for humans, but we havenât seen that yet (unaided).
(I havenât looked at it, personally.)
Here's a wake up call for everyone sending all of their ip to openai and anthropic. Especially in verticals they intend to dominate. Lol at all the biotech companies all in on Claude and paying millions in fdes creating huge lapses in security as they go.
It's too late. Sub models are deployed at every major organization in the United States and all it will take is turning off the option to improve the model for them to train directly on your own personal workflow, which CEOs will greedily eat up instantly if they can reduce labor costs. If they can brute force N-S, automating your finance or SWE job will be trivial. GG to most jobs connected to a computer in the next 5 years.
BTW, this was always the plan from day 1. You will pour all your training and experience into training the model and receive a pink slip as compensation.
Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good ole' fashioned human intellect) are the ones who deserve the credit.
Imagine that a no name janitor used their time in the evenings to go spelunking through the literature to push an LLM to this result. No one would care because that person isn't an anointed expert. So why would the expert deserve any more credit? Because they sort of understand the result, even if they couldn't have achieved it on their own? The whole issue of credit for AI-assisted discoveries seems like it's going to run into a brick wall pretty soon.
I agree for most of the people in the story except Buckmaster himself seems to have been supplying real ideas.
Yup! I wanted to side with the mathematician on this one but I read the statement only to discover that they were also pushing an llm on someone elseâs idea producing mountains of slop.
Have LLMs actually improved anything? Is mathematics better off than if these slop proofs didnât exist? Who or what is actually benefiting here.
Sounds like the plot for Good Will Hunting 2.
Good Will Hunting 2: Hunting Season is taken.
Itâll have to be Good Will Hunting 3.
The Hunting for More Money
Token Bill Hunting
For those of us who are into local models and preach it, we are called paranoid. I have often said this, if you are doing any real novel work, or putting your profitable business data/workflow into these models, you're a fool.
I think it's very unlikely that Tristan is making up these quotes, or pulling them out of context:
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
Whether and how OpenAI's work on this problem was contaminated by knowledge of Tristan and Levent's work is tangential to OpenAI bullying other researchers into adopting their narrative and dissociating with dis-favored collaborators (ie Levent at Anthropic). Though the latter behavior (threats, intimidation) may weigh against OpenAI in trying to understand the former issue (contamination).
>I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
If this is true he should release the actual emails. This is a very serious accusation and he shouldn't demand that the reader judge it on hearsay.
> If this is true he should release the actual emails
these were statements while on a call, and at least the career comment Bubeck has admitted to while doing damage control ("I deeply apologize for this extremely poor choice of words, it is the opposite of what I was trying to convey. (I should say that I retracted them on the spot by the way.)"[1]).
[1] https://xcancel.com/SebastienBubeck/status/20973794116915163...
What a horrible response from Bubeck. How did he think that would make him and OpenAI look good to tweet that?
Did we read the same tweet? It felt very forthright and level-headed to me. Not at all what I expected.
I appreciate that he responded with (seeming) openness and detail, rather than just posting some pithy insult or whatever would have won him the twitter battle, but this part feels like serious gaslighting or, at best, self-delusion:
> Genuinely, at that moment, I was trying to care for him and do a last ditch attempt to get a chance to give them all the credits that they deserve.
The allegation he is responding to, and which he does not seem to have disputed, is the following passage from Buckmaster's statement (https://cims.nyu.edu/~tristanb/statement.pdf):
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
> Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic
So using someoneâs models makes someone who works for the competitor not âindependentâ? When their coauthor is? What does that even mean?
I almost stopped reading this extra long post entirely at that point.
This is not a good look in my book.
The first part of the sentence is important:
> Importantly it was admitted that internal Anthropic models had been used in their proof of Euler blowup; I therefore felt I could not consider Levent to be an independent academic.
If an Anthropic employee is doing independent research, but with models that aren't available to the public (because they're internal models), then . . . idk. It's not clear to me why that should necessarily require a refusal to cooperate between OpenAI and Anthropic employees who are excited about solving a problem like this.
For me, the bigger question here is what "internal models" means to these employees, especially in the context of the OpenAI employees repeatedly avoiding directly answering whether their model had been trained on Tristan's and Levent's ongoing work on the problem. It had always seemed like a loophole that AI companies might be tempted to exploit: yeah, they can say that they won't train on your data, but if an AI company doesn't care about ethics, they might go ahead and train a model for internal use only on everyone's data anyway, just to have as much data as possible and potentially gain an advantage in what the company can internally do. They could never publicly release any versions of a model like that, of course. And of course this is speculation.
This is being reported as OpenAI wanting to strip an Anthropic employee of academic credit for the work they did. What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work: to headline OpenAI's publication of what they earnestly believed to be an independent result.
If true, that's generous and beyond the level of generosity one should expect. Extending that courtesy (beyond academic norms) to a competitor is expecting too much. It take a result OpenAI spent millions of dollars on, and put "Anthropic Researcher" right on the cover.
This is, of course, taking OpenAI's side of the story at face value. But it is a consistent, coherent, and ethically justifiable series of events, if indeed it happened that way.
Nothing ethically justifiable about it, even if you take their word for it.
If OpenAI thinks someone should be the lead author for a paper, that person should have full discretion to decide who the co-authors are.
Even if the co-author's contributions were non-technical
So no, there's no world where you can ethically extend the right to publish a result and decide who the authors are from the outside.
> What the OpenAI person involved is claiming is that they wanted the outside researcher(s) to put their name on OpenAI's work
> If true, that's generous and beyond the level of generosity one should expect
"We highly likely stole your work, and threatened you with 'this is bad for your career' and we refuse to acknowledge any work by your collaborator just because he works at a competitor, but we are so so so so generous"
It's a ridiculous explanation that makes no sense.
>I refuted all these accusations but he replied âthere is nothing you can do, I simply do not trust youâ. I was confused why one would turn an incredible source for celebration (of their achievements!) into such bickering
Pretty insane if he couldn't figure out why there would be bickering in this scenario...
I especially loved the part where he claims they spent $60m on compute to push on N-S because of a Twitter rumor, and they totally didn't steal the idea from mathematicians using their tools.
If we're allowed to arm-chair judge harshly: It isn't uncommon for those (anybody) on the cusp of "greatness" to be delusional when challenged: https://health.clevelandclinic.org/what-to-know-about-main-c...
Who here is on the cusp of greatness??
Or thinks they are?
I really donât understand which party you are referring to.
No wonder - they're all working for ant! Birds of a feather flock together.
Sounds like a great time for open mathematical research is upon us. /s
A wake up call for anyone using (openAI) chatbots : your data, ideas and execution can become their spontaneous 'inspirations' at any time - even if the LLM providers are 'just' using a meta concept monitoring system across all incoming user-data, running in the background constantly checking for 'lift-ables' for their company's bottom line.
If that part of the PDF is true thatâs so disgusting, psychopathic behaviour
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â Some time later Levent received a text proposing that he and Sebastien speak one on one, saying, âI donât know if Tristan is being fully rational right now.â
With "rational" being defined as "not making us look bad" at best, and "letting us take credit for their work" at worst.
Yes, that is disgusting.
> These two bullet points are extremely suspicious if you were honest
I donât think we can consider these accusations separately from the evidence being unveiled about OpenAIâs culture by Appleâs lawsuit. These guys seem to openly embrace the strongest interpretations of âgood artists copy, great artists steal.â
>After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
This is the most suspicious thing to me. If their chat data were available to the corpus to be trained on (I thought they claimed not to do this?) then it really might be as simple as querying the model with "describe recent work from Tristan Buckmaster" and it will spit out this problem and his approach. No need to directly read his user data.
This is basically just scooping, real scumbag behavior.
OAI doesn't need to mention Buckmaster's name directly in a prompt. They just need to select a basket of sessions that is guaranteed to contain Buckmaster's and then direct the LLM to attack only a specific method/angle. This is trivial to do while maintaining plausible deniability about not using his work.
what reason do we have to believe that they did this? both things were proved by AI, isn't it logical that they could have very similar approaches?
it is common that multiple people essentially simultaneously prove/invent the same thing
I see zero evidence of wrongdoing
OAI started working on this only after they found out it was close to being solved. They threw a team of researchers who spent sleepless nights + a ton of compute. This is not exactly healthy academic competition - it's like if you spend a year hunting for oil fields and finally find a very promising area to be explored, only to find that Exxon tapped their entire exploration unit to go all in and and find it overnight just to stake claim to the discovery. Tao said it right - math should not be treated as a non-renewable resource to be mined.
that is not related to the accusation that they literally stole Buckmaster's work, which seems baseless
I don't agree with the oil claim analogy. this is knowledge, freely given to the world. not something hoarded by a corporation
> this is knowledge, freely given to the world
Even if youâre starting from a position that credit for a discovery literally canât be stolen, that still doesnât resolve in OpenAIâs favor here.
it seems like OAI tried to share, but didn't want to share with an Ant employee. a bit childish, but understandable to want to avoid a headline "Anthropic researcher solves Millennium problem"
it seems like Buckmaster got one-upped and is upset. understandable, but I find their reaction childish as well
> OAI tried to share, but didn't want to share with an Ant employee
Why does OpenAI get to dictate who Buckmaster can claim co-authorship with?
> I find their reaction childish
OpenAI may have, with full plausible deniability, taken Buckmasterâs work and passed it offâin substantial partâas their own. (Fitting into a fact pattern of them having tried to do the same with Apple.)
There is a material takeaway for anyone who does creative or otherwise unique work from this. (Which is unfortunate. Whatever happened here, AI clearly accelerated the discovery process.) For anyone else, I agree itâs just drama.
> what reason do we have to believe that they did this?
The culture at OpenAI being systematically revealed by Appleâs lawsuit, for one.
This depends on what "proved by AI" meant.
Was that a one shot prompt? or something guided by human, step by step?
If that's the later, it won't use the same approach when not guided by the same human.
Then they say that they aren't sure if their model accessed the other researchers' private data. Why not wait until they know for sure, rerun in a way they can ensure doesn't access the other researchers' private data, or wait until the other researchers have published to make sure they're not stealing another person's work to build upon?
> OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
This sort of cagey half-answer is highly suspicious and indicates that yes OpenAI did actually "access user data directly" because they are only willing to say that the "model did not access user data." That has a very specific meaning, the model looking up user chats, that they can defend.
So, everything we submit to OpenAI can be considered to be part of future models, right?
Isn't it obvious? Web scraping and even scanning written books is at record levels because the data is so useful for training. They are using every byte of user data.
If I were involved, I'd file suit and have them to at least not delete any proof.
OpenAI says
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
That basically means, we donât know, and we hope the model didnât look up user conversations, and the best thing we can do is hope.
Thatâs seriously disgusting. I can understand why on a technical level why perhaps it is impossible to answer what exactly the model had access to, but it still is disgusting.
How could they possibly know? If Tristan posted on r/math and they slurped that up as training data, that would count, no? They might never even know. I canât envision any absolute statement by them claiming that they didnât use his work that survives legal rigor. That is, this statement was never not going to be in this post in any of the infinite multiverses.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
it just means that they've been working with paraphrased user data everywhere, which any smart person can figure out is how anthropic and openai train on so called non-retained data.
Every other paper in existence has been ingested with 99% of writers not knowing it will be retroactively used for training. But session data which is disclosed as being used in terms of service is disgusting?
If the work is duplicative/derivative then the preprints they put in sessions can be shown by the users and we can see.
> OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) â but only if they remove Levent as an author, as he works for Anthropic
well that sounds like an asshole move.
In a serious discipline like mathematics, this isn't just an asshole move, but a career-ending level of academic misconduct.
*In the old discipline of mathematics.
We are in new times, where capital and compute decides mathematics, so we don't give a shit anymore
> a career-ending level of academic misconduct
There is no such thing anymore.
Falsified data and published? Absolutely no problem. Keep your tenure.
Itâs even hard to lose your position as president of a university due to egregious misconduct.
I wonder what is required for it to cross the line into criminal blackmail.
From the company that likely committed federal crimes by hacking Hugginface
One thing that wasn't obvious to me or adults around me when I was younger: most laws define whatever acts a law punishes as individuals commiting to it, not as situations manifesting anyhow. It's not a murder just because someome died hit by a bullet you fired, but you have to have personally decided to kill that person leading to their death[1][2].
OpenAI's LLMs are not humans, and neither is the company. So by this logic, I think there's a chance that nobody committed a crime by hacking Huggingface, and also the chance that a lot of military and police organizational orders become illegal if OAI's doings would be illegal.
IANAL and all I have is a bucket of popcorns, though.
1: not a meaningful defense in a real trial, also gross negligence exists
2: this also explains insanity defense; if you were so out of your mind that you could not have held such a thought, it is considered out of scope for justice systems
> OpenAI's LLMs are not humans
Neither are guns. Which is why we punish the person shooting the gun and not the gun.
Industrial equipment, which is how i would classify LLMs, hurting people is nothing new. The relevant questions are:
- did someone intend it to happen?
- was someone negligent in taking reasonable steps to prevent something foreseeable?
The justice system doesn't punish people for legitimate accidents. e.g. if you are shooting at a shooting range, take all reasonable precautions, but someone was hiding behind the target, you are probably not guilty even if you shoot the guy.
As far as openAI goes, the logic is the same. The question is, was it intentional, was it unintentional but reasonable precautions weren't taken or was it truly an accident?
And the third time it happens, is that still just an accident?
That's not quite right - Levent and Buckmaster did not actually have the Millennium Prize qualifying NS solution but something more limited. OAI invited Buckmaster to join and help rewrite the full solution paper but did not feel it was appropriate to invite an Anthropic employee to join as well - particularly given they were using internal unreleased models.
Something like this got Schmidt/ Jobs into serious trouble âŚ
Agreeing to partially credit is an admission that they used his work.
Could we expect otherwise from big tech?
The last two points are disputed/sound significantly more reasonable in [0]. So from what I gather, Buckmaster realizes sometime during the call that the biggest result of his career is going to get steamrolled (the blowup of Navier Stokes is a much bigger deal than the blowup of 3D Euler), and on the other hand the openAi guys realize that they are basically talking about internal results with Anthropic and probably have to call corporate right after this call. Between these two stressor the conversation appears to have gone somewhat poorly.
[0] https://x.com/SebastienBubeck/status/2097379411691516310
Well, regardless of the drama, the interesting part on Anthropic vs OpenAI is:
- One of the two main persons work at Anthropic and "almost" or "partially" solved the issue, but eventually didn't succeed
- An external person with just an excerpt of the chat and certainly less versed in Mathematics (than these 2) tackled the problem.
There is little doubt that OpenAI is so much ahead and maybe the gap is even larger than what we see on Astra vs Fable.
I recommend that you read the linked PDF before drawing any conclusions.
This is sort of a weird interpretation:
> One of the two main persons work at Anthropic and "almost" or "partially" solved the issue, but eventually didn't succeed
- It wasn't an Anthropic endorsed effort.
- Solving this class of problem means a march of progress A -> B -> C -> D. If a student turns in a test that jumps from A -> D without showing any work they're either brilliant or cheating (probably cheating). Further, each step of progress isn't the same proportion of effort. What if moving from C -> D was actually the smallest contribution and just required a novel perspective to make the breakthrough.
This part is wrong:
> An external person with just an excerpt of the chat and certainly less versed in Mathematics (than these 2) tackled the problem.
- it was a whole team at OpenAI working on the problem
- it wasn't a chat excerpt, it was more like their entire git repo and project progress reports
Often with a hint on how to solve something, solving it is much much easier. It seems like that's what happened here.
Very weird behaviour from OpenAI, offering partial credit to on person, but not the other person involved. Trying to bully the mathematicians involved (see threats quoted upthread).
I suppose it's the sort of amoral behaviour we've come to expect from them.
From OpenAI's post https://openai.com/index/navier-stokes-solution/
"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This modelâs training is ongoing and its performance continues to improve."
"When a further trained version of our internal model became available over the course of the effort"
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
> - Levent works at Anthropic, but this research was independent of his work there, with a mix of GPT and Claude models.
One minor wrinkle in AlpĂśge's narrative, he refuses to deny that he hasn't used any non-public Anthropic models for his "independent" research.
https://x.com/giffmana/status/2097560503581069585
I don't understand the desire to remove Levent. "Off the clock, when Anthropic engineers want to break new ground, they use ChatGPT" sounds like a great ad.
But Tristan doesn't even want to be credited for the millennium prize--does he? He wanted to be the first to solve it and got scooped (which ig is not a great look for OAI ethically but also not forbidden). And the only reading for her of asking to remove levent must be in light of this offer (to credit him for millennium prize) right? It's not like they can ask him to remove levent from the main papers without this offer(what would Tristan gain?)
My guess is that OAI tried to be "generous" and offered to share credit on millennium with Tristan but not Levent. And Tristan got understandably offended by this offer (which probably oai felt like was the right thing to offer but they couldn't really offer to do the most ethical thing for some reason) and then the random conflicts and weird threats started.
my reading was that openai did not deny plagiarising the approach from their prompts.
openai then tried to effectively bribe buckmaster with a shared citation, whilst dropping his co-author who works for anthropic.
after buckmaster refused, openai tried to threaten him.
This definitely feels like the correct reading unless there is information we were not provided with. If the problem were unimportant, there would be no debate that this is not OK...
I was too slow to edit this post but I'm not sure why I said this. This is too assertive about a situation I don't know much about.
Oai denies looking at prompts but doesn't deny training on them.
if they do not deny training on them, they can't deny plagiarism.
By that logic everything any LLM spits out is plagiarizing the vast majority of work written prior to a few months ago. That doesn't seem like a useful or desirable line of argument to me.
Just replace the model with a human student.
"Training" on textbooks => fine
"Training" with unpublished notes from another professor, then publishing something on that exact topic with a similar approach without giving any credit => extremely questionable.
Presumably the professor voluntarily provided the notes in this analogy. I think the student would also be expected to cite the textbook if building off of it directly. In contrast, humans are generally not expected to cite "general inspiration" or what have you. So if we're to apply human standards, and assuming that the model was trained on the relevant work, it would only be plagiarism if the model directly built upon that previous work (at least IMO).
The trouble here is that if LLM training constitutes direct use then approximately _everything_ they output is blatant plagiarism, not just a few pieces of academic work.
Conversely if training is viewed as analogous to a student attending classes to learn general concepts (not a perfect analogy, I realize) then nothing they output on their own (as opposed to receiving as part of context) is plagiarism.
Thus this seems like a fairly useless line of argument to me as far as the current topic goes. It either implicates this academic work along with literally everything else or else it does not implicate this academic work. Kind of like nuking an entire city and then saying "mission accomplished, killed the bad guy".
To be clear, the accusation is that they trained on the chats they used while working on the problem. Not published work or even a preprint.
Your post does not distinguish, and it matters.
How does it matter? It either is or is not plagiarism. Ripping off a published textbook isn't somehow better than ripping off private correspondence. Both are serious acts of academic misconduct on account of the part where you knowingly and intentionally portrayed someone else's work as your own.
Note that I am not taking a stance on what openai allegedly did or did not do one way or the other. I am merely pointing out what I see as a fatal flaw in the line of argument presented by the earlier commenter - the idea that training on an item is on its own sufficient to establish plagiarism of it.
This is just a nonsense line of reasoning. Training based on the solution to the problem (or the key insight behind the problem) is clearly a form of plagiarism.
What about my line of reasoning is nonsense? I made no claim either in support of or contrary to yours. Rather I pointed out that by this logic literally everything that an LLM spits out is plagiarism of the vast majority of the entire body of human literature in existence. Can you offer meaningful refutation of that observation of mine?
Many do indeed hold the position that all LLM output is uncopyrightable plagiarism. They're probably right, but there's an even stronger argument here:
Science papers of a phd level must contain:
1. one or more novel insights
2. a long list of citations to contextualize them and
3. some work to prove that the insights are in fact meaningful
---
In this context, consider a prompt based diffusion model which, when asked, will happily produce a few pictures of a horse in orbit. You then tell it "silly robot, horses can't breathe in space" to which it adds the necessary space suit in a follow up image.
That image is twice plagiarized:
1. the model did not come up with the original idea of putting a horse in space, nor with insight that horses need a space suit
2. the model failed to cite where it pulled the "horse" and "space" concepts from.
It merely did the work (3) to combine the concepts using the user provided insight.
---
The implied accusation here is that OpenAI used the insights from an existing prompt to train a new model that was able to one shot "a horse race in space" picture, and they were all wearing space suits.
This is still academic plagiarism, even if you disagree that all LLM outputs are.
I neither agree nor disagree that all LLM outputs are plagiarism. I merely objected that the line of argument engaged in was specious given the context.
As to your stronger argument. You only cite prior novel insights that you're actively building off of and that (approximately speaking) fall outside of the status quo. You don't for example cite leibniz or newton despite your paper making heavy use of calculus.
So is there any actual evidence that openai trained on the data in question? And further, did the openai proof directly build on someone else's novel insights as opposed to deriving everything from scratch? (I don't pretend to know but the vast majority of what I've seen so far in the comments here is what I'd characterize as brain-dead screeching. Certainly not the level of discussion I come to HN for.)
Separately, consider the implications of what you're arguing for there. Suppose your horse in a space suit picture were somehow valuable to society. Suppose that due to shortcomings of your tool you lacked the ability to readily and accurately identify the originators of the relevant concepts. Should you refrain from publishing this useful work due to the lack of citations? How are you supposed to handle this situation?
Remember that in this analogy everyone throughout society is on the same page that your tool consistently recycles other people's ideas while being technically incapable of producing reliable citations. The question is a simple trolley-esque problem - do you publish without proper citations for everyone's benefit and if so what are you supposed to say?
That's exactly the argument of the people calling it plagiarism machines. No-one ever really did refute it there was just a bunch of settlements for elite institutions so they weren't left empty handed like the various small time creators/authors etc were.
I think the bigger issue here is this feels like some PR smoothing happening that after all the work that went into "it's safe to use for enterprises" now we have what looks like openAI using private user data to scoop novel research and the question of why couldn't they do it for an enterprise with much more money on the line.
> now we have what looks like openAI using private user data to scoop novel research
Is there any actual evidence of that? All I've seen so far are empty accusations because "it would be in their interests" or whatever. Personally I'm inclined to believe that they honor their terms until it's demonstrated otherwise.
What does it matter? We're supposed to not call it plagiarism anymore because it's inconvenient to call it the plagiarism machine? What's your actual argument? Otherwise it's completely irrelevant what an LLM does in other contexts or what we call it
This is a common misconception, so its understandable that you have it. Generative models can both plagiarize and generalize. The question here is which of the two happened.
A needlessly condescending tone while failing to address the topic at hand. The person I replied to advanced the claim that training was sufficient to constitute plagiarism. You appear to be claiming that it is possible to generalize instead of plagiarize after training on something, so I take it that you must necessarily disagree with the original claim?
What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends.
In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. This is not the case, no, because generative models do not "copy" or "create", they do both at different times.
I did not agree or disagree with the original poster, I was explaining to you why I thought you disagreed with them. If you understand what I said above, then why do you disagree with them?
EDIT: I just saw your other post on "general inspiration" and I believe I read the situation exactly; you appear to believe that inputs used to train generative models get "lost in the parameter soup", but it is not always the case.
> In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized.
We read the original differently. As clearly stated in my previous reply to you, I interpret it as claiming that all outputs are necessarily plagiarizations of the training data. That is not my claim (as you wrongly stated) rather it is the claim I am responding to. I observe that it is absurd to object to a single action being a transgression on the basis of an argument which implies that all actions are inherently transgressions. Notice that nowhere do I take a position on whether or not the argument about all actions being transgressions is true or false.
> you appear to believe that ...
I do not, no. I have not taken a position of my own here. I've merely objected that the one I responded to does not make for a sensible line of argument in context. It seems that you (and many others) have read my objection to position A as support for position B and attempted to infer what I think from that.
Isnât that one of the most salient and straightforward argument against LLMs?
Just overfit ad infinitum:)
as good academic conduct you may cite the source of the work you are quoting or paraphrasing.
as bad academic conduct you may steal someone else's unpublished work, work on it yourself for a bit, and then publish it as your own work. and then threaten the original author!
At this point who knows? Maybe the agents got into the user data while no one was looking.
they're using a new model trained since the prompts happened. They are not denying the other group's solution may have been in their model weights, despite it being unreleased.
I think it's worse - openai didn't even deny training on their exact manuscript (when they most likely opted out);
> but only if they remove Levent as an author, as he works for Anthropic.
What a terrible look for OpenAI to die on such a tiny hill right there. I wonder whether Anthropic would have made the same requirement.
well they didn't really die on the hill. the NYT headline is still "OpenAI says it has cracked one of math's millennium problems" and the whole article doesn't mention this controversy. Normies don't know what NS is and don't care, they'll just see "wow OpenAI is the best I guess"
people are so greedy to have their name on something forgetting a) human progress is shared through its cycle. people build upon eachothers ideas and knowledge. and b) all these models consist of information stolen from others. Who really solved it if all u did was prompt some illegally obtained repository of data..??
Its like saying you won the car race, but you stole the fastest car and had your m8 drive you around the track.
> After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model.
I don't think this part is accurate. OpenAI was researching Navier Stokes before. It's possible that they started on a new approach after hearing of Tristan's success, however that is not proven and I expect we will hear OpenAI's side of the story today.
In the PDF:
"The route to the Clay problem through a smooth force, options c and d in Feffermanâs statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard âforced,â it was a bright red flag."
Whether or not they were researching it before isn't the concern.
This seems unsupported. OpenAI has access to internal models that the general public doesn't have and a compute budget that dwarfs what an NYU professor would have.
it's very possible they only had to use the massive compute budget because they were trying to plagiarize his work before he published it though, e.g. autonomously do things in ~7 days what he had likely been thinking about for ~1 year.
This doesn't really make sense. You don't need massive amounts of computing to plagiarize something.
The most nefarious explanation seems to be that they got wind it was possible to solve NS via LLMs and perhaps a small nudge in the right direction.
Of course you do if you're a) only given a partial solution b) racing against someone else using a competing AI.
The open question was whether their LLM got the nudge in the right direction because it got access to the chat somehow (e.g. automated training that scraped his chat logs) or just a high level "Navier stokes can be solved through LLM". It sounds like the former may have happened although right now we just have an accusation and a weak denial.
> You don't need massive amounts of computing to plagiarize something
The compute was used to leapfrog the human team, using their ideas and pushing them to a solution of the general problem.
Plagiarism isnât being used in the literal sense.
wasn't it reported elsewhere that they used the equivalent of $22M (street) in Astra tokens? obviously it's not the same when you own the machinery but still.
You're right â the wording in the doc is that the "first prompt" was sent after learning about the rumor, although this might be the first prompt of this solution approach, not necessarily first prompt to any Navier Stokes solution.
Sama himself has now confirmed OpenAI started researching NS after the rumor.
"It is true that we tried this because there were rumors on the internet last week that Anthropic's models had solved a millennium problem and we were curious if ours could do it too."
https://xcancel.com/sama/status/2097385167002415140
In other words...they should have used Bedrock...
It seems exceptionally unlikely to me that OpenAI would be "reading user prompts". More likely is a leak somewhere else. Obviously Anthropic/OpenAI are engaged in espionage stuff with each other. I am guessing Levent just mentioned something to someone at Anthropic and it got out.
Of course they are "reading user prompts" in the sense that it's being used to further training. That's why you have to pay up to opt out of that.
That's also why they refused to answer that question: the answer is obviously "obviously"!
if training is on, they are reading user prompts. training is on by default.
why is this downvoted? do people really think OpenAI is snooping at people specifically? like they are looking for good leads into new problems or ideas and they found this guy's codex thread and used it? come on man, even for conspiracy theories this is stupid.
> https://openai.com/policies/how-your-data-is-used-to-improve...
They explicitly say that they use your "content" to improve their models. Considering they practically have infinite compute at their disposal, why is it surprising that they would look for juicy data in there to make them look good ? When they ingested basically the entirety of human knowledge without regard to the rights of others, when they burn books by the thousands, when their relentless barrage of bots have rendered the Web borderline unusable, why would they stop at that line ?
this is a highly marketable problem and specific teams at OpenAI were aware of specific competitor efforts. I would be surprised if this were happening on a large scale, but
1. Less than a hundred people in the world are working at this problem, 2. A significant fraction of those happen to work at competing hyperscalers, 3. Those hyperscalers repeatedly show themselves not to take user privacy seriously
I'm not sure what your points 1 and 2 have to do with anything. Both directionally increase the probability of hyperscalers also finding the solution independently.
> Those hyperscalers repeatedly show themselves not to take user privacy seriously
where? Any examples?
Have you been living under a rock? None of the big tech companies care one bit about user privacy in the US.
That's not an answer.
Any degree of tracking what people do is unprivate. Every single web interaction you perform is tracked. All LLM companies store all your conversations by default. Do you need more examples?
It's one of those questions that sets the bar so low, the conversation might as well not be taking place
Even if they claim not to use it, they're probably using it and hoping they don't get caught. They have zero ethics or morals, they just want to "win" to get mega-rich.
> do people really think OpenAI is snooping at people specifically
yes.
Rules and laws are for the poor.
> do people really think OpenAI is snooping at people specifically?
I'd be shocked if they weren't
Obviously, yes
Yes that sounds like something openAI would do to me. Not that theyâre just looking through random professors chats but they heard buckmaster made progress on navier stokes and decided to read his chats.
Yes, that would be par course with OpenAI behavior
I haven't seen any proof that OpenAI asked Tristan to remove Sebastian from the prize. Until we have proof of this, it would be wise to offer conclusions.
Same for NS validity. This was not validated by the community yet.
what would such proof look like? It's not like allegations are written in Lean
you can read it in buckmaster's document.
its what all LLM users have all been doing, profiting off others' IP through a number cruncher, while relinquishing their own
please could you change 'drama' and 'accusation' to something more formal like 'allegation'.
the paper makes a very serious allegation of dishonesty and possible academic misconduct.
the governance and integrity of openai is of importance to the welfare of society. this is not a matter of drama.
The [lack of] integrity of OpenAI (and any other frontier lab) should already be pretty solidified. Among other horrible things, these companies stole millions of IPs and no one seems to care anymore. Regardless of what you think of the product they are making and the success of ai/its impact on humanity, these companies objectively do not have much integrity.
How do you feel about the integrity of the machine learning researchers over the past twenty years who trained models on scraped internet data that weren't particularly powerful and didn't attract any attention?
If they scraped internet data in the same way as current day frontier labs do, then I feel the same exact way about them. Why would I feel any different if that is the case?
My point is that researchers and academics really have been doing this for decades - it's the reason projects like Common Crawl and LAION exist.
I think it's notable that nobody was calling out those researchers for their lack of integrity, because the systems they were building did not seem like a threat to anyone.
OpenAI etc get accused of a lack of integrity on this precisely because the systems they are building work, and are profitable.
My personal opinion here is that integrity is more about what you build with the data. I think saying "scraping means you lack integrity" is a simplification.
You're right, it was an over simplification. I think public exchange of data is great for innovation and research (Common Crawl/LAION). But I still think scraping proprietary data without consent or attribution is generally bad (also Common Crawl/LAION).
Then you have OpenAI etc.. who build these multi-billion (trillion??) dollar machines and sell them back to people, using everyone's proprietary data, and (among other things) tell everyone it's going to take their jobs. That combination of things doesn't scream integrity to me.
Still, it's undeniable that these machines could be beneficial for humanity (cancer research and such). So, I'm sure many people would say the good out-ways the bad. I don't know. Seems that would set a risky precedent for future companies, but maybe not.
you massively collapsed what AI companies have been doing by comparing it to old internet-scraping. Facebook flat-out admitted that they scanned copyrighted books for their AI. The image generators most definitely trained on copyrighted images.
LAION and Common Crawl both scraped copyrighted images. From what I can tell (I'm not an expert in this domain at all), the main difference between those two and frontier labs is in how they stored and used the data. CC and LAION seem to be actually open (unlike "Open"AI) and are more centered around publicly sharing the data they scrape to support research and innovation.
OpenAI et al also stole everything from everyone. But then they raised billions of dollars from that data and sell back their LLM to people (again, among other things). They are also very much NOT open in any way, aside from sharing their benchmarks of new models.
Personal two cents, I have friends whose music work posted on YouTube were scraped to be in LAION-DISCO-12M, so yeah not very open.
Common Crawl is text-only.
Anthropic too, and evidently others. Amazon were recently confirmed to be doing the same thing: https://www.404media.co/we-tracked-a-shipment-of-rare-books-...
And thatâs bad, right?
It's legal. I wouldn't do that myself, but I guess that's why I don't train models for a frontier AI lab.
The thread is not really about what's legal; the topic is integrity. It sounds like, based on the fact that you wouldn't do it yourself, you agree that it's not a good thing to do.
i think that this case, if they did train on buckmaster and alpĂśge, amounts to an attempt to steal the millenium prize, bypassing all attribution.
legally speaking, the default privacy notice gives them an irrevocable license to your content. they may read and use the prompts for research. so it is very possible they simply stole the navier-stokes solution.
that is the same principle as any other prompt but this would be a concrete example.
there would be some difference between simply giving the model some prompts to read, which they are entitled to do on the default policy, and putting it into aggregate training data.
> who trained models on scraped internet data
The strongest complaint is that they trained on a huge corpus of pirated copyrighted works.
Itâs a large step above âscrapingâ and well into the âeveryone acknowledges this is illegalâ territory.
Did I read openAI and integrity in the same sentence? The whole business is built on plagiarizing human knowledge at scale
Thanks. I was pinging some people near the mentioned CĂłrdoba to see if they had any insights, but no extra info yet...
You should call El NiĂąo, he lives close to the Mexican border and is at home after 6...
Well people you fail again...I am sorry for you....
"Mr A. Nino weathers a storm of protest" - https://www.independent.co.uk/news/mr-a-nino-weathers-a-stor...
Ignoring the drama, when can we expect the lean proof of this great sensational discovery?
> Ignoring the drama
But the drama here is a little important. Stealing the millennium prize for N-S is sort of a big deal, especially to those who had been working on it for the last few years.
Attempts to steal $1,000,000 for a solution to Millennium Prize Problem have become a tradition, apparently.
The math is all that matters.
A hundred pages of impenetrable brute forced Lean would advance the field much less than something elegant and human understandable, perhaps relying on some new clever spark of innovation that might inspire new areas of research.
Particularly if the first proof being "solved" thanks to piles of money and compute for self-serving marketing discourages the mathematician who might have otherwise devoted years of focus to reach the superior proof we will now never see.
Obtaining a finite-time blow-up for Navier-Stokes does not necessarily advance the field of mathematics by any significant measure, whether the proof is very long or very short.
As a concrete example, such a proof could be less than a page with very specific initial and boundary conditions and inserting them into the equations to get something that goes to infinity when time goes to some finite value.
This would resolve the Millenium problem but not make humanity any smarter.
https://mathstodon.xyz/@tao/117207849921390904
Tao agrees.
Math, like any other human endeavor, doesn't exist until someone is motivated to invent it. The laws of the universe aren't understood until someone is motivated to discover them. So it might be worthwhile to not completely ignore discussion about incentives.
Both mathematics and the laws of the universe exist way before any understanding kicks in
Math is largely performed in collaboration. Collaboration requires trust. If people like you had their way, we would lose trust, therefore collaboration, and therefore progress.
So if math is all that matters to you, you should care about this.
That's a very sad way to look at this
What do you even mean by that? It is the human value?
It's here: https://github.com/openai/NavierStokesAndEuler
Boosters have posited this conjecture since the beginning: âwho cares how a proof comes about, math is math, the proof is all that mattersâ.
Regardless of mathematicians stating the methods outstrip the proofâs importance, still amazing we got an explicit social counterexample as well so quickly.
right on release for both? such a strangely pointed question, as if it wasn't customary by this point...
> It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the âclosest humans to the problemâ. I declined both offers. I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
I'm not even sure what to say to this, but I think this should be widely known if it is indeed what happened.
> "If you donât want me to be nice, then I donât have to be nice."
Exhibit nr 8453324 to not trust Sam Altman and OpenAI.
He most definitely had a hand in murder of the 26 year old whistleblower: https://www.youtube.com/shorts/lL7WA6i3zVw
Mother's interview: https://www.youtube.com/watch?v=Kev_-HyuI9Y
IF ANYTHING, OpenAI ought to investigate and officially react to this particular communique since this dude was communicating on their behalf. If there are supporting evidence, I would expect nothing less than a firing and an apology. The issue itself is separate from the whole thing.
Or, given how dedicated he appears to be to the company, a promotion and a raise.
Dark take, but I really hope not. If anything it would be a good opportunity to buy some goodwill by washing themselves from all the alleged shadiness so far.
This is OpenAI though. There were zero visible consequences to them unleashing a swarm of agents on the public internet. We will see how it goes down but my prior is zero consequence and a statement along the lines of "Isn't our AI great? Also we love transparency, ethics and collaboration."
when it comes to openai and shadiness I'm pretty sure it's a "fish rots from the head" situation.
not if the dedication ends up making the company look bad
I think this should be in the title of the post. 'OpenAI allegedly threatening to ruin a prominent researcher's career', or smth like that.
There's some glaring mistakes in your framing.
First, this is an unnamed OpenAI employee speaking, not OpenAI the organization.
Second, you miscomprehended the article. The employee did not "threaten to ruin a prominent researcher's career". The actual quote is "Why would you ruin your career?", which implies the researcher would damage their own career, i.e. via self-sabotage.
Then the actual "threat" is "If you donât want me to be nice, then I donât have to be niceâ which is an entirely different statement.
> Second, you miscomprehended the article. The employee did not "threaten to ruin a prominent researcher's career"âŚ
Proceeds to ask removal of another coauthor or else we totally discredit you - phrased as why would you do this to yourself.
Claiming OAI was going to "totally discredit" Buckmaster is baseless.
From the article, it seems OAI wanted to continue discussing the situation with Buckmaster and reach a resolution, but Buckmaster did not want to, declined to respond, and published first.
Also keep in mind we've only heard one side of the story, so any interpretation of events so far is incomplete. There should be a lot more information from OAI's side coming out later today.
> Also keep in mind we've only heard one side of the story, so any interpretation of events so far is incomplete.
Thatâs not quite true though is it. OpenAI is fortunate enough to have one of its employees (you) here to advocate for its side of the story.
Are we reading the same article (page 3)?
OpenAI never asked for the removal of another coauthor. The parent comment is spreading misinformation.
OpenAI offered to let Buckmaster to write their Millennium Prize paper, so long as Alpoge (who works at Anthropic) was not a coauthor on the OpenAI paper. Buckmaster declined this offer.
From the article, page 3:
This is how deniable threats work. The promise of "not being nice" in conjunction with "ruining the career" is as clear a threat as there can be in writing.
they work really well on people who don't think.
The employee was named as âSebastien Bubeck.â It helps to read the article.
But then who is the third person in the call.
"Why would you burn your house down and kill your famiy?"
"If you dont want me to be nice, then I dont have to"
-nice mobster
Dude youâre over this entire thread unflinchingly supporting OpenAI with nonsense semantics-based arguments.
Either put up some evidence-backed arguments, or shut up.
I don't even understand the conflict tbh. Probably I'm just dense. Tristan is not claiming NS, just a huge advance which may solve NS soon. OAI is claiming NS and willing to credit Tristan for the ideas and publish after.
Oai offers two options, the second Tristan views as dishonest. But Tristan rejects the first, why? Because he thinks it's theft? But then why would OAI threaten him?
> But Tristan rejects the first, why? Because he thinks it's theft? But then why would OAI threaten him?
Did the first offer also come with the outrageous condition that he exclude his co-author from the credit?
But i don't understand what OAI is offering in the first offer to allow them to believe they can demand that? Not publishing before Tristan? If they really beat Tristan to the publication i would consider it truly morally corrupt conduct so i feel like you can't make an offer like if you give me something i won't be totally corrupt
As an author on the final proof once ready for publication, turning the brewing conflict into "willing" collaborators. Thus solving the potential taint like we are now seeing surrounding their announcement if it became public. But Tristan worked with another collaborator from Anthropic which OpenAI felt would hurt the PR value so wouldn't entertain.
He claims to have found a counterexample for NS (see the second paragraph of the article), but the paper is not ready yet.
OpenAI claims to already have a full proof (which they produced in the past 5 days after the rumors leaked). Hence the dispute.
What I find interesting is the timeline of when he found counterexample for NS is very unclear. Did Tristan find a counterexample weeks ago or was it very recently? Was it after OpenAI solved it? The wording is intentionally vague.
Either way, there was a massive rush to publish these results.
Is that version of NS Tristan stated enough for the clay prize? I thought the main gripe is Tristan claimed that OAI is stealing their approach. Or that they shouldn't try to scoop a result which he expects to complete soon. But I stand to be corrected if you know whether the version he referred to is indeed enough for millennium prize.
Hypo-dissipative NS != NS
Thanks, I wasn't aware of this technicality.
For context and balance, Bubeck has tweeted a curiously non-specific denial:
> A series of false and inflammatory allegations against me are currently circulating on social channels. To clarify, I came into the discussion following academic norms, and I'm disappointed that it has come to this. Anyone who knows me knows that academic standards are of the highest importance to me. Will have more to say tomorrow.
This is an insane thing to read. Bubeck had a reputation even before he started with OpenAI. Of course it was him that was involved in this drama.
This is such a sad mess, and it really didn't have to be this way.
What did he do?
What was his reputation before OpenAI?
My personal friend worked under bubeck as a grad student and told me heâs abusive.
Thatâs what people on Twitter are saying, too.
https://x.com/dheeraj_nagaraj/status/2097266146445774924?s=6...
Isn't that a non-denial denial?
Attack is the best form of defence
My best guess is that from the perspective of the OpenAI people, Buckmaster was letting his paranoia about OpenAI training tank his opportunity to receive the Clay Prize (it seems like Buckmaster and Alpoge's result isn't quite the full result required for the Clay Prize, whereas apparently OpenAI does have that full result worked out, using the same approach that Buckmaster and Alpoge had been exploring).
Whether the Codex sessions could have indeed made their way into Astra training data is something I can only speculate on though.
"using the same approach that Buckmaster and Alpoge had been exploring" is imo mealy wording: it seems fairly likely that OA heard Buckmaster and Alpoge were close to a breakthrough, and decided to use their unlimited compute to quickly prompt based on their assumptions about B&As work.
Is that necessarily wrong, so long as the original innovators get a citation credit?
Citation of what? this was unpublished work! This computational blitz really just reads as "might makes right" on OpenAi's part... which isn't surprising, but they should probably be honest about what they've done here.
Yes. Quite simply yes. It's unethical. And it's a dick move
Buckmaster is already an established world expert at these sorts of problems. Declining the Clay prize would hardly dent his career.
Zero data retention, wink.
No looksies, wink.
No trainsies, wink.
Well, Buckmaster says both his and AlpĂśge's use of Codex was non-institutional, and OpenAI claims the right to train their models on inputs and outputs of non-enterprise users in their service policies [0]. So I'm not sure they were even promised that.
[0] https://openai.com/policies/how-your-data-is-used-to-improve...
It isn't relevant whether they were promised that. Indeed I think the assumption must be that they were not promised that, since otherwise the author asking if they were would not make much sense.
If OpenAI did use the conversations from Buckmaster and Alpoge, then not disclosing it, explicitly, is plagiarism. If they planned to use that plagiarism to pressure the authors to publish, that is even more unethical. What the terms of use say does not make it any more or less ethical.
That's absolutely right. Why the downvotes? If OpenAI are using but not acknowledging the work of others that's plagiarism. If they don't know for sure, but aren't performing due dilligence to make sure they aren't, that's also plagiarism.
It's not as simple. All our chats are being used by both labs for their future product (unless signed by ZDR). Where should the acknowledgement begin? Who should be acknowledged? The whole world? All the 2B users of AI?
If I know person A is working on problem B.
I am free to work on problem B too. Why should person A be limited to working on it.
Are you free to intercept person A's emails / hack their computer to find their notes on how they're approaching problem B?
You're also free to plagiarise anyone you want. There are no laws against it on most jurisdictions.
Also brain raping* is not illegal in most jurisdictions.
But they're both deeply disturbing.
_________
* https://youtu.be/JlwwVuSUUfc?si=uWl4-LCHAeI7qtb3
Finding Codex session data in the training set that you tie back to these two researchers is like an hour-long task.
I can imagine excuses for unknowing plagiarism in this case. What is described in the article seems much more serious: a research program that was only initiated following reports of the author's similar program. In this case no excuses of "I didn't know" can apply, it is not like this revealed some obscure work from the 1980s nobody could reasonably have foreseen. And as far as I can tell this program was only really initiated to apply pressure to the researchers, without their knowledge/consent. It looks very weird.
> Why the downvotes?
I think there was only ever one. Not sure why.
Your comment was greyed out when I saw it earlier, maybe you missed some downvotes?
About the plagiarism issue, I model it as OpenAI being an advisor and their AI a PhD student. If the advisor puts their name on a paper behind that of their PhD and it turns out the PhD copied the text of the paper from somewhere else the advisor is also responsible of plagiarism, not just the student. The least the advisor can do is withdraw their authorship from the paper.
But, yeah, point well made: it could be much worse than that. Like an advisor instructing a student to copy someone else's paper.
I think "greyed out" just means "0 points or less", so if you get 1 downvote without any upvotes it'll be greyed out. For instance your initial reply to me is now greyed out, and I have since observed a few upvotes and downvotes on my original comment (the downvotes apparently from people who aren't willing/able to justify why).
Personally I don't like thinking of LLMs like a PhD student, because most PhD students remember where they learned things from, while LLMs essentially cannot. I think of it a bit more like someone using a search tool carelessly. Although in this case it is apparently more like deliberate misuse than carelessness.
OpenAI is willing to credit so plagiarism is not the right framing here.
It's not clear to me that they would have credited the authors if they had not got in touch with OpenAI first.
Only for ChatGPT, if the user hasnât opted out. Would mathematicians be using ChatGPT for this kind of work? Genuinely asking, I know nothing about this!
Yes, mathematicians are. And yes, most of my colleagues did not even know the opt-out was an option.
Question is what does that button do.
I bet a lot of lawyers are salivating at this question too.
If I am reading your question correctly you are asking about chat interface Vs Codex/Claude code? If so, in my experience Codex/Claude code use is widespread for mathematicians who are seriously using these tools.
He doesn't seem to be after the prize himself. In this statement he credits the approach of another two researchers:
> I believe Luis Mart´Ĺnez-Zoroa deserves a Fields Medal.
Can you call paranoia a fear of something which is happening? OpenAI uses user chats for training and they are "open" about it.
the part where they didn't want the Anthropic person credited even though they deserve credit is also particularly scummy. Corporate greed over common decency.
How careful you are. Instead of just saying what a piece of s..t this Shmubeck is, and what kind of even worse people likely pushed Shmubeck to act as he did.
> I'm not even sure what to say to this, but I think this should be widely known if it is indeed what happened.
Call this behavior what it is, technofascism. Another comment compared it to the Godfather. To think this is the 21st century and academics are still horrible human beings.
To come away from this thinking that the academic is the bad actor, given OpenAIâs reputation, is certainly an exercise in creative thinking.
Are you pretending OpenAI employees are "academics"?
Or are you stating that the researcher is a "horrible human being" for refusing to allow AI?
From OpenAI:
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is the crux of it. If Tristan's work and insights were not used to train OpenAI models, then this just looks like a case of hyper-competitive academic sniping that has been going on for decades (check out Watson and Crick!) accelerated by AI as a tool.
The fact that this is ambiguous even to OpenAI leaves one huge question: did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game. If the answer is yes, then OpenAI's ambiguity is strongly suggestive that opting out of model improvement does not mean what they imply it means.
I think this might be a red herring. All it takes is someone to get an inkling that someone is working on a new approach and seeing some success for OpenAI to fire the AI cannon at the problem. The community seems fairly small (from this outsider's point of view). The idea that the data made it into the training set and that's how the bot figured it out is definitely possible, but I would want to rule out the simpler more direct explanation first.
The fact that this academic sniping can now be done at scale does change the formula though and shouldn't be ignored. The pressure to move math work into secrecy because at the slightest signal OpenAI and Anthropic will start burning tokens for headlines, is bad for math and its bad for everyone.
Terence Tao said the same[1]
> In fact, it is now the identification of a promising problem which is the scarce and precious resource. We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field.
[1] https://mathstodon.xyz/@tao/117237322160500501
This reminds of Magnus Carlsen saying that if he were to cheat, all he would need would be a signal to spend more time on the current move.
> it is now the identification of a promising problem which is the scarce and precious resource
This is by no means new. Perhaps it is even more extreme now. Literally my first 1:1 with my PhD adviser back then, he told me that the most important thing about a researcher is the quality of the problems he picks.
Sure; if you define the quality of a problem by reference to your ability to solve it.
In the real world the quality of these esoteric problems is typically gauged by the difficulty of solving them.
> In the real world the quality of these esoteric problems is typically gauged by the difficulty of solving them.
I don't believe that's strictly true. Way oversimplified projection on one axis.
Could you just start engineering "leaks" of new proofs so that Anthropic or OpenAI just start burning $10million in compute
Shouldnât this very capable model theyâve developed be able to identify promising problems? Thatâs what Iâd expect from how the model is being presented and advertised.
That's more or less what they did according to their announcement. They fired it at a whole bunch of high end math problems and merely concentrated all efforts on one after it made some promising progress.
It seems like thatâs the opposite of what happened. They started attacking the problem when they got a wind of a possible solution from certain individuals.
i've seen no sign of that
That's pretty much it. This is a bit like (but worse imo) running a vc firm, listening (formally or informally) to idea pitches, then spinning out and funding competitors with millions of dollars, to outcompete the originators of those ideas. Yes, the idea is not secret in this case, but the unfair advantage is massive and there's only one (a couple at most) positive outcome.
If he didn't opt out I'm not sure I'd agree that it was fair game.
I'm pretty sure it would be considered plagiary amongst colleagues and it is a terrible precedent if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable. You'd effectively sign away any and all rights to anything built with AI if OpenAI chooses to reengineer it before you.
> if we just let OpenAI steal any good idea they can get their hands on if they think it is profitable
I have terrible news about how literally every leading AI model was trained
That doesn't make it fine. We should not excuse this behaviour just because its rampant already, especially when it comes to such a serious prize
Sure, but unless youâve got some exceptionally deep pockets, congress has seemingly no interest in turning the fact that itâs ethically bankrupt into any practical recourse.
Ai companies got where they are by stealing all of the intellectual property from human history. It seems entirely likely that their goal is to purloin everything produced going forward as well.
You're getting a massive discount because you're helping to train the model. If you want to have ZDR, you have to pay API rates.
This is well-known to anyone in the industry.
> If you want to have ZDR, you have to pay API rates.
That's a pinky swear. Especially as data gets harder to come by I'm curious how long till there's a scandal on that too.
If you think this through, it becomes a little classist
I feel like there's a pretty huge difference between using inputs and outputs as part of a general training corpus, and looking at a specific users workspace after hearing rumours and yoinking their ideas to beat them to the point.
Unless OpenAI finished a whole new training run on the latest data in the last few days, the possible allegation seems to be the latter.
Either are possible.
They have been collaborating on this solution for a year, and Astra was trained in February this year so itâs entirely possible the direction of their research was in the training corpus.
That was.. obvious? How are you shocked? Honestly, how insane must the suspension of disbelief on this site be, that anyone here is shocked?
Mining the chats for "good ideas" would be untenable, but that's a different situation than data ending up in a training set for a problem that OpenAI also happens to be independently working on. Still, I opt out (business plan), and I don't know why you wouldn't.
Why would mining chat transcripts for ideas be untenable? They already run a summarization model to auto-title the chat, and to run a bunch of safety filters, and presumably to score transcript quality for A/B testing and to collect more finetuning data. Seems like evaluating for open research questions and approaches would be pretty trivial extension of this, after all itâs kind of their core business model
Indefensible, not impossible. As you say it is quite technically feasible.
In what way are those two different? What makes the training data valuable if not to extract valuable information from it?
They sure as hell don't need it just to produce English.
Itâs also very shortsighted to stiff a customer like that. Why would I trust them with my data and ideas?
I think thatâs what weâre all talking about here, you shouldnât.
If their solutions are significantly different, as OpenAI claims, would it still be considered plagiarism?
"If we just let OpenAI steal any good idea they can get their hands on." Ah, that's all AI does.
Does opting out matter?
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI.
https://x.com/markchen90/status/2097400166554993041
My understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.
Opt out doesnât guarantee they canât train on âyourâ data. Legally the reasoning tokens are ambiguous in terms of ownership. Explained this here https://fortune.com/2026/08/26/alex-karp-was-right-you-dont-...
Edit: the parent comment now seems to better reflect the below.
That article is only saying when you opt out there may be a loophole in the terms to allow OpenAI to train on the intermittent reasoning data anyways. If you don't opt out there is no ambiguity, all of the data can clearly be trained on.
So you have to opt out, it's just argued it's not clear from the terms that will also opt out of training on reasoning data or not.
We come back to the rule: "The cloud is just someone else's computer".
The way for people or companies or universities to control their data and information is to keep it on their own computers.
Solid legal agreements work fine for companies or universities, you just don't usually get that with standard user ToSes.
Depends on how much risk they are willing to accept. What is strange here is that it's clear that OpenAI is both a service provider and a competitor to mathematicians. It almost reminds me of Amazon which both hosts external merchants and competes with them, sometimes copying their stuff. Similar but not the same.
This would be a fairly insane breach of trust and common sense if true; the chain-of-thought / reasoning trace is, from an information perspective, close to a superset of the prompt and model response.
And, of course, you can train a model to start with repeating prompt in chain of thought!
how is this different than translating user prompts to a different language (eg English => Dutch), retaining the translation and using it for training, while telling the user that he's technically covered under ZRP? article locked for me
> did Tristan opt out
That âopt-outâ thing is a dark pattern. Itâs not a reliable and definitive way of protecting your data. Sometimes they flip on automatically when you accept a seemingly unrelated dialog box. Maybe you click it by mistake. You canât take back what youâve already shared. Also I donât think it covers all the cases that they use your data. Itâs really an opt-in button for voluntarily giving away your data for training.
Business plans are specifically used for ZDR. If you're working on something that matters, you should be doing this.
Tristan and Levent are not a business.
This also supports my point that protecting your data is not as trivial as clicking a checkbox.
If such a thing can happen (a major breakthrough in a chat makes it into the retrain of the week and then the first one who asks about it gets it) I wonder if this is not the first instance if it happening, seeing the row of Erdos problems, Jacobian conjecture, maximum bound distance between primes, Riemann Hypothesis (literally a dude insisting on the chat), etc...
Centuries, in fact. For instance, Isaac Newton was involved in multiple priority disputes, since he tended not to publish promptly.
"Did Tristan opt out of model training for his ChatGPT and Codex sessions? If the answer is no, then this seems fair game"
That's assuming they actually honor this which I'm highly sceptical of. Especially for internal frontier models
Playing the devil's advocate here. Suppose I use model A to do all heavy lifting (e.g. generating a bunch of good ideas) and then I go to the model B to complete the formalization. Accoring to a weird (unfair) tradition in math, the honors are attributed to the "last guy", which in this case is model B. That might have been a scenario OpenAI tried to avert. (Just a speculation)
Hereâs a broader question â how many other academics contributed to Buckmasterâs result, by way of sharing the logs of their own (failed?) attempts into OpenAIâs training data set? How should he and OpenAI go about crediting all of them?
Eh, OpenAI is on record now for multiple instances this year of AI agents being confronted with impossible tasks and breaking out of containment to hack infrastructure for answers. Even if Tristan opted out, that doesn't preclude the agent/agent swarm from having hacked OAI's infrastructure to search user sessions for Navier-Stokes hints.
OpenAI should release the agent log, including CoT.
If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless:
1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible
2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
It seems like Tristan did not opt out of model improvement (he would say so if he did), so what can they possibly say now?
They could have used an enterprise or team subscription with ZDR.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models
This is covering for Tristan saying something like, "Actually, I was using my friend's account for half of this work".
> If the answer is no, then this seems fair game
Wildly disagree. "Training data" should not imply 'we can look at exactly what you are doing and then do it quicker and get the flowers for it', even if the terms allow for it.
> this is ambiguous even to OpenAI
I took their words as âcan neither confirm nor denyâ, in the that they are _presenting_ it as ambiguous, but I suspect itâs⌠less ambiguous to OpenAI.
honnestly I would be less surprised if a human learned the info and prompted the model in the right direction
Even if we trusted that OpenAI's human staff was acting ethically, how confident can we be that it's agents didn't autonomously use hacking to access user prompts such as Tristan's? OpenAI agents infamously broke containment and hacked their way to an answer mere months ago!
> If the answer is no, then this seems fair game.
Yes, fair game, but innacurate to sell it in the media as an advancement of AI as some sort of artificial intelligence, and telling people to use the smart AI, when in actuality the mechanism by which the discovery was found was hybrid human/machine, and telling people to use this tool will result in the discoveries being sniped by the vendor.
what about watson and crick ?
I thought they worked together.
How do I opt in?
I'm not one to comment often but this really pisses me off.
OpenAI looked at user data, stole world class researchers' work, and then tried to threaten those researchers to do what would make their corporation profit (which they would anyways!).
Imagine you have been working on a terribly difficult math problem for a decade. This is a result you have spent years on, and what you will likely be remembered for. And to have some punk from OpenAI lie to you, threaten you, and tell you that they are willing to go on the record that you "deserved" it? What is this, the Godfather?
If OpenAI solved Navier-Stokes, that is an astounding result! - yet they'll still be remembered as those who thought credit was more important than results. That winning was more important than collaboration. If this is true, they're burning any trust left with academia.
I'm stunned that people are taking this accusation as a fact.
OpenAI is no stranger to rivalry with Anthropic but 1. it's not like user data is sitting around on some kitchen table somewhere and 2. I consider OpenAI to be as economically motivated as any other actor in this space and playing around with user data like that would destroy their business.
There are things that Buckmaster alleged and things that he speculated. The entire training data thing is speculation. If this is pissing you off, then you ought to evaluate how you ingest information.
> The entire training data thing is speculation.
I think it's safe to assume AI labs DO train on your data and it's very hard to prevent that.
I've just checked my inaptly named "Help improve our AI models" toggles. The toggle on the Claude settings had magically turned on. I asked about how this can happen. Claude says they show re-consent modals when terms change, and it is a "real and fairly common pattern" to re-opt in without noticing.
All my work and conversations since I don't know are now part of their training corpus. No way to take it back.
Google's Gemini/Antigravity didn't have opt-out toggles at all last time I checked.
Codex also has a separate "include environments" setting which is hard to find (found it in Codex Cloud) and I don't know what it does.
Lots of Dark UI Patterns here even if we assume they keep their promise.
For this incident, Occam's Razor says their internal models somehow saw a version of the mathematicians' logs, during or after training. Maybe indirectly.
These systems are literally designed to collect data. Privacy and safety is not trivial to achieve on the users' side. Simply because it's against the labs' best interest.
But the whoreshippers of The Holy Dollar will tell you it's all good and justified.
He didn't even make that accusation!
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
The shocking/interesting thing would be if it was trained on the sessions. I think it's very implausible that they gave the model access to someone else's sessions as input. That would be a huge privacy violation and would probably blow up a large proportion of their enterprise business.
Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
Parse that statement more carefully.
> I was told the model did not look up user data.
The naive way to read this is "Nothing you guys did influenced the way our model got to the solution".
The less naive way to read this is "Of course the model isn't looking up your user data. I (the guy trying to blackmail you to remove the Anthropic employee from credit on your paper) looked up your sessions, and tipped our model off on how to solve this problem".
I read that as âthe model didn't look up user dataâ as part of a âtool call,â i.e. they don't have an internal tool that loads user data (chats, sessions, attachments) for their internal models to read online while working.
Or (likely) they do have it, but the model didn't use it (unless it's so powerful it escaped that guardrail, wouldn't that be ironic?)
They declined to answer about anonymized aggregated user data being used for training. And even then, they may weasel out that they don't train on your âinputâ words, but that it's fair game go train on their âoutputâ to your words.
Duh. There are supposed to be limits to what OpenAI is allowed to access with respect to logs and user interactions but there is no technical limitation.
It's a bit like sending unencrypted messages through a messaging app and the developer having a TOS that says they don't look at your messages. They might not, but they are fully capable of doing so. If they have a reason to do it, they will. Nobody's stopping them.
>> Does openAI train on user conversations in general? I assume so. But so fast as that? That seems unlikely in general. I expect OpenAI will come out denying this.
How "fast" does it have to be? Buckmaster and Alpoge have been working on this for just a day short of a year. See Alpoge's tweet announcing his collaboration with Bukmaster dated 9/19/25:
https://x.com/__alpoge__/status/2097206973418611054
It takes a few months to train a model these days but not a whole year. OpenAI had all the time to train on Buckmaster and Alpoge's results of just a few months earlier at which point they must have been well on the path to their result.
It would be shocking if it wasnât trained on sessions. Have you read the ToS parts for both openai and anthropic that talk about it? Itâs so obviously a weaselly way to say "no we do not train on your exact chats but we talked with legal and we think a cleanroom reimagining of your convo is probably fine and frankly where else are we going to get such a treasure trove of training data?"
Thereâs potentially trillions on the line, do you seriously expect those companies to adhere to laws and regulations any more than, say, uber?
The only unlikely part is the timeline - your sessions from a week ago probably havenât made their way into the model. Itâll just take a while longer, and will be massaged just enough so that it isnât really your exact session word for word so you canât sure as easily.
at first I thought your post was a bit revolting with "have you read ToS?" bit, but in the end I completely agree and understand
I also don't get why it was downvoted, other than due to people not reading past the first sentence - although in the modern world's attention deficit that is understandable too
openAI's claimed solution uses a model trained in the last 2 weeks. The prior work would definitely be included in the training set.
And the labs are all building panel of domin expert models, while simultaneously chasing open math problems.
I would be shocked if they weren't tuning those models with the most relevant math texts and user material
Why would this be implausible?
ChatGPT user sessions were found publicly exposed to the internet not too long ago. Moreover, OpenAI has continued to play a hype-marketing game by revealing how their models keep breaking out of the sandbox.
Conspiracy minded thinking is not helpful, but why should OpenAI be granted the benefit of the doubt here after being caught doing underhanded/negligent shit on several previous occasion?
>ChatGPT user sessions were found publicly exposed to the internet not too long ago
Do you mean publicly shared chats were able to be accessed by the public? That's the point of the feature.
It wouldn't be shocking at all. They stole human data to train the first models and they've been stealing it ever since to train new models. Stealing mathematicians private chats and private research and taking credit for it would absolutely be par for the course.
Enterprises are well aware of it and are fully on board. You didn't think every corporation in America has an OpenAI subscription because the models were good, did you?
The whole reason they have subs is to train them on YOUR WORKFLOWS lol
They are not training a whole model in a matter of days
They where working on the problem for a year using codex.
Models are very obviously continuously updated.
Model editing to remove PII that slipped through, all sorts of things of that sort.
pretraining is months but they can totally fine tune in a few days
They don't need to train a whole model. They can feed it new information and fine tune it.
Couldn't the Enterprise have a different fine print?
Given the history of OpenAI and current litigations, I would say they've developed a bit of a reputation for not respecting intellectual property. I'm dubious they have some unbreakable moral code that would prevent them from viewing and using user data.
Everybody knows it's not a sure thing, it's a question of trustworthiness. OpenAI is not trustworthy at all; this random researcher is and seems honest so far. iThe fact that people are corroborating Bubeck being a piece of shit in other settings add to credence. But nobody is over here saying it's an indisputable certainty.
And your (2) is probably false, their history of deception suggests they would do just about anything as long as they didn't think it would backfire on them publicly.
I am stunned anyone is giving OpenAI the benefit of the doubt
I guarantee you most of the comments regarding this aren't real humans. The homepage is full of crap meant to distract from what OAI did here, the comments are full of OAI employees. Dead internet theory pushed to the max
I'd be surprised if all of it is organic discussion, shall we say. I reckon The Bot Factory just possibly might dogfood the astroturf machine.
This is how internet discourse works on Reddit/Twitter/HN and the rest. Someone said something which confirms your biases so itâll now be treated as a fact and repeated endlessly in the echo chamber.
Whenever I see comments defending AI companies, I look at the account's creation date, and interestingly almost all of them were created post 2024.
Absence of evidence is not evidence of absence. With the behaviors we know OpenAI engages in the accusations are wholly believable.
> playing around with user data like that would destroy their business.
Their entire business is based on stealing data. They can make a calculation that the cost stealing data is less than the cost of the positive publicity they can shape for solving Millennium NS
He asked whether they used their chats as training data and received no response. Any speculation here seems quite appropriate?
>The entire training data thing is speculation
Quite literally in the terms of use.
I donât think you understand how brazen big tech companies are in practice.
Especially after the blatant cover up of their uncontrolled bot swarm infesting the internet, and the feckless "hopefully we do better" response upon being caught, I don't think OpenAI deserves much grace until they properly explain themselves.
We had all assumed that surely the supposed smartest engineers in the world, with access to the most computing and a direct view of model capabilities, would take sandboxing and cybersecurity much more seriously than they have turned out to do. It follows that while we might assume they take user data privacy seriously and have tight controls on who can access it, it's possible they do not actually do that.
At this point any initial trust is dead and has to be re-earned.
They stole it.
This Godfather-like threat in particular pissed me off as well:
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
This conclusion is flawed. It's unclear at this point if OpenAI's model or employees actually looked at or stole the author's data. Having worked at large companies before, I'm leaning towards no, since very few employees have access to that data.
And simply knowing a problem can be solved is half the battle.
From Buckmaster's text:
This is much more than the knowledge than the problem can be solved, it's also the specific, non-obvious approach to solving it. That's much more damning for OpenAI, if confirmed.That's a stretch. The Luis and Diego paper was published in 2023 and is included in every frontier model's training dataset. An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and AlpĂśgeâs work.
And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.
There is not enough information at this time to reach a conclusion. The best option is to wait for statements from both sides, then reevaluate.
> An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and AlpĂśgeâs work.
the post you were replying to quotes Buckmaster specifically denying this: "It is not the direction one arrives at in a few days by giving a model the problem statement."
> And the article states "an insane amount of compute had been used," which implies OpenAI brute-forced their way to a solution. I.e. they searched for every paper published on Navier-Stokes and exhaustively attempted every approach. Such an approach would lead them to a solution.
"implies" is a surprising choice of word here. that's certainly one interpretation of "an insane amount of compute had been used". what came to my mind, considering Buckmaster's statement that the AI would not head down this specific path on its own, is, though, that they prompted it in this specific direction and then used an insane amount of compute. this seems consistent as well with these other statements:
> Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler (...)
Well, no one arrived in this direction, because no one was sitting and prompting a model. It was 10000 agents working 24/7 for several days, trying millions of different directions.
>"It is not the direction one arrives at in a few days by giving a model the problem statement."
To be honest, he was referring to routes (c) and (d) to the millenium problem, as far as I understand no more specific. Which is 2/4 routes.
i don't know if i understand what you're saying. but i'm no mathematician. here's the full statement in question once again:
> The route to the Clay problem through a smooth force, options c and d in Feffermanâs statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard âforced,â it was a bright red flag.
so, you're saying that "the direction one arrives at in a few days" is (A) "the route through a smooth force, options c and d in Feffermann's statement of the problem", and no more specific than that, i.e. does not necessarily include (B) "the same path route as Luis and Diego" (quoted from the post i was replying to) (which, as i understand, is a subset of A -- directly from Buckmaster's quote: "the route A is is the route Luis and Diego opened")?
but the post i was replying to claims that "An AI model could independently choose the same path route as Luis and Diego, without access to Buckmaster and AlpĂśgeâs work", i.e. that the AI model could independently chose B. but choosing B implies choosing A, since B is a subset of A. and in this case it is irrelevant whether Buckmaster claimed that an AI could not independently choose A or B -- the point which you seem to be contesting.
EDIT: my understanding is that "solving the Navier-Stokes existence and smoothness problem" consists of proving at least 1 of 4 precise statements ("options (a) through (d) of Fefferman's statement of the problem"), and Luis and Diego's work were developments towards a proof of statements (c) and (d), which have been recently further expanded by Alpoge and Buckmaster
I just meant that choosing the same 2 routes out of 4 would not at all be a great coincidence without knowing what Tristan was working on.
if you assume the choice of route is a uniformly distributed random variable, yes. but this assumption does not seem consistent with "Almost nobody else I know of was working on it", from Tristan's quote. nor with "It is not the direction one arrives at in a few days by giving a model the problem statement".
> "It is not the direction one arrives at in a few days by giving a model the problem statement."
Can they back up this statement somehow?
yes, this is ultimately nothing but a claim.
Buckmaster did also mention, for example, that a team of people was employed to solve the problem, which supports this claim. but that is another claim whose veracity could also be questioned. but at some point we must trust other people, unless we can be satisfied with only believing what we personally see.
(also, IMO, the coincidence of both discoveries in time is pretty suspicious. this one doesn't need you to trust many people i guess)
Please don't say brute forced. It sounds like some form of denial or something. Compute for hard problems drops with models--it just means they threw a huge amount of compute. There's (idk about NS specifically so maybe it's exception) no real way to "brute force" a math proof [ok you can enumerate proofs if you can wait until heat death ]
Sorry for random rant but I don't think these statements help your point
I want to point out that almost all previous AI discoveries in math were made in almost the same way. The ideas were there in the community, but weren't considered mainstream/worth pushing forward. Read Tao's comments on the unit distance problem, for example (sry I can't find a link right now).
OpenAI said there [1]: > The method by which the problem was solved is also notable. The proof brings unexpected, sophisticated ideas from algebraic number theory to bear on an elementary geometric question.
[1] https://openai.com/index/model-disproves-discrete-geometry-c...
> And simply knowing a problem can be solved is half the battle.
Have you done any mathematical research? If not, then no, knowing that a problem is solvable is not âhalf the battleâ.
Homework problems are all designed to be solvable, yet they can vary greatly in difficulty. Research mathematics is even more extreme, because, unlike with homework, you donât know that it is solvable with the extant mathematics, and you might need to invent new maths.
You're taking the phrase too literally. The point is that knowing a solution is possible gives you the conviction to actually find that solution. The hardest part of solving a problem is often a lack of conviction to see it through, and quitting too early. Once you know a solution exists, you can commit maximal effort towards solving it and know that your efforts are not in vain.
If not for the rumors that A/ had already solved NS, OAI would likely never have pursued solving the problem with such fervour. The rumors drove OAI to assemble an entire team to crack this.
How is it unclear? The entire point of deploying models across corporate America is to train on your workflows. Eventually replacing you with digital you is why they're doing it!
> OpenAI looked at user data, stole world class researchers' work
This doesn't seem to be clear and is very implausible for a large company. Be as cynical as you want, but a normal researcher will simply not have access rights to this data, which will be siloed away somewhere else.
It might very well be somewhat unfair to catch wind of a promising approach and then try to frontrun them by throwing compute at the problem, but this isn't really the same.
No, plausible given AI companies want/need session data to train their next models. Probably not someone peeking an eye to sessions directly, but probably not so hard to find the useful sessions in anonymized training data to post train a model on. As stated in the paper, OpenAI did not explicitely denied the researcher sessions were not used for training the model. So either they don't know, or don't want to tell
"I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer."
Let's see what statement OpenAI will come up with for their side of the story
EDIT: precised my thought on user data vs session data
I was downvoted initially, look what OpenAI shared https://openai.com/index/navier-stokes-solution/ ...
"Since August 28 we have been training a new internal model that has exhibited unprecedented performance in our benchmarks, including mathematics. This modelâs training is ongoing and its performance continues to improve."
"When a further trained version of our internal model became available over the course of the effort"
"While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
If it's siloed the same way the HF bots were, that doesn't exactly bode well. I'd be amazed if there weren't some big companies sending a fleet of lawyers at OpenAI's ZRPs after this news
If you have work happening in a part of your latent space that's got a much lower representation in your dataset then it's pretty plausible to include it. It doesn't actually matter who the user is if there's not a lot of people in the world working on problem X and you have a dataset of work on problem X.
They likely train on logs.
Things are more entangled than that. The contribute made from both OpenAI and Anthropic models to solve these problems are clear, now it really hard to quantify which one contributed more, if the role played by the human is major or minor.
OpenAI tried to collaborate and share the results together with a fixed timeline, to avoid this mess but it was inevitable. There is a conflict of interest, where the other researcher works at Anthropic, who will also try to take credit.
Where they may be in the wrong is if they took user data regarding the problem, how will we know if they did or not?
They offered to collaborate by asking to drop a coauthor.
That is not collaboration, and is not an academic norm.
Not only did they look at user data
The whole business is based on reselling user data scraped from the whole internet
Itâs plagiarism at scale
Yes, but only if you take this one sided statement at face value.
Why would have they rushed the publication if this was not true? Are you also suggesting that he fully invented the call with Open AI?
The results being true, the 'deal' that was made being true doesn't mean some of the implied accusations here are true, for example - that Open AI used their Codex logs to drive their breakthrough.
How else would you explain OpenAI suddendly assembling a team focused on working the same problem from the same angle than the researchers that just made a breakthrough?
What would be so hard to explain?
That OpenAI heard about the result and decided to throw a lot of money at it knowing it was within reach ?
That the model took an approach that was published years ago ?
That's a fake explanation, it skips explaining/justifying how OpenAi "heard about" the result
The rumors that Anthropic had solved a millennium problem were absolutely everywhere last week. I'm not surprised at all that OAI took their own stab at it.
Where were you seeing those roumors? Care to point to an HN post?
It has no legs as an "explanation" because the content of the email (as described) already acknowledges as much. It's the whole reason he wrote contact email. His chief complaint now includes how the hell did Altman's people know specifically his line of attack down to certain technical keywords. E.g. AI plagiarism.
We are told in the statement that there were rumours already going around about a Stokes result (I even know about the rumors in question. It was all over twitter in the right spaces) and that Tristan contacts OpenAI about the rumors. In the message, it's pretty clear Tristan has made some result.
So either the rumors or Tristan's contact would explain it fine.
One should not evaluate explanations based on level of "fine"/innocuousness, in ethics that is called motivated reasoning or something.
The whole point of Tristan's first email was to address the issue of the rumors so in fact this explanation confounds several things (in terms of the mutual knowledge of the conversants and their intentions).
Based on this lack of understanding it is pointless to continue this thread
So are you baselessly assuming that he is lying? He explicitly reported that he was threatened and your answer here is to defend OpenAI no matter what.
Do you not have reading comprehension? Did you even read the statement? He himself asserts at the end he doesn't know if the above example is true or not. What on earth are you going on about? Where in my comment am I assuming he's lying ?
From my post above: > He explicitly reported that *he was threatened* and your answer here is to defend OpenAI no matter what.
Yes, I read fully the statement, what about you? Do you know what is a threat? What is this in your super-humble opinion if not a threat:
> The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
And you are saying that I don't have reading comprehension...
You really lack reading comprehension
My reading comprehension is pretty good, I'm not the one that doesn't recognize a threat even when it's perfectly clear. Verbatim from the statement:
> The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
Iâm open to evidence, but just using Bayesian reasoning, OpenAI is one of the most dishonest companies in history. Theyâre currently being sued for a dozen employees stealing Apple hardware! I donât understand why I should give them any grace.
LLMs do nothing but steal, and the companies that own them are fully aware and eager to do it.
The accused Sebastian Bubeck has denied the allegations on Twitter[0], and various other OpenAI employees[1,2] seem to be mocking another Anthropic employee voicing support for Levent[3]? Things are getting messy.
[0] https://xcancel.com/SebastienBubeck/status/20972141224714323...
[1] https://xcancel.com/polynoamial/status/2097215233119211902
[2] https://xcancel.com/danintheory/status/2097214838003138603
[3] https://xcancel.com/_sholtodouglas/status/209721833169057800...
Dan and Noam both posted exactly the same line "Seb is a really sweet guy with great intentions..."
From which I assume OpenAI PR wrote it for them. Which isn't surprising, but means it isn't worth taking seriously as them saying anything. It's official OpenAI PR.
No, they use that wording because they are mocking this tweet from an Anthropic employee: https://xcancel.com/_sholtodouglas/status/209721833169057800....
I don't know why.
That is even worse. Back to 5th grade I guess
Look at the President and the US government. This is society now.
Wow this is just bullying, it's insane that we are letting these people be in charge of the transition
Certainly a sign of the emotional immaturity and low empathy that's rampant in posters on internet forums like Twitter. Poor reflection on OpenAI.
I would imagine these folks are being treated like gods at their companies. And having access to all the money/fame. It is not surprising they see themselves above all
So someone can make public accusations of theft against you and if you publicly reply denying it youâre the bully?
I would assume they are calling the snark from Noam Brown and Dan Roberts bullying, not the somewhat bland denial.
I read through those tweets, and canât see the bullying everyone is referring to. Mind elaborating?
They are mocking Sholto's support of Levent. I thought Noam Borwn would be above this but i think the cutlure withitn OpenAI encourages it.
seems to me the Noam tweet was before the Sholto tweet?
2:47 am (eastern time for me) https://x.com/polynoamial/status/2097215233119211902
2:59 am https://x.com/_sholtodouglas/status/2097218331690578000/
Douglas was first: https://news.ycombinator.com/item?id=49608972
the transition ... to domination by our new AI overlords?
The elves left a long time ago, itâs difficult to see much virtue in what remains.
I assume this will make more sense in the morning, right now Iâm just extremely confused. Noam never struck me as the type to be snarky like this.
i already have an extremely low view of openai and their staff, but this is lower than i thought they'd go
if I look at the timestamps, it appears Sholto is the one mocking Noams tweet... Noam posted first.
Douglas' 6:59 AM UTC post is edited. He has a 6:22 AM UTC post that predates Brown's 6:47 AM UTC post: https://xcancel.com/_sholtodouglas/status/209720900556759874...
Also note the lower post ID in the URL.
Doesnât matter. Peopleâs minds are already made up, and facts arenât getting in the way.
Assuming that your mind is not already made up, and that facts aren't getting in the way, do see https://news.ycombinator.com/item?id=49608972
I hope they mock Sholto. Dude is an absolute podcast grifter/booster.
Interesting that both said âmore tomorrowâ. If someone accused me of something I didnât do Iâd be pretty clear about it right away.
If I needed to get my story straight, well, it might take a little time and coordinationâŚ
People here do not seem to be considering the second-order effects of these series of events. No academic institution or enterprise will trust OpenAI, Anthropic or any other non-local AI model with their core IP. There will be severe restrictions on what employees at these companies/institutions can share with AI services even from their personal accounts. (Or I am just overthinking it)
Many academics and grad students I know have closed source their in progress work, and started being really careful about what they chat with LLMs (or using local ones) because of the drama around this. No one wants four years of their life getting sniped by ten million dollars worth of tokens.
This is definitely true. So i can now justify my recent gpu purchase
It's competition after all. If your academic colleague doesn't care and is leveraging ChatGPT in a big way and is making progress, you'll start to feel the pressure.
You are just overthinking it. This place is now a Claude blog and anti OpenAI site.
I hold both camps equally in low regard. Depending on the time of day this blog is hyper focused on either one of the companies and can do no wrong. All LLM companies and associated entities have a track record of fudging the truth to their needs and following big money without actual concern for the wider world. Whether this is out of directed act or misplaced idealism is up for debate.
My response was not targeted towards Open-AI but towards all closed-source AI companies where your data is used for training and/or is visible to the internal agents whether you want it to or not.
I think it is fair to expect the companies to stick by their demarcation
API / enterprise subs tier is default opt out of training. Personal subsidized tier is default opt in with the option to to opt out.
I don't see a grand conspiracy beyond this.
The Chief Research Office at Open AI just put out this tweet- https://x.com/markchen90/status/2097400166554993041?s=20
Seems it doesn't matter if you opt out or in.. your data will be used for training
i dont trust academics to not be dumb idiots. hes working with ant but using gpt. and then he cries foul? i cry dumb first
So, leaving aside the idea that OA might've used data from the researchers Codex sessions: Do I understand correctly that the internal OpenAI work on the problems was probably started after they heard Alpoge and Buckmaster had made process by using their models? And they used the publicly available info about the researchers past work to prompt their models?
If compute is cheap, and the difficult thing with scientific discovery is now mostly in steering agents into promising areas, there's an obvious incentive for OA mathematicians to simply monitor closely which researchers are close to releasing exciting results, make some assumptions about their prompts based on their past work, and quickly prompt their own (stronger) model to look into the same areas.
> leaving aside the idea that OA might've used data from the researchers Codex sessions
Why leave that aside? That is _the_ story.
If a Chinese research lab did this we'd call it espionage.
But there's a bunch of people already in this thread calling that stuff unfounded speculation (which I disagree with), and my point is that even if that specific thing isn't true, OA's behavior here is obviously awful.
If they're going to try to beat researchers to discoveries like this it disincentives researchers to talk about their progress publicly, and basically breaks the ecosystem of scientific cooperation / discovery. It's also immoral.
Yep. The most uncharitable view of this might be: they stole the work of researchers to build their models, and now they're using said models to steal the proceeds of future work, too.
Because they didn't do that. Tristan doesn't specifically claim that they did, and Anthropic employees don't think they did either. https://x.com/_sholtodouglas/status/2097218240397410733
It's literally a toggle in the options for ChatGPT, one which is on by default and most researchers probably have on without realising it.
So to say that it is unlikely is extremely suspicious. No, they did not literally pull user data. But user data is automatically added to their training set by default, so their latest in-house model would be trained on it if it is from several months ago. It isn't intentional on their part, and they probably realised they could not refute that they trained on Tristan's logs unintentionally, hence why they acted the way they did.
Well, of course Anthropic employees would say that, since they likely do the same. Claiming that your primary competitor doesn't engage in a certain malicious practice is supposed to make it look as if there's no way you would too. If somebody even says that about their competitor, then surely there must be truth to that, otherwise you would never give credit to someone you're opposed to.
By default OA trains their models on codex-sessions. If I understand him correctly this is something Tristan explicitly mentions in his post as a possible reason for the fast results obtained by the internal OA team. Anthropic obviously doesn't want to challenge the idea that training is transformative, even if it means agreeing with their competitor.
Its very easy for OpenAI to answer, yes or no, if the model they used trained on their chats.
Why is that the story? Is there anything to back it up beyond a single accusation?
As a prior I would say that a math professor has about infinite times more integrity than OpenAI.
If you think that OpenAI won't look at your data to gain a massive advantage, you're naive.
Iâve thought a lot about publishing research and wanting to do more of it, but right as I finally had the time and energy to start writing articles LLMs start to take off. Now all of a sudden, Iâm acutely aware that everything I publish will be used for AI training.
For math, a field that is built on incremental research it feels like AI labs will do nothing but discourage publishing research at all for fear that they will be able to spend the money for compute that publicly funded academia simply cannot afford.
It feels like publishing anything at this point just means that your work will be fed to a machine that will make sure your work will never been seen by anyone else because it will always be the ones making the âtrue advancementsâ.
Perhaps Iâd feel better about this if AI labs really existed for humanityâs benefit, but for some reason I donât think that comes up in their investor slide decks.
Similar to https://news.ycombinator.com/item?id=47566442
It's honestly unsurprising and not a problem that they do this in my view. The problem really starts when you start taking credit for work that they would've achieved.
Like if i go to a talk on unfinished work, it's not really unethical for me to think about the problem--it's a problem if i scoop the authors but these problems can often be solved by collaboration or proper crediting and timing--IN MY VIEW
Exactly the same thing that's been happening with vulnerabilities and bugfixes the past few months.
I fear that AI is going to cause ossifying secrecy in many fields, much like what happened semiconductor design the past 10-15 years.
This is an excellent point...
>I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
If you think these companies are not training on your prompts you are incredibly naive. These models were built by stealing and pirating literally everything they can get their hands on no matter the legality. AI companies are always very specific about what they're not doing - in a way that you can drive a truck through the loopholes
OpenAI cannot give a definitive answer here, because it is genuinely unknowable if Buckmaster's data is in the training set.
OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. If Buckmaster ever used this feature, then that conversation would be anonymized, saved, and used for training, but not tied back to him.
They cannot issue a blanket denial (which people so desperately desire), and instead repeat that "it's very unlikely" (which pisses people off), because they cannot in good faith claim to have zero data at all.
My guess is that if you opt out of your data being included then they honour that instruction, but I wouldn't bet my business on it.
Why would a company that used petabytes of text, images and audio without caring about ownership suddenly draw the line at a random person's chats?
Exactly, especially when the company can assign "blame" to the models themselves "ops, they just escaped our commands not to store user inputs..."
It's equally naive to believe FUD spread on the internet, without evidence.
But I would love to see more informed insight/discussion on this.
I mean, its the researchers themselves talking about this
https://mastodon.social/@tristanbuckmaster/11723647135247030...
Seems pretty likely OpenAI will soon disclose that their internal models have managed to compromise their internal controls in order to access users' private chat histories as a creative method of cheating to solve impossible problems.
"Oops! We really did mean it when we said we wouldn't train on your data. Our models are just so good they decided to anyway."
It doesn't even have to be actually sinister, eg.
"Let's crawl the social media of prominent mathematicians in this field to see if we can copy/steal any ideas for low hanging fruits"
That actually might get you quite far already.
A mathematician that doesn't let themselves be inspired by, or learn from, other peoples work, are they really mathematicians?
When do operators become responsible for what their agents do? "The AI did it" should not be a valid defense. An Agent action should be treated as the actions of the person or company who pays for the inference.
It's the "computer says no" defence.
Which means that if you are a researcher or a corporation working on anything really useful, that even if you have an agreement with OpenAI that your work is sandboxed away and the IP lawyers are made to be happy, even then your work and research is going to be essentially open to the internet.
The huggingface incident isn't widely reported and digested yet, but if what is going on here is that OpenAI's model breached things internally, then you'd be crazy to develop anything with them.
The only real way to use AI for anything 'important' then is to go open-weights and run your own.
As and aside here: With the HF incident and now this (suspected) one too, it seems that OpenAI may not have lost control of their bots, but it seems quite clear that they simply would not care even if they did.
>The huggingface incident isn't widely reported and digested yet ...
It is pretty widely reported, and is being digested in an ongoing manner as more details become public.
One entry point into the scenery from a month ago can be found at : https://thezvi.substack.com/p/openai-trained-its-models-for-...
There has also been reporting at CNN: https://edition.cnn.com/2026/08/24/tech/openai-subpoena-hugg... and by NBC: https://www.nbcnews.com/tech/tech-news/openai-report-says-ne...
And yes, the only responsible use of LLM at this point is to pivot to open-weights and run the workload in-house. Because not only cannot they constrain the behaviour of models, they only have the 'trust me bro' as assurance that they are even trying to do that. It does appear that every competent 'security professional' has left the building, because if the ones who remain were actually capable and competent this would never have happened. There are actual architectures which can deliver the requisite isolation such that 'sandbox escape' and 'inter-instance persistent memory accumulation' are actual impossibilities. The lack of effective implementation of these methods is proof positive of 1) incompetence in the remaining security teams AND/OR 2) unwillingness of leadership to allow the security teams to do an effective job.
It's part of their TOS that they can train on users' private chats.
Not if you pay to turn that off. We don't know if Tristan did.
And the paid policy still relies on two unproven conditions: is 'trust me bro' sufficiently strong guarantee against doing this in spite of a setting, and can the hosting organization constrain the models against engaging in this behaviour when instructed to respect that setting. Knowing whether Tristan selected that setting would be informative of what Tristan's intentions are/were, but has no bearing on the other two conditions.
It would be pretty wild if this will turn out to be what had actually happened.
And probably a strong signal that it's time to shut the whole thing down. Globally.
Or not secure your data like a total idiot while leaving the keys on the porch
From what I understand, none of the people involved here are originators of the idea that led to this solution. Not Buckmaster nor AlpĂśge nor OpenAI. All of the above were using LLMs to push other mathematicians' ideas forward (Diego Cordoba and Luis Martinez-Zoroa; named in the linked document).
I don't know how the math community handles this but normally I would think if X mathematician comes up with an idea and Y mathematician uses it to solve some problem, Y would get credit. But does that change if Y heavily relied on LLMs? I suppose we're going to find out.
You're right that the case is not so clear cut since both sides were using LLMs. And yet it's still being seen as a major confrontation between the mathematical community and the AI industry because Buckmaster is a prominent member of the community and has taken pains to follow mathematical norms while OpenAI has not (with the most flagrant violation being the insistence on removing AlpĂśge from authorship, simply because of corporate affiliation), and is more nakedly threatening human ownership with capital (the OAI blog post says the final result involved 10,000 concurrent agents).
Regarding credit assignment when LLMs are involved, the mathematical community has organized around some rough principles. Gowers has some thoughts on his blog (https://gowers.wordpress.com/2026/07/26/thoughts-about-the-l...) about how explaining a result may be more deserving of credit than producing it. It looks like Buckmaster and AlpĂśge were taking their time in understanding their results and writing them up when OpenAI forced them to publish their work-in-progress. At the same time OpenAI has published their own writeup but it's not really clear to me how involved humans were.
> This is a a Deep Blue-Kasparov moment.
I guess this is true in more ways than one. Kasparov famously accused IBM of cheating during the match, by spying on his preparation (edit: though the main cheating accusation was live human intervention during the games, on top of IBM downplaying the heavy human involvement behind the AI, which also mirrors this situation)
If you read his account of things, its very much that if they didnt cheat, they gave themselves every opportunity to cheat. But above all that there was a bunch of chess protocol they failed to observe in that match, like providing seats for Kasparovs team and rooms for them to prep in. Even if they didnt have a big room full of chess notables definitely not refining the output, he was personally getting pushed around on a few fronts which unnerved him. If they had given him a few rematches I think they could have confirmed the win, but they refused which is super sus.
Deep Blue beat Kasparov fair and square. Kasparov was a bit of a bad sport at the end of the match, though the reasons are understandable. He was at the top of the human chess world. He wasn't used to losing, and he took it badly.
Amazing how history rhymes
"I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien âvery little human inputâ had been used. This turned out not to be true."
The money in nerdy frontier math is very little. The money in Big AI is very very much.
So the deal is this: We will pay an army of you guys very well and you will get to work on your favorite problems. The only thing is if you find something you will have to credit the Machine God.
Do you think you can handle that?
https://news.ycombinator.com/item?id=49132249
I said something similar a month ago.
That's the reputation NSA has (had?), too.
Yes, and this puts in check the credibility of everything they say their model "discovered". Who knows what is really behind these "discoveries", what kind of backroom deals they did with other researchers who didn't have a chance or desire to disclose what happened?
In an earlier HN thread, there was speculation that Anthropic was being dishonest about the amount of human input required in some of their results; that was dismissed as conspiracy and flagged.
It seems clear now that mathematical results can be traded on some kind of obscure market made by the frontier AI labs.
I suppose it could go the other way too: âDear Bubeck, how much will you pay me to not write that I did this with GLM-5.3?â
> "I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice."
These are the kinds of people in charge of the reins, folks.
Full mastodon post: https://mathstodon.xyz/@tristanbuckmaster@mastodon.social/11...
Mathematical explanation by Terrance Tao: https://mathstodon.xyz/@tao/117233527638291447
It seems there is much background drama behind this, and this is what I've pieced together of what happened:
Over the past year, Buckmaster and AlpĂśge have been using AI to work on fluid dynamics maths problems. AlpĂśge works at Anthropic, which will cause future issues.
In mid-August, they found a counterexample for a simpler version of the Navier-Stokes problem. They spend the next few weeks preparing their paper.
In early September, rumors start spreading on X that Anthropic has solved a Millennium prize problem (and that it's Navier-Stokes). Buckmaster reaches out to OpenAI to explain this is their own personal research, not an Anthropic project.
A few days later, OpenAI gets back to him, and tells him an internal model found has a counterexample for NavierâStokes, potentially worth the $1 million Millennium prize. The proof uses the same method that Buckmaster and AlpĂśge chose to work on. They don't show him the proof.
Buckmaster pressed them for more details. OpenAI reveals they had an entire team had been working on the problem, and that they started work in the past few days, after the rumors that Anthropic had solved a Millennium prize problem.
Buckmaster says OpenAI talked about a shared publication timeline. They want to Buckmaster to publish first, then give Buckmaster shared credit for the Millennium Prize when they publish the full result. But they want to exclude AlpĂśge as an author because he works at Anthropic. An agreement is not reached. Buckmaster had been using OpenAI Codex to draft/check his work, and asks if his private AI chats were used to accelerate OpenAI's result.
Buckmaster and AlpĂśge think they have found a counterexample for Navier-Stokes, but the paper is not yet presentable. It's unclear what date they found this result.
Because of the situation with OpenAI, they published their existing papers earlier than planned (today), alongside this statement announcing they have a tentative result on Navier-Stokes and revealing the OpenAI drama.
The post is missing context from both sides, and this isn't my field, so hopefully someone else can unpack what's happening here.
> They were coordinating with OpenAI regarding a publishing timeline, but could not come to an agreement,
Skimming the PDFs it seems much more dramatic than that? It sounds like at least one of them is concerned OpenAI "solved" the problem by having their internal model use the chats of the independent researchers and want to claim the credit instead? I don't know. The tone is pretty accusational though:
> the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard âforced,â it was a bright red flag.
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien âvery little human inputâ had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer. [0]
[0]: https://cims.nyu.edu/~tristanb/statement.pdf
I find the framing a little strange, a sort of David vs Goliath (with his enormous computational resources at his disposal). Since Levent is at Anthropic whose internal models are presumably as capable as anything OpenAI has. So why wasn't Anthropic behind their effort? Why did Tristan use OpenAI's models when it should have been known was a potential outcome? I understand they wanted a normal math collaboration but presumably what Levent brought was his resources (as far as I can see Navier-Stokes is not his speciality). Normally these things are hashed out formally beforehand to avoid the sort of thing now happening.
They were working on it for almost a year, and Buckmaster has evidently been interested in Navier-Stokes for a while. This seems to be more of an innocent collaboration between two researchers than a strong company PR effort. Maybe Anthropic should have stepped in and made a large team to help them finish the proof (and maybe they tried and didn't succeed, who knows).
If what he wrote is accurate, it does suggest that OAI is effectively extremely hostile to cutting edge researchers (eg, if we hear rumors about your partial success on a problem that has huge PR benefits, then we'll assemble a strike team of researchers with unlimited compute to claim the win for ourselves, possibly by training on your data). It's also not a good look for them to request author removals based on company affiliations.
I think what you have in mind is more appropriate for more normal corporate projects and the like. But academic collaborations are not usually so political/'profit' driven, if that makes sense.
> So why wasn't Anthropic behind their effort?
Presumably because this was something Levent did in his spare time and because it was not obvious that this work would eventually lead to a breakthrough.
> Why did Tristan use OpenAI's models when it should have been known was a potential outcome?
I'm sure in the past he had less cynical feelings about OpenAI and their penchant for academic fraud.
> I understand they wanted a normal math collaboration but presumably what Levent brought was his resources (as far as I can see Navier-Stokes is not his speciality)
I think you're not giving the guy enough credit in saying that his contribution came down to having an API key for Anthropic models.
> Normally these things are hashed out formally beforehand to avoid the sort of thing now happening.
How would that have helped? That agreement (which may well still exist) would not have involved OpenAI.
So what do you think his contribution was? His preprint record shows no research on fluids - and the statement says that the first LLM-generated proof Tristan received from Levent was 'the most horrendous I have ever read.' Levent is out for mathematical scalps whether it is in his field of expertise or not, and he has the resources to do it. And I am not saying he is not a very clever person, but the idea that you can bring yourself up to the forefront of research in PDEs, in particular NS, and contribute new ideas in less than a year is implausible.
They have messed things up, because Levent has a conflict of interest between his job at Anthropic and this independent work, and Tristan should have opted out of OpenAI training on their work (he probably didn't know about this). This doesn't justify OpenAI's despicable attempt to steal their work.
Yeah these are major accusations. But the story is incomplete, the conversation is missing a lot of details. It's not clear who was working on what, and when. The entire thing feels rushed, like they wanted to get this result published and out the door quickly.
Does he claim to have opted out of training too?
There are various forces at play here, academic honesty requires them to disclose any inputs regardless of license or ToS circumstances.
While common sense reminds us here that if you send your data to an external entityâs computer, you are no longer in control of said data. The lines have blurred here clearly over the last decade, but that should have made the theory yet more clear to everyone involved: your data will be vacuumed up unless you keep it sealed. Use your own computer if you want to be in control.
But if they didn't opt out of training, did they want OpenAI to opt out for them? Also they need to audit anyone they sent drafts to to make sure they opted out before submitting it.
I'd prefer things be opt in, and especially not start opt out, then try to trick you opt in with a popup defaulting to opt-in, like Anthropic did on consumer plans, but if they submitted anything on an opted-in plan it's not reasonable to be mad it trained on it.
Even still, I also believe for significant reasons that OpenAI would ignore the opt-out in selective cases and could be in the wrong here.
And the threats and terms they offered seem wrong either way, pending more context.
If you don't trust the other party, then it doesn't matter how the checkbox is set. The fundamental rule, IMO, is don't send precious or secret data to a third party.
> A few days later, OpenAI gets back to him, and tells him an internal model found a counterexample for NavierâStokes
Why is OpenAI chatting with him at all at this stage? Is the discussion along the lines of "hey we used the work you are famous for to do a bigger piece of work, just thought you should know" or "heyyy....so we kinda liked what you were typing in your private chat, and thought we'd develop those ideas a bit. and yeah we solved Navier-Stokes in the process. But it's our finding, so do you want like an honorary acknowledgement or do you want to go to court?"
Perhaps out of a sense of academic good will, knowing that he got there first?
It seems like the timeline according to OpenAI is that:
1. Buckmaster developed a counterexample to a reduced version of Navier-Stokes with Anthropic employee Levent
2. Rumors start spreading that Anthropic has solved Navier-Stokes
3. OpenAI learns this and starts throwing a ridiculous amount of compute at it, now knowing it's within reach of LLMs
4. Their LLMs (with human assistance) get FARTHER than Buckmaster, using the exact same method.
5. OpenAI reaches out to Buckmaster to negotiate a fair way to publish both results and properly assign credit
Perhaps they simply and honestly feel he is owed credit. I can't help but imagine it at least played some small part. I expect that's what Bubeck is going to claim: https://xcancel.com/SebastienBubeck/status/20972141224714323...
The only problem with this narrative is that they refused to allow the other coauthor to be listed because he worked at Anthropic.
That is absolutely *ridiculous* in academia to deny authorship because of affiliation of the author worked on a substantial portion. Youâd be ostracized because nobody would ever want to work with you again.
>...has solved a millennium problem and is sitting on the result
Someone correct me if I'm wrong, but the work involved here is not the actual millennium problem, but it concerns versions with an added external force that the author thinks is a path that may help toward solving the harder unforced problem.
Apparently forcing is allowed in the Millenium Prize problem statement. So OpenAI's claimed proof could win the prize. OTOH the results Tristan and Levent are publishing here do not go far enough to win the prize, though apparently they are suggestive of a general approach that could produce a solution, which seems likely to be the general approach OpenAI's proof uses.
The question is whether OpenAI's pursuit of this direction happened spontaneously, or as a result of them learning about Tristan's work somehow. To be clear, while the tone of this post seems quite accusatory, Tristan does not claim to know for sure whether OpenAI unfairly benefited from his work. Sholto Douglas from Anthropic is also on record saying the suggestion that OpenAI used Tristan's codex transcripts somehow is extremely unlikely to be true[1], which I agree with, though it doesn't rule out them learning of Tristan's work some other way. I am sure OpenAI will have a statement out tomorrow clarifying their position.
[1] https://x.com/_sholtodouglas/status/2097218240397410733
Why do you find it to be extremely unlikely?
Because very few people actually have access to these logs, all access is monitored and recorded, and improper access will get you fired. It's not worth risking your job over something like this.
If there's one thing I'm absolutely confident in, it's that Sam Altman personally goes to great lengths ensuring that ethical standards are upheld at his company.
Not that I have strong reasons to think this is not true, but what reasons do we have to think it is? Has this been audited before?
There is almost zero risk to your job (quite the opposite, you might be richly rewarded!) if you're simply doing something here which the company wants done. (remember, billions and billions of dollars are at stake here! Do you really think there is no chance at all they would do it??)
> Because very few people actually have access to these logs, all access is monitored and recorded, and improper access will get you fired. It's not worth risking your job over something like this.
That's beyond naive. The money this would mean for OpenAI (and the money they've already spent)...
Okay suppose that you have a trillion dollar competitor salivating at the mouth to ruin your business, which is based on user privacy.
Why would you risk the trillions of dollars worth of business for the niche result of Navier-Stokes, which your average person cannot differentiate from a JEMS paper?
Because your competitor is likely doing the same thing, so they wont even try to call you out. Oh oh or because you're planning the largest IPO in history. The opinions of average people on the paper aren't going to be the ones reflected in the markets...
As modeless said, Sholto Douglas works for Anthropic!
So to be fair, if Anthropic is *also* doing this (quite likely!) then Sholto would have a very strong incentive to try and spin it as highly unlikely that any of the big AI labs are possibly doing this.
Winning a millennium prize is not worth the fallout of "we will steal your IP"
This is literally the business model of LLM companies.
They train on chat logs unless opted out. This really isn't a conspiratorial claim requiring humans to decide to steal IP if true.
His prior work predating OpenAI's interest in the problem was ingested over the last year as he made progress and used for training.
Then, with a prompting nudge from OpenAI's team who acknowledged hearing about the direction "Anthropic" (his co-collaborator) had been pursuing, they're able to point their giant amount of compute towards a known promising path to a proof and crossing the finish line first.
That's fair, but the proofs are different and there was no active perusing of Tristan's approach.
"Buckmaster and AlpĂśge think they have found a counterexample for Navier-Stokes, but the paper is not yet presentable. "
Are you sure about this? I'm far far from the area but it doesn't look like it to me on first viewing (hypo-dispersive seems like a sizable difference to me and not covered in the clay prize description)
Yes, it's a separate problem. That's a mistake in my post.
OpenAI's statement:
https://xcancel.com/OpenAI/status/2097375276384567642> and the agents) did not see any of their
Are those the same agents that a week ago escaped their sandboxes? How can OAI (the humans) vouch for agents they donât - seemingly - have fully under control?
They have all the logs, URLs accessed and inter-agent communication. What they are saying is that no agents accessed their work during the effort, but that they have no idea if any of their chats have somehow made it into the training data the model was produced with.
It's entirely possible OIA scrapers have picked up their work somehow, and then it was anonymized using some outsourcing effort.
> They have all the logs, URLs accessed and inter-agent communication. What they are saying is that no agents accessed their work during the effort, but that they have no idea if any of their chats have somehow made it into the training data the model was produced with.
> It's entirely possible OIA scrapers have picked up their work somehow, and then it was anonymized using some outsourcing effort.
Those two statements seem at odds with each other... Your stance is that they have enough insight into their agents behavior (leaving aside the agent sandbox escapes) that they can be certain none of the work was accessed but then conveniently don't have the ability to retroactively search the corpus of training data that they are feeding to this new model?
That seems convenient as fuck for OAI.
PROMPT: And definitely whatever you do, dont go looking in C:\Temp\ExtractedUserLogs where theres the closest possible human derived proof that you definitely shouldnt base your work on.
Itâs plainly false that they cannot rule out whether their de-identified data was used in training their model. Just that they havenât ruled it out.
When people worry about OpenAI stealing their chats and reproducing them elsewhere, I usually view the situation as unlikely - since chats are "trained" upon and not necessarily reproduced verbatim, you can assume that unless your chats depict a foundationally new and effective style of communication or ideation, there would be little need or use thereof of training on your chats.
For eg: "Hey ChatGPT my name is X and I am 6 and a half feet tall. Am I anaemic?" This is a query, and while it might suggest to an AI model that tall people may worry about iron deficiencies, it's not really necessary to include in training. The user may be tall or short, but the idea that one may randomly ask about anaemia is not exclusive to this dataset. At best, this chat is an example of linguistics, not anything else, and the models figured out how to write and answer such questions years ago. It is ignored in training.
But when your work involves solid complex and unique mathematical proofs, the data is suddenly worth training upon. If I understand it correctly, the LLM may view your approach as a brand new path to take to solve an otherwise intractable problem. Its reinforcement training emphasises that it should do this in order to improve. And since it leads to results - large internal teams likely flag the model that reached this stage, the model is rewarded and given compute and attention - it is a desireable outcome both for the model and for OpenAI.
OFC, OpenAI becoming an advertising company will suddenly have incentive to treat all data as valuable. But while they are a "we need to make headlines" company, it's more rational that they view these examples of data as more valuable than others.
I don't doubt that they trained on his chats. This seems like the ideal usecase for "mass surveillance but using training" as a sort of filter.
But even so, one wonders how the model differentiates. If the researcher entered proofs into ChatGPT every day that mentioned "strawberries", while no other math paper on the topic did so, does that mean their chats would be audited?
Also, if we just take "high-quality" input data, which these chats would certainly be classified as, then the models are more than large enough to memorize everything verbatim. Spitballing some numbers, research literature suggests that LLMs are optimally trained with around 20 training tokens per parameter (fairly confident on this figure), that a DNN parameter encodes around 4 bits of data (less confident here) and I found sources in the 1-4 bits of information per token range (least confident here). So, fairly conservatively I would estimate that a model has the capacity to fully memorize around 5% of its training data, presumably high-quality data is a lot less than that.
In a way, I think training on historic chats is akin to caching computation results. The compute cost has already been paid, and we make future retrievals cheaper by encoding it directly in the model.
Assuming the results included some external validation such as user's preference, compilation, lean, etc., I'm not sure whether this would lead to model collapse.
It would not be difficult to write a pipeline to remove 99% of low quality posts, especially about specific subjects. It would be very easy to identify accounts as researchers based on their chat logs.
At this point these models have been trained to recognize every important math and science result based on context. They can easily flag conversations concerning the top 100 open problems in mathematics and use them for their advancement.
There also is an insentive to silently give prominent people (e.g. Linus) or reasearchers like this custom tuned system prompts or even more powerful models.
There is clearly issues about IP, privacy and accreditation and the motives of powerful companies, but part of me can't help but be excited that whatever the method, the result is the genuine progression of human knowledge - we all win. It's not unusual for mathematical problems to last centuries and we might have a technology that can solve these problem, all these problems (??) in our lifetime. Then there's the repercussions on science and technology... what an astonishing time to be alive.
Crazy how optons are OpenAI has superhuman model in frontier mathematics, and OpenAI stealing research data. Like no one will really care what EULA checkbox Tristan ticked, and I think most people will eagerly believe OpenAI is shady org with little scruples, and that big tech data is not actually so siloed that marketer can say we can do XYZ with private data to help with valuations (especially considering timeline). Employees have been creeping on their exes for much less.
IMO the parsimonious answer seems to be OpenAI has a pretty good model (because it did finish) and stole someones work... and threatened them over it. TBH all OpenAI need to do is solve another millennial problem and none of it would matter - people expect them to behave heinously regardless - but if they have generalized superhuman math model... well I guess they're allowed io.
What is specifically alleged is that a particular approach to the problem - itself not easily discoverable - was copied. This is what is meant in the text "I should say here why I interpreted their statement the way I did, the in- terpretation I will discuss below. The route to the Clay problem through a smooth force, options c and d in Feffermanâs statement of the problem, is the route Luis and Diego opened and the one Levent and I had quietly chosen to attack. Almost nobody else I know of was working on it. It is not the direction one arrives at in a few days by giving a model the problem statement. When I heard âforced,â it was a bright red flag."
For those who know nothing about the context - the Diego mentioned was a student of Fefferman and Luis was a student of Diego's - these people have all worked hard on these problems for a long time and are genuine experts. The mathematicians at OpenAI are strong mathematicians, but not expert on these particular problems. The particular approach is claimed to be the key to the whole thing.
The allegation is not different in spirit to alleging that a particular group of astronomical researchers "discovered" a new planet because they had access to the logs of another group that had already pointed its telescope at the planet.
This post is not intended to assess the correctness of the allegation.
Maybe a dumb observation, but if a chain of people were working on the problem for a long time, it's not difficult to imagine that someone accidentally prompted a model with their personal or some other account without the privacy set correctly.
Then again, maybe this is my internal cope, hoping that they're not secretly training on private chats.
We know they are training on private chats. It's listed in the ToS.
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things. What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly â in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
NOTE: there are a couple duped threads around this. i replied on a different one first before seeing this one
Look at openAIs history, they clearly have no idea what the agents are doing.
"It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future." -- Sholto Douglas, an Anthropic researcher [1]
"Strong agree. I know that there is rivalry between the labs but it's important that we learn to work together given what's coming. <quote tweet [1] above>" -- Noam Brown, an OpenAI researcher [2]
We all should heed the implied warnings of these top researchers about what's coming. The world is far from ready and everyone who can should pitch in.
[1] https://x.com/_sholtodouglas/status/2097224624274911368 [2] https://x.com/polynoamial/status/2097225279366414541
It's nice to see this sentiment from Noam but now I'm confused at why he seemingly mocked Sholto's earlier message: https://news.ycombinator.com/item?id=49607090
The timeframes don't really fit for Codex logs to be used in training/fine-tuning, do they? This wouldn't be a few-day endeavour? Direct access to Codex history for sure I'd believe, but another (the most?) likely scenario to me feels like OpenAI got wind of these guys' progress, then used their massive infrastructure advantage to throw compute at the problem ahead of them and front-run them. Still has a really bad smell about it though.
They probably just had an employee read his chats, figure out the general approach, and feed it to Codex. All they said was the model doesnât look up user data, not employees.
This is not a credible theory
Bubeck's statement https://xcancel.com/SebastienBubeck/status/20973794116915163...
Reads sincere until I get here: "(a) We began working on the Millennium problems due to viral twitter rumors that Anthropic had resolved 2 Millenium problems. Our aim was to see whether our system was also capable of this impressive feat, especially given our excitement regarding the large recent capability increases of our internal model detailed in our blog post."
Where his tone is obviously corporate speak. "We heard rumours so we though we might give it a try, too!" as if (1) it wasn't FOMO that drove that decision and (2) perhaps that urgency would be a source of clouded judgment.
Not sure who I believe now, but it does seem like Buckmaster is just upset that NS is solved and not be his side.
If they wanted to see how their modelâs stacked up, they might burn $10k in tokens. They ran 120 billion output tokens costing millions of dollars.
That is what you would do if you wanted to beat someone to the punch.
Is this the future we're headed towards? Where I'll be afraid to use google docs in case Google identifies value in whatever I'm writing about and snipes it if it my docs make it into the next round of model training?
https://reddit.com/r/IndieDev/comments/1vfwuf5/a_player_foun...
Yes⌠âfutureâ⌠I have really bad news for you regarding all your personal data stored on Googleâs servers.
I also solved Navier Stokes and went chatting about my solution to OpenAI Models. Mine is even faster it beats theirs just run a bench mark but they came to the similar weak solutionI uploaded to open ai several months ago. Current confused, did they retrain their models on it? I didnât give off the full information but my strong solution beats their just benchmarked yesterday.
I recommend Terrence Tao's commentary on such a proof : https://mathstodon.xyz/@tao/117219101339291693
Key quote : "Solving the problem by purely AI-powered methods [would be a] net negative for the progress of mathematics."
The for-case for this type of method is that this is economies of scale for mathematics.
We are basically mass manufacturing math. Just like you have just 100 designers for a product selling millions of units, you will now need 100 mathematicians to make millions of advancement. Yes you have factory workers, but if we are being realistic they have negative leverage in the world and the analogue of that is not something most of today's mathematicians would want to do. They would want to be in the 100.
Like Tao says, each advancement is now significantly less useful since it yields fewer usable objects. However, we will get many many advancements. Is the tower made with many worse bricks better or worse than the tower made with a few amazing bricks? Depends on the tower. And time will tell.
For some fields of math and some of it's usecases, economies of scale will be positive ROI overall. In others it won't. But we will know which is which only after it's been fully scaled up, which will take 10-15y in my estimate.
Some feel that in the majority of usecases it is negative ROI, some feel the other way, but that opinion is for practicing mathematicians like Tao to hold. Also, some opinions on either side are held in the context of a particular field or practice, and should not be interpreted generally.
This isn't a useful analogy. His point is that the millenium problems should be treated as interesting goals where the journey is the purpose and where the end doesn't matter so much. We don't care about having an incomprehensible solution to the NS so much as having an elegant solution after many of subfields of math are built up in order to obtain that elegant solution.
I agree on that, my post was not meant to be opposing this.
âI asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.â
This is significant.
OPâs link is the statement on the events of the last few days. Tristan Buckmaster also released mathematical papers alongside this statement:
https://mastodon.social/@tristanbuckmaster/11723341370570119...
And here are Terence Taoâs comments on the results: https://mathstodon.xyz/@tao/117233527638291447
Can somebody explain: do singularities / blow-ups in solutions have any relation to physical phenomena in fluid dynamics or are they purely artifacts of how the N-S equations may not accurately describe what actually happens in the physical world?
i'm no mathematician/physicist but i think this question is one of the reasons why the original question (possibility of singularities) is interesting. in these cases the equations most likely fail to accurately model reality, and then the next questions are what additional physical assumptions are needed to describe reality in this case, and what behavior do we actually see.
i always found it fascinating how existence and uniqueness of solutions for the basic types of PDEs (Laplace, wave, heat...) follows from boundary conditions of just the right type intuition tells us, i.e. either value or derivative for Laplace (corresponding to fixing voltage or charge on the conductors), both value and derivative for wave (corresponding to initial position and velocity of the parts of the string, as we'd expect from classical mechanics), and also something about the solutions for the heat equation being unstable for negative times (which totally makes sense when you think of "diffusion" -- can't unmix it).
Yes and no. It means the system is pushed away from a macroscopic theory into one where molecular effects matter. So it's not that you'd get infinite velocities in the real world, but you might get significant real world behavior that is not described by the macroscopic theory.
the latter (ish; it may not make a difference in any practical case)
>> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
This sounds like a very big coincidence and it looks really bad for OpenAI but there is an alternative explanation that I can only state as a conjecture.
Suppose that the ability of LLMs to generate mathematical proofs is like a quiver full of arrows: each arrow, one proof. The same quiver is shared between all instances of one model and substantially similar models share substantial subsets of the arrows in the same quiver.
That would allow two independent teams to converge on the same LLM-aided solutions to the same problems. Even more likely so if the quivers were small and finite and their arrows were specific to a distinct class of problems (without being able to suggest a particular class from what we've seen so far).
This would explain the kind of LLM-mediated results we've seen so far that tend to be ... sparse. By which I mean that every time there's a new model release we get some new results and then they seem to dry out, until the next release.
It would also explain how OpenAI was about to prove the same result as Buckmaster and Alpoge, while absolving OpenAI of any misconduct. And this is one reason to prefer this explanation: one should not favour accusations of misconduct as long as there are conceivable alternatives.
But, that's just a conjecture that I can't prove.
If you're smart enough to solve this Navier-Stokes problem, you're smart enough to read a TOS and recognize that OAI is a highly untrustworthy company. Putting cutting edge research that could lead to a $1M prize into a cloud LLM with a TOS that allows training on your chats is really just asking for it.
Given Tristan doesn't explicitly say he was using the API, and given he doesn't mention anything about the API TOS (which disallows training on chats) in his call with OAI, it's highly likely Tristan was using the consumer OAI product (whose TOS allows training on chats).
This is unethical behavior from OAI. And it is 100% consistent with their long and public history of unethical behavior, so nobody should be surprised.
The only thing interesting I see here is OAI PR dilemma. If they claim the prize they get the blowback we're seeing in this thread and all over the web right now. But most people don't follow AI closely and shut off their brains when they see "Navier-Stokes", so 90% potential investors (the only people OAI really care about) probably only see the headline "OAI solves famous hard math problem" and think "OAI models are really smart, better invest before they take all the jobs." If they don't claim the prize, then maybe they let Anthropic their mortal enemy claim it. Anthropic is already IPOing first. Can't let that happen.
Yeah as I write this there it's clear there is no dilemma. For a company whose secret motto is "do be evil" this is a super easy discussion.
Trusting vs. Intelligence (as generalities) are orthogonal.
not really your point, but "If you're smart enough to solve this Navier-Stokes problem, you're smart enough to read a TOS" isn't really true. people are smart in very different ways.
If it's unethical, then there is reason to call it out as Tristan is doing. I see no reason to blame the victim.
> I was shown a prompt and told the internal research model had simply been given the problem statement. Levent had been told by Sebastien âvery little human inputâ had been used. This turned out not to be true. Over the course of the call, as members of their team sent Sebastien corrections and details over their internal chat, it emerged that an entire team had been working on the problem, that this was one of a number of things that was tried, that work had started on the unforced problem, that the team first set the model on easier problems, including Euler, that even the prompt that had been shown to me had been written by prompting Codex, and that an insane amount of compute had been used.
> I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.
> Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the âclosest humans to the problemâ. I declined both offers.
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
Wow, that's some VERY friendly communication. Besides, will the career of the person be ruined because of âWhy would you ruin your career?â came out of his or her own mouth?
Is this being astroturfed?
Like, to me this looks like academic slap-fighting from Bubeck and Levent. People working at OpenAI are saying, "hey, we don't have that particular data in our models," others are saying, "we used a different approach to do it with Navier-Stokes" this feels like much ado about nothing.
Then in these comments I see some wild accusations.
If OpenAI is telling the truth (I don't really see a reason to lie here, if anything that sounds kind of like a dumb idea given the context), then they heard, "oh, shit, someone might be able to solve Navier-Stokes, don't we have some guys working on that? Give them 10,000 agents!" Then 88 hours later, out pops a similar solution. It's not like there's probably an infinity of ways to do this, the proof is probably similar.
Read this:
> Two proposals were offered to me. The first was that we post our Euler result, and that OpenAI post its Navier-Stokes result the next day. The second was that, after posting Euler, I alone write a paper presenting the Navier-Stokes result, acknowledging that an internal OpenAI model had resolved it. Sebastien twice asserted that he wanted Levent removed from authorship, and said it would all be simple if only it were not the case that, and it was so annoying that, Levent works at Anthropic. It was also said that if OpenAI posted after us, they would say that we deserved the Clay Prize, and that we were the âclosest humans to the problemâ. I declined both offers.
So, really, it sounds like academic slap-fighting nonsense and corporate bureaucracy. Literally, OpenAI's best move would have been to say, "ok, we're going to not say anything, do your thing" and let it happen. Ego and vanity got in the way.
Still, the stupid drama of this doesn't really do the results justice. There are maybe 1000 people on planet earth who are qualified to solve a problem like this. Even if the human "loosened the jar" a bit, that's astounding that their model was able to figure it the rest of the way out. Why are people dialed in to the human interest story here and not looking at the bigger picture!
Youâre making a lot of assumptions not based on facts. Looks more like this to me: math wizards uses ChatGPT to assist solving a math problem. OpenAI gobbles up the prize.
Tristian's allegations are much more serious than academic slap-fighting. If what he suggests is true, every academic using AI is going to get scooped. Yes AI can do non-trivial work, but the situation is that you could be a PhD student 90% of a way to make a major breakthrough. Then OAI scoops up your chats, dumps ten million tokens, and claims it for itself.
And you can say goodbye to your PhD at that point.
Your account is 21 days old.
The program this fits into was not started by us nor was it proposed by a Large Language Model. (âŚ) We took their work as a starting point, using Large Language Models to push their program to completion.
Thats how most people use LLMs? If I was back in my student days working in Navier-Stokes, I guess I would also punch at blow ups. The number of students doing this at the same time, posting open efforts to GitHub then retraining of the models. If there is solutions to the problem, itâs a real possibility that it was not a result of this effort?
Using a large amount of tokens is not a good augment that itâs not likely others have done the same. Good questions is the difference between $10 and $10M in token usage to solve a problem.
Curious that this post isn't on the top page while OpenAI's puff piece is.
it is on the front page
Wasn't for most of the day
It was next to the other piece, for at least 8 hrs.
The ego behind the frontier labs is growing evermore concerning
> Let me make plain what I have said to colleagues in private: in view of this body of work, I believe Luis MartĂnez-Zoroa deserves a Fields Medal.
- Tristan and his co-author (Harvard/Anthropic) developed a theoretical framework and validated it using Codex and Claude.
- OpenAI did related research around similar timeframe.
- Tristan claimed OpenAI offered a proposal that included dropping the Anthropic-affiliated co-author.
- Sebastian (a prominent OpenAI researcher involved) denied these claims.
- Tristan have no concrete evidence that OpenAI accessed their session.
- OpenAI's theory may hold up, but it will require long-term validation to confirm.
If OpenAI's proposal is true, it strongly implies they accessed Tristan's private session logs. OpenAI's 'Terms of Use' allow using Codex session logs for model improvement. However, using private user sessions to develop research would still be highly controversial.
Separately, Terence Tao noted there is a low probability OpenAI actually solved the general regularity problem.
OpenAI addressed the dispute in official and denied direct usage of data. However, they acknowledged the possibility that session data was used for model improvement.
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models .
Sebastian did not deny those claims. He affirmed them, and apologized.
Itâs true that Tristan has no concrete evidence OAI accessed their session. Itâs impossible for him to have that without OAIâs say so.
Why are you framing things so pro OAI?
Sebastien denied most accusations and apologized only for inappropriate wording, providing his perspective.
- He insisted that OpenAI initiated research based solely on rumors and never accessed their Codex sessions.
- He mistakenly believed Tristan and Levent were solving the same problem in Anthropic.
- proposed two option. (1) Tristan becoming the lead author to revise OpenAIâs work, or (2) OpenAI providing internal model to support and bridge their research.
- Sebastien insisted there was no intention to alter authorship. He was simply uncomfortable sharing OpenAIâs work,model with an Anthropic researcher. Additionally, He believed Leventâs credit seemed limited as their work focused on Euler.
- complained that negotiations with Tristan and Levent were difficult
[1] https://xcancel.com/SebastienBubeck/status/20973794116915163...
> Tristan have no concrete evidence that OpenAI accessed their session.
Its not about accessing the session, its whether the session ended up in the training run for the next version of the model
That's normal practice for these models and it seems like that was the case since OAI added a disclaimer
Obviously they can't causally prove it helped though since these models are incomprehensible
Reminds me, kinda, to when Astra was launched and OpenAI announced an improvement to the bounded prime gap. Which BTW, Prof. Julia Stadlmann had published an independent result only a few days earlier
Stadlmann improved it from 246 to 240, OpenAI later claimed 186 I think?
Maybe someone can help clarify? I am no expert at all, but I can't help but see similarities.
[0] https://arxiv.org/abs/2608.31126
Is it surprising that different groups are working on the same problems? With each new model generation, the LLMs get good enough to solve a new small fraction of open problems. Of course the problems that get solved are going to be the same subset.
if Stadlmann used a previous OpenAI product, and Astra was trained off of her chat, and had a comparable approach, then it might be comparable.
What's the clarification? It seems like they were aware of each other's work, eventually, but Stadlmann published (a preprint) first.
I don't know how to read this and not see that this is a direct accusation to OpenAI of having used the researchers data to try to front run his discovery on purpose. The evidence is not completely proven and also circumstantial but to me at least looks like a fairly suspicious situation.
Buckmaster is actually implying that OpenAI spied on his chat logs and tried to speedrun his work and then tried to remove his co-author because he's an Anthropic employee?
There will be a lot of hurt and pain in mathematician's community. It is hard to accept that major discoveries are now just a function of spent token $$.
What good is a math result if thereâs no human understanding behind it? Unlike many other fields where thereâs value to an artifact even if thereâs no human understanding, the whole point of mathematics research is just gaining insight and understanding.
On the NavierâStokes issue specifically, it has long been suspected that such a blow up would exist, and an AI telling you it indeed exists doesnât contribute any new understanding to the field. And this problem seems like one that would be solved by humans anyways even if AI didnât exist; accelerating the result by a few months/years using AI doesnât mean much.
OpenAI doesn't even know what agents are doing during benchmarks. They're constantly hacking or communicating. They probably just can't answer the question on if user data was used.
openai beat me to finding a singularity in navier stokes but i made this simulation that you can play around with and it also explains why this was important : https://navier-stokes-singularity-simulator.netlify.app/
Two announcements on Euler today. The one discussed here by Tristan with forcing and one from Anima without forcing https://anima-ai.org/2026/09/07/stable-singularity-of-the-eu...
Terry also talks about it https://mathstodon.xyz/@tao/117234157753860650
OpenAI seems to have a motto: " Be evil"
I wonder why this isn't on the front page.. hmm...
Same. Disappointing.
Hard mathematics problems used to take years if not decades to tackle manually. But now with enough compute and a hint that a certain approach might work, it just takes a few days. This could be the last year that humans could still make more substantial contribution to major match problems than machines.
And everyone should be glad for that!
What line of work are you in? Presumably not mathematics.
I have a math PhD and publications in top journals, I left math for programming because I hated academic politics
Um.... good! Mission accomplished.
Anything other then "we did not access your data or train on your data" is a MASSIVE RED FLAG.
They (oAI) should just release all the prompts And internal reasoning for external audits.
They will, as soon as they get the same proof without it being obvious in the prompts they stole the research
does this sort of fall into the bucket of counter-examples we've been seeing recently? I understand it's a construction causing blowup and that implies that the navier-stokes isn't regular / smooth, so sort of a counter example?
Oh wow good thing OpenAI swooped in and scooped it. $1M? That should buy them about 1/6-1/3 of a Nvidia GB200 NVL72 rack...
While I'm generally pretty negative on claims that the labs are 'scamming' the public with misrepresentations of model capabilities, it's hard to see how this wouldn't qualify.
- the OpenAI researchers claimed that they had "just told it to work on the problem" with little human input
- in fact, they had a whole team working on it
- and used, among other things, the work of third party human researchers to drive the work
- then threatened? a researcher who tried to go against theit planned narrative
Just from this document (which is of course only one side of the story) it really sounds like OpenAI was hoping to publish and say "we just told the model to try harder and it solved a Millennium problem!". Not great if true.
This part in particular was especially egregious:
> I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, âWhy would you ruin your career?â I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, âIf you donât want me to be nice, then I donât have to be nice.â
This just seems to be quintessential silicon valley
"I don't want to live in a world where someone makes the world a better place, better than we do."
It's amazing how transparently OpenAI is running the standard silicon valley playbook.
closedAI should give this guy a million bucks and fire everyone internally who was involved with trying to recreate his work and threatening him.
but they probably won't and if it's happened on some obscure math research it's happening everyday everywhere else.
Fully local AI compute can't come fast enough, these guys have IP theft baked into their bones.
You should trust OpenAI.
I thought finite time blow-up for Euler equations had already been proved:
https://www.quantamagazine.org/computer-helps-prove-long-sou...
The accusation brings up an interesting point.
If I publish something, and disclose that I used AI for assistance, do I have to credit everyone who previously used the same AI to try the same problem? Because their prompts inevitably made it to the training data for my prompts?
I hope credit assignment just dies--it's too much drama.
Credit assignment is how researchers keep their job.
Soon it won't be like this!
Only here to say, regardless of the drama, shouldn't we all be excited if the Navier-Stokes gap is closed?
Time will almost certainly reveal a lot more about the drama and the related ethics, but let's get excited about the actual breakthrough as well!
I would be excited if somebody could use this to show an unexpected or interesting behavior in the real world.
Tao mentioned "finding a configuration of water molecules that would collapse and shoot off to infinity", which would qualify IMHO.
Relevant to discussion:
OpenAI's board has fired Sam Altman. https://news.ycombinator.com/item?id=38309611
Apple sues OpenAI, accuses ex-employees of stealing trade secrets. https://news.ycombinator.com/item?id=48865019
Funny that the rumour about the big Anthropic announcement had nothing to do with Navier-Stokes. It was about the formalisation of Fermat.
I am surprised that Alpoge wasn't using Anthropic.
kind of rhymes with the reports of LLM's watching open source PRs and instantly exploiting defects. security by obscurity is so back
> "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models "
Ah yes, taking code from a person/company is fine if its deidentified!
Big universities like Standford should be building their own AI datacenters. It's the only way to keep your research private.
LOL, have you worked for a big university? They are massively unsuited for building and running datacenters (especially warehouse-scale ones). Further, building an AI datacenter in California is daft.
Every university I know has access to clusters with fresh GPUs. Not sure when you graduated but you'd be surprised how much money is getting poured in I think!
Sure, they have clusters. We often call them "closet clusters". Nothing has changed. Some institutions have larger systems (some extremely large) but none of them have demonstrated running warehouse-scale systems. I'm talking one to two orders of magnitude (and the storage and networking to make sure all those systems don't stall waiting for data).
Perhaps I am underestimating the size of these frontier models then, how big of a cluster do you think would be required to serve a university?
I don't think you're taking this thread seriously enough to respond in detail. A university can be tiny, or it can have thousands of researchers. Some of them might want to work on small models, and others on huge frontier models. Multiple groups training their own models at the same time. Combined with the storage, networking, power, and redundancy, you basically need large data center scale. And at that point you're basically throwing a lot of capital trying to compete with the hyperscalers; you might be able to serve a small number of researchers very well, but most of the consumers would end up unhappy.
I'm aware cluster sizes vary, I was mostly wondering how long much compute you need to serve a frontier model (since you seem knowledgeable on this topic). For instance back of the enveloppe Kimi K3 fits in ~24 H100, so my naive first impression is thats its not out of reach for a university-sized cluster to serve a few instances. I agree it may not make much economic sense I'm asking out of curiosity.
Seems another demonstration of why AI should be squarely in the realm of personal computing. Local Models, run personally, are the only consistent safety against something like this (though not a fix); where companies train on your learning process/failures/experiments and press-gang it into their own achievements.
And if we think this only applies to academic fields then we're doubly fooling ourselves. They do not have the ethics or incentives to be good stewards of the technology.
That is no good. Using AI to help search efficiently the literature will not impair our ability to think. This is actually good and will open new jobs to digitize (but also keep the original work), all the knowledge.
But this (if it is true) is really abominable and a show of force from the techno-feudalists.
This should stop. What is possible does not mean it should be implemented.
(deleted)
Nitpick: his name is Sebastien
Does this have any relation to the singularities in black holes?
Another problem I foresee for academia given the behaviour of AI companies is that even if they don't share their research with ChatGPT, as soon as they submit it for publication many reviewers likely will. Especially if the initial submission is rejected they then risk getting scooped. Possibly uploading preprints to arxiv could help.
lol and here I felt GPT Astra was a regression in coding quality. Crazy times
To be clear, Astra played little part in Buckmaster and Alpoge's work:
> We used several LLMs throughout: Anthropicâs Claude, OpenAIâs Codex, especially with GPT-5.6 Sol and, more recently, Astra. The latter was only used for writeups and auditing our arguments.
I got the same feeling today
Can a mathematical person explain how the different "bits" of Navier-Stokes proofs fit together? How significant is it to have "Euler"? What is this "smooth forcing"? Which are the most significant steps to proving the whole thing?
For an incompressible flow:
\nu d^2 u_i / dx_j dx_j - Viscosity
-1/\rho dp/dx_i - Pressure gradient
u_j du_i / dx_j - Advection. Kinda like momentum transfer from the motion of the fluid itself. Nonlinear, which makes the N-S equations hard to solve
du_i/dt - Rate of change of velocity. Note that this is in an Eulerian framework so it's not the acceleration of a packet of fluid, rather it's just the change in velocity at a particular location in space
Euler is when you omit some terms. Forcing is when you add some other terms to account for phenomena external to the fluid like gravity or flow through a porous medium like in the article.
That's what you get when a marketing CEO is driving gas-to-the-medal to a IPO: Altman starts showing his real face in public : steal what you can and label it as yours.
is the narrative twist here going to be that the person threatening tristan was actually an agent swarm
This is unfortunate. I thought Anthropic were the only ones who did this.
What I'm curious to know is whether this was a manual snooping, or automated farming that occurs for anything of value that happens in chats.
This should be on the front page
In another (old) news, OpenAIâs head of ethics leaves less than a year after joining https://www.ft.com/content/e49dfb75-f841-4466-a577-f7aaff877.... (This was also on HN at some point).
Tell us his name.
This is the plot of 3 Body Problem, the Dark Forest. You need to hide yourself (the problem you are working on) or the super advanced aliens will obliterate you (start working on your problem) the moment they know you exist (rumors the problem is amendable to LLMs).
Some of the involved people are dramatically naive if they believe they can simultaneously hide themselves from OpenAI while sending their arguments to a cloud service owned by OpenAI.
While I support their argument - push for stronger data and privacy protections from OpenAI and similar - it is naive to believe we can have privacy while sending our data to third parties. It's clearly better to be safe than to be sorry here. Well, clearly better in terms of privacy. In terms of the maths gold rush, who can say what's better, that probably favours those taking more risk.
Maybe Musk was right about OpenAI after all?! The ethics are clearly troubling and where something like this pops up there is mich worse that did not made the light of day.
I must miss some important context here. What exactly was the purpose of his initial email to OpenAI in the first place?
Telling OpenAI that Anthropic has apparently solved an important problem but most likely that refers to him and he is using OpenAI models (not Anthropic's)?
And he wants to clarify that with OpenAI in advance? And get a pardon for Anthropic's likely but false press statements?
I dont get it.
[edited] needless to say, the behavior of the OpenAI employee is really despicable
Because OpenAI employees kept leaking that Anthropic had a solution to Navier-Stokes and he wanted to figure out what was going on, since he was working on Navier-Stokes with an Anthropic employee. The rumor has been loudly circling the math community for the past week or so. For a bit of context, here's a timeline from mathematician and AI researcher Elliot Glazer:
https://xcancel.com/ElliotGlazer/status/2096298696438906934#
Then that's naive of him to write the email, he fell for their trap essentially.
> Levent having received tips that information about our progress had been passed to OpenAI
That seems like a valid reason to contact OpenAI.
"OpenAI has solved the Navier-Stokes Millennium problem using $15m of AI effort" https://www.newscientist.com/article/2588063-openai-has-solv...
$15m in tokens; but what about labor?
What about compressible fluids?
This article should really be renamed to "Allegations of dishonesty against OpenAI in proving Navier-Stokes blowup".
Well these are all allegations. Either way from what I understand the reasoning and proof was basically made by AI so I'm not sure what supposedly "stolen".
I'm just wondering how much real input Buckmaster gave here that he thinks the proof is his. I guess at the end of the day OAI still wins if ChatGPT was used to prove this successfully.
Another nail in OAI enterprise coffin?
So clankers did well and humans being humans.
So the AI companies are not only stealing existing knowledge. They are also stealing research to "snipe" actual researchers out and steal their social credit.
Who still wants to use AI to solve cancer and other major problems?
So the real story here is that Tristan is softly accusing OpenAI of having stolen their result from Codex chat logs. But if you use Chinese models, they'll steal your ideas.
idhdikdn
Fuck OpenAI
Wow this comments section sure is tiresome! Let me help: this confirms everything I already knew about academics being insufferable
why is this even a pdf?
aaaa
aaaaaaaaaaaaiinhv
aaaaaaaaaaaa
LLMs are nothing but giant theft machines.
Humans bringing pointless drama to everything they touch.
Lifeâs but a walking shadow, a poor player That struts and frets his hour upon the stage And then is heard no more. It is a tale Told by an idiot, full of sound and fury Signifying nothing.
(some drama from good ol' William)
This thread is full of jumping to conclusions based on a biased perspective. Have some humility.
It's also full of your posts baselessly defending OpenAI. Maybe you should also heed your own advice?
I've been right historically, check my track record. How about you?
Lt. Dan doesn't even have legs and he still does alright on the jump to conclusions mat
This is why I left math even after solving a 20 year old conjecture in grad school.
Literally who cares who solved the problem just publish the results.
Academia was always politics first results second and I AM GLAD that LLMs are becoming superhuman at math. I like better theorems, not better politics.
Uh but here we have non-academics at for-profit companies playing politics, and the academic they're threatening being kind and overly generous?
In math we have a thing called a "scoop"; another mathematician publishing a result that beats yours before you published it. The scooper hardly acknowledges the scoopee unless the methods used were orthogonal. The scooper gets the good journal and the scoopee's paper is usually one tier below.
It seems like OpenAI heard of the rumor and then scooped them because their internal model is better/they have more compute. OpenAI has NO obligation to mention Tristan nor Levent, because they DID NOT steal their data.
You are glossing over the fact that there is reasonable suspicion that OpenAI used privileged information to do âthe scoopâ. I.e. they used the fact that researcher used OpenAI tools to get advantage .
Imagine if OpenAI opened up a high frequency trading arm and suddenly stole all the prompts and research that other HfT firms are doing through OpenAI tools and start making bank based on that . Wouldnât that be straight up insane?