This story is someone on bluesky screenshotting someone on twitter without a link. The person on twitter screenshotted the source on the fediverse without a link. This is getting daft.
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."
There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".
I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
We are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing.
Or at least: tell me I should be careful/worry about those particular things.
Personally, I wouldn’t assume it was lying. To me, dark patterns (like manipulative wording) imply that:
1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest.
And also:
2) someone in the c-suite or marketing has decided to mitigate that through some dark pattern.
If they never intended to give the user that control, they’d probably just not give the user the control in the first place. Giving them the option and not honoring it would either imply they were that incompetent or careless with fate protection, which seems most likely if this is all true, or an even more cynical approach to tricking users into thinking it’s not used for training to get them to share better shit. But if that was the case, why bother with the sleazy dark pattern? That seems a little cartoon villain-y to me. I suppose the in-between option would be that they decided at some point they were no longer willing to honor it and didn’t want to deal with the inevitable PR shitstorm of removing it. I could definitely see that happening in this industry, these days.
> To me, dark patterns (like manipulative wording) imply that:
Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.
The author of the original mastodon post, Andreas Thom, acknowledged that he had not opted his data out of being used for training until June 29 of this year. He spends most of the post lashing out at OpenAI for not being transparent about whether his data was trained on (when the answer is obviously yes).
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.
If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
Hopefully OpenAI can solve good problems where no one can make a claim that they stole their idea where the methodologies used and the techniques used are so out of the ordinary that the achievement is respected. Furthermore they should work on new novel solution on the old problems to lay these issues rest, thus by improving their models they can avoid issues of academics accusing them of using their work additionally the academics should also demonstrate their unpublished work is significant enough to have solved the problem . This is a bit subjective but it is also objective for the person with domain knowledge.
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
also, none of it means anything without the lawyers to back it up. Just like you can be a pedophile in the highest office of democracy and escape persecution.
Reminder that there are degrees of "trained on conversations". From John Schulman:
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
The most successful companies of the last decade have precisely been ... selling usage data.
Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.
Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.
I mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal.
Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.
It seems that in this case OpenAI are suggesting that the researchers whose work they scooped were using OpenAI models with an account setting that allowed OpenAI to train on anonymized prompts.
It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
If Sam Altman tells you the sky is blue, you should double check. I certainly hope nobody believes him when he claims controversial things from which he stands to benefit.
Considering how Apple today announced that its products will
contain a spying-on-conversation anti-feature by default, it
seems reasonable to assume that all this spying is primarily
done to train their AI model; and secondarily also to gather
information about The People.
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
> in this case too they didn't actually have the solution
Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.
I would almost expect training to overweight conversations with novel scientific and mathematical implications.
> the only reason they can't definitively say no is that for privacy reasons
They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.
Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
> that opted-out user data was used for training
Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.
Original posts
https://mathstodon.xyz/@andreasthom/117240535270608201
https://mathstodon.xyz/@andreasthom/117240536885387540
https://mathstodon.xyz/@andreasthom/117240537520615623
This story is someone on bluesky screenshotting someone on twitter without a link. The person on twitter screenshotted the source on the fediverse without a link. This is getting daft.
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
https://news.ycombinator.com/item?id=49643556
Quoting:
> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."
I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
https://news.ycombinator.com/item?id=49643513
How do people become that trusting?
The phrasing itself is guilt tripping
Or at least: tell me I should be careful/worry about those particular things.
https://x.com/thsottiaux/status/2097746417012166816
1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest.
And also:
2) someone in the c-suite or marketing has decided to mitigate that through some dark pattern.
If they never intended to give the user that control, they’d probably just not give the user the control in the first place. Giving them the option and not honoring it would either imply they were that incompetent or careless with fate protection, which seems most likely if this is all true, or an even more cynical approach to tricking users into thinking it’s not used for training to get them to share better shit. But if that was the case, why bother with the sleazy dark pattern? That seems a little cartoon villain-y to me. I suppose the in-between option would be that they decided at some point they were no longer willing to honor it and didn’t want to deal with the inevitable PR shitstorm of removing it. I could definitely see that happening in this industry, these days.
Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.
https://mathstodon.xyz/@andreasthom/117240535270608201
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
source: https://x.com/johnschulman2/status/2097440545853637108
I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.
The most successful companies of the last decade have precisely been ... selling usage data.
Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
Welcome to the party, with the rest of humanity.
Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.
It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
AI is becoming more evil by the day.
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.
I would almost expect training to overweight conversations with novel scientific and mathematical implications.
> the only reason they can't definitively say no is that for privacy reasons
They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.
Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
> that opted-out user data was used for training
Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.