If one of the researchers forgot to check "don't train on this" a single time, the pretrain is going to know what they were working on. The question of whether they remembered to check "don't train on this" every single time is super concrete, and embarrassing to _both_ sides if the answer is "no, he didn't read the eula."
These models will absolutely remember a brilliant insight that appeared a single time in the pretrain corpus, because if they couldn't-- they would get a slightly worse loss. I don't know what this wishy-washing "well maybe we trained on it but we didn't read it" is supposed to mean.
Just an FYI, despite the wording on this, my "don't train on this" checkbox never changes (I don't need to re-toggle it, or change it on every device, etc...). This post reads as if its something that needs to be done repeatedly. Also if you turn on the advanced security option, ChatGPT locks to no training and cannot even be toggled on if you wanted it to.
Everyone I think assumes training in the general “it affects the distribution” sense. But if it fishes precise novel insights out as alleged then it makes these products dramatically less valuable. At least the non zdr ones
This is a problem with every technology product. Before Facebook and Gmail had strong institutional controls, any employee could just open anybody's email and private messages on a whim. Now things are a bit more strict for random employees, but if the company as a whole wants to do something, you bet they can.
Can you opt out of sharing data for “analytical purposes”? I think you can’t. And what does “analytical purposes” even mean?
I think people straw man the whole data argument on “opt out of training purposes” and ignore the mandatory analytic purposes. Which probably includes figuring out trends and important research problem data to steal from users and even business ideas.
I don't think there is as much IP theft risk on the analytical side. That to me implies they're tracking behaviour (what you click etc) more than taking the raw data
It looks like clicking "don't train" might not matter. Per Mark Chen, the Chief Research Officer at OpenAI,
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
He doesn't make a claim that implies "don't train" might not matter in the way you might think.
They cannot rule out that the data was trained because they have no per-user provenance tracking through the training pipeline once data is de-identified..
The entire point of de-identification is the inability to know the source of data. If the researcher forgot to hit "do not train" then that's that..
The only thing they have is a coincidence and the fact that the LLM may have used the training data that then researcher technically may have agreed to share.
Whether or not that's smoking gun of anything is hard to say. And the fact may remain that the proofs are significantly different, we do not know.
It's possible that OpenAI was mining prominent researcher's chats for inspiration to tackle these problems, but it's also entirely possible the leak came from his collaborator's end as he works at Anthropic and I'm sure there's plenty of corporate espionage going on between those two. That would also explain why they didn't want to comment on the source of the prompt.
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things.
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)."
https://openai.com/index/navier-stokes-solution/
The fact that this link is a Reddit thread should probably be taken as evidence that we don't have a clear picture of what happened yet, and speculation is rampant.
Alternatively, they might train in a different manner than what you expect.
They might have determined that your canaries (whether it is a unique string, URL or similar) have a low signal-to-noise ratio and not worth following up.
Unless you provide more details, it's difficult to ascertain whether your conclusion is correct.
If one of the researchers forgot to check "don't train on this" a single time, the pretrain is going to know what they were working on. The question of whether they remembered to check "don't train on this" every single time is super concrete, and embarrassing to _both_ sides if the answer is "no, he didn't read the eula."
These models will absolutely remember a brilliant insight that appeared a single time in the pretrain corpus, because if they couldn't-- they would get a slightly worse loss. I don't know what this wishy-washing "well maybe we trained on it but we didn't read it" is supposed to mean.
(Example of Fable knowing the content of a deeply unimportant LessWrong post I wrote: https://claude.ai/share/a907b46c-bf7b-4fca-9c71-8582cf8507cc a working Navier Stokes solution would be way more salient)
Just an FYI, despite the wording on this, my "don't train on this" checkbox never changes (I don't need to re-toggle it, or change it on every device, etc...). This post reads as if its something that needs to be done repeatedly. Also if you turn on the advanced security option, ChatGPT locks to no training and cannot even be toggled on if you wanted it to.
I don’t use ChatGPT anymore, but when I did, the training opt-out setting was frequently silently resetting itself to off.
This is a good FYI. I guess my microsoft trauma is showing.
If true it could have quite an impact.
Everyone I think assumes training in the general “it affects the distribution” sense. But if it fishes precise novel insights out as alleged then it makes these products dramatically less valuable. At least the non zdr ones
This is a problem with every technology product. Before Facebook and Gmail had strong institutional controls, any employee could just open anybody's email and private messages on a whim. Now things are a bit more strict for random employees, but if the company as a whole wants to do something, you bet they can.
Can you opt out of sharing data for “analytical purposes”? I think you can’t. And what does “analytical purposes” even mean?
I think people straw man the whole data argument on “opt out of training purposes” and ignore the mandatory analytic purposes. Which probably includes figuring out trends and important research problem data to steal from users and even business ideas.
I don't think there is as much IP theft risk on the analytical side. That to me implies they're tracking behaviour (what you click etc) more than taking the raw data
It looks like clicking "don't train" might not matter. Per Mark Chen, the Chief Research Officer at OpenAI, "Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
https://x.com/markchen90/status/2097400166554993041
You might be misreading him.
He doesn't make a claim that implies "don't train" might not matter in the way you might think.
They cannot rule out that the data was trained because they have no per-user provenance tracking through the training pipeline once data is de-identified..
The entire point of de-identification is the inability to know the source of data. If the researcher forgot to hit "do not train" then that's that..
The only thing they have is a coincidence and the fact that the LLM may have used the training data that then researcher technically may have agreed to share.
Whether or not that's smoking gun of anything is hard to say. And the fact may remain that the proofs are significantly different, we do not know.
It's possible that OpenAI was mining prominent researcher's chats for inspiration to tackle these problems, but it's also entirely possible the leak came from his collaborator's end as he works at Anthropic and I'm sure there's plenty of corporate espionage going on between those two. That would also explain why they didn't want to comment on the source of the prompt.
OpenAI's release explicitly says No. But then also caveats that with "we cannot rule out that de-identified data derived from their usage of our products" impacted things.
What's most striking to me, and what may or may not be true, is the "we cannot rule out" bit.
"We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced)." https://openai.com/index/navier-stokes-solution/
That disclaimer could mean maybe Astra inferred your theory from the pattern of clicks on the Accept button.
Anyways, the moral of the story is: Not your GPU, not your data.
They are currently being sued for trade secret theft so it seems likely.
The fact that this link is a Reddit thread should probably be taken as evidence that we don't have a clear picture of what happened yet, and speculation is rampant.
I'm surprised solutions weren't trained off their Office365 or Google Docs..
[dupe] Discussion: https://news.ycombinator.com/item?id=49605915
And currently: https://news.ycombinator.com/item?id=49613262
Yes.
A rare exception to Betteridge's law of headlines
100% their data was stolen. You can't enter any private data into AI and expect it to remain private.
I have put hundreds of canary strings into my conversation history of all the big providers, and so far the 'hit rate' is very low.
Really looks like they aren't yet training on conversation histories effectively.
Alternatively, they might train in a different manner than what you expect.
They might have determined that your canaries (whether it is a unique string, URL or similar) have a low signal-to-noise ratio and not worth following up.
Unless you provide more details, it's difficult to ascertain whether your conclusion is correct.
Can you explain this a little more?
[dead]