That's not what your Chief Research Officer, Mark Chen, says on X:
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
"If they opted out of training, then we definitely did not train on them."
Per OpenAI's privacy policy, they use de-identified data to improve their products. From Mark Chen's comment, improving products includes improving ChatGPT and Codex in a holistic way. Improving models in a holistic way sounds a lot like training to me.
> Per OpenAI's privacy policy, they use de-identified data to improve their products
That’s not inconsistent with what you responded to. They use your data unless you opt out. If the user doesn’t opt out, their de-identified data is used to improve their products.
It appears than you can only opt-out from having OpenAI train models on your data. There isn't an option for opting to exclude your de-identified data from being used to improve OpenAI products.
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company." - Mark Chen, Chief Research Officer, OpenAI.
My understanding is that even if you opt out but then press thumbs down or give other feedback you are implicitly or explicitly or whatever giving permission to them to look at that chat alone.
I have opted out from data sharing, and when Claude asks me for feedback on a session it then asks if its OK to share that data with Anthropic.
I'd assume an opt out is an effective opt out. An opt out that is ignored by Anthropic is a breach of contract, not something they would do casually, esp. given the high turnaround and animosities between their own employees and ex-employees - and the labs. All it takes is one pissed off whistleblower to open a can of worms.
Occam's razor applies. The mathematician did not opt out from data sharing. OpenAI vacuums up all such data into training data sets. If OpenAI genuine does not easily know if a given session went into the actual training data set its probably due to the complexity of the data pipelines - not everything ends up impacting the model weights, after all.
It looks like clicking "don't train" might not matter. Per Mark Chen, the Chief Research Officer at OpenAI,
"Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company."
He doesn't make a claim that implies "don't train" might not matter in the way you might think.
They cannot rule out that the data was trained because they have no per-user provenance tracking through the training pipeline once data is de-identified..
The entire point of de-identification is the inability to know the source of data. If the researcher forgot to hit "do not train" then that's that..
The only thing they have is a coincidence and the fact that the LLM may have used the training data that then researcher technically may have agreed to share.
Whether or not that's smoking gun of anything is hard to say. And the fact may remain that the proofs are significantly different, we do not know.
For folks that are thinking this is just semantics or qualia, the site mentions this:
"We can also see activation in the visual cortex during fMRI studies, the area in the brain that processes images from the eyes, further suggesting actual visualization."
Well, all that says is that something that is activated by seeing something can be activate by other means (and that some lack this). What does it mean to "see something" in the mind's eye? For me it triggers a similar response to if I saw it, but I certainly do not see it. I see a lot of people that believe they have aphantasia just because they don't actually see things when they imagine them and this "aphantasia test" should be truthfully answered by "everyone" that they do not see it.
On the other hand, I can do hypnagogic meditation where I am on the verge of falling asleep, and something like visual dreams start before actually falling asleep, and in that case I can actually see things as if I saw them with my eyes, and I can look at details and so on.
On the other hand, when asked to visualize a doctor I can't!
And that's why I'm not sexist (lol), when asked gotcha questions all imagined doctors are sexless lines on an unwritten page that exists in the soupy pre-semantic peircean firstness of my brain.
A boy and his mother are in a serious car accident. The mother dies on scene and the boy is rushed to the hospital where he needs urgent life-saving surgery. But the male surgeon says "I can't operate on this boy - he's my son!" How is this possible?
Maybe I misunderstand, Why is it an arguments against qualia? Qualia doesn't deny a neural substrate, just that qualia isn't reducible to the neural activity, being that one is objective and the other is subjective.
It's sort of the difference between logical supervenience and metaphysical, or weak emergence and strong emergence. Qualia being something additional to the neural description.
Here's my problem with this: take an fMRI of a bunch of professional jugglers while they imagine juggling. Now take an fMRI of a bunch of non-jugglers while they imagine juggling. I promise you they will be strikingly different, with the professionals having all sorts of real-life coordination areas lighting up.
Conclusion: some people are a-juggular and are wholly incapable of imaging juggling. Poll: are you a-juggular or not!?
I hope this illustrates how we're missing a very sensible third option here: visualizing is a skill that some people are better at than others, but (almost) anyone can learn.
Learning to juggle involves easily observed objective activities that other people can check, and which can be learned that way--you can watch someone else juggling, read book that describes the tossing of the balls, etc.
How do you do any of that with visualizing? If I'm trying to learn to juggle, I can watch someone else juggling and say, I want to do that. How do I watch someone else visualizing and see what to do?
Yes, it's harder to evaluate whether someone can visualize well, but that does nothing to invalidate my point.
Have you actually dedicated significant time to trying to learn how to visualize? I suspect that if you did, you'd realize: oh, wow, maybe I can actually get better at this. If most people who give it a serious effort also come to that conclusion, I feel like we've pretty much settled the question. So it is testable, just not as easily.
> Have you actually dedicated significant time to trying to learn how to visualize?
I haven't had to. I don't have aphantasia. Visualizing doesn't feel like a learned skill to me. It feels like something I've always known how to do. Of course if I pay more attention to what I'm doing, I can visualize in more detail than if I just do it casually. But I've never had the feeling that, oh, if I practice this and learn I can get better at it, the way, for example, that I can get better at playing the piano by practicing.
Where I'm stuck is, if someone else who couldn't visualize asked me to teach them how, I wouldn't know what to do. Whereas, if they asked me to teach them how to play the piano, I'd at least have some idea.
If we go by "fool's errand" as "needless or profitless endeavor", https://arxiv.org/abs/2303.11156 may already be a good enough answer to the paper you cited, so their work is already laid out for them. The green token idea is thoroughly attacked with much more effective techniques than those in the original paper through "recursive paraphrasing". Among some hypotheses in the paper, one is particularly interesting:
>These experiments provide empirical evidence that more advanced LLMs can lead to smaller TV distances. Thus, based on Theorem 1, reliable AI text detection would become increasingly difficult
https://x.com/markchen90/status/2097400166554993041
reply