Hacker Newsnew | past | comments | ask | show | jobs | submit | hellohello2's commentslogin

You guys should seriously offering a clear way of working with (semi-)confidential data for particulars. Regardless of what is actually done internally, toggling off an opt-in isn't reassuring enough, which is why people are having these worries.

Option 1 is opting out manually. Option 2 is business / enterprise plans, which opt out by default.

Any ideas of things we could do to make it clearer?


Option 2 feel reassuring enough, but is out of reach of particulars.

Option 1 is not. In part because it is opt-out (will it turn back on on its own like my Facebook privacy settings?), and not always respected (sending feedback can mean your chat is used?). Also because disabling "Improve model for everyone" is very vague.

There simply needs to be a setting like "my data is confidential", in which case there clear guarantees like there are for ZDR.

As an example, I've seen people speculate that while input prompts and output tokens are discarded, thinking traces are retained for training, which could leak information. I doubt this is true, but it shows that the policy is not unambiguous and reassuring enough to remove all doubt.

Thanks for asking.


I was thinking about this some more, and perhaps the best solution for subscription plans would be to charge more for real privacy. In which case breaking that privacy would be committing fraud. Just a thought.

Generative models copy training data verbatim and also generalize, the two are not mutually exclusive.

You can have a look at the literature on exact copying in image models if it interests you, but just online we often see online examples of agents outputting code that already exists, even if its not the common case.

I very much doubt OpenAI points the model towards a specific conversation, but these trillion parameter models can very much "remember" their training data. For instance I can ask GPT to summarize my papers from their title alone, without looking them up, and it works decently.


Of previous, not concurrent work. Science is friendly competition, and spying on others is unfriendly.

To offer a counterpoint, Airpods Pro are a great product that significantly increased my quality of life. Depending on how sensitive you are to noise then having easily accessible/socially acceptable noise cancelling is incredible. Perhaps you are too much of a normie to get it ;)

Every university I know has access to clusters with fresh GPUs. Not sure when you graduated but you'd be surprised how much money is getting poured in I think!

Sure, they have clusters. We often call them "closet clusters". Nothing has changed. Some institutions have larger systems (some extremely large) but none of them have demonstrated running warehouse-scale systems. I'm talking one to two orders of magnitude (and the storage and networking to make sure all those systems don't stall waiting for data).

Perhaps I am underestimating the size of these frontier models then, how big of a cluster do you think would be required to serve a university?

I don't think you're taking this thread seriously enough to respond in detail. A university can be tiny, or it can have thousands of researchers. Some of them might want to work on small models, and others on huge frontier models. Multiple groups training their own models at the same time. Combined with the storage, networking, power, and redundancy, you basically need large data center scale. And at that point you're basically throwing a lot of capital trying to compete with the hyperscalers; you might be able to serve a small number of researchers very well, but most of the consumers would end up unhappy.

I'm aware cluster sizes vary, I was mostly wondering how long much compute you need to serve a frontier model (since you seem knowledgeable on this topic). For instance back of the enveloppe Kimi K3 fits in ~24 H100, so my naive first impression is thats its not out of reach for a university-sized cluster to serve a few instances. I agree it may not make much economic sense I'm asking out of curiosity.

Hm. Looks like its possible everyone behaved terribly here unfortunately. :/

I remained impressed by ChatGPT however!


Everyone?? No, most definitely OpenAI.

But they have learnt their lesson, next time they won't reach out to who they stole it from, they will publish first.


Seems like OpenAI did a boring normal corporate thing (find out your competitor made a breakthrough, try to replicate it) and then when the other mathematicians found out OpenAI had beat them to Navier-Stokes, they decided to lie about what happened because they were upset they didn't get to make the big breakthrough themselves.

You are entitled to your opinion, but it isn't supported by the information already available.

The available information is essentially that story.

OpenAI's account: they heard a rumor that a Millennium Prize problem had been solved, so they tried to do it themselves and succeeded. Then they contacted the other researchers and were surprised to discover those guys hadn't actually cracked it, but offered the one of them who's not an Anthropic employee a co-authorship anyway. The conversations got testy.

Buckmaster's account: totally unsubstantiated accusations of plagiarizing from chat logs and plainly false accusations of OpenAI trying to get Alpoge removed as coauthor of a thing he was not an author of in the first place, and threats to ruin people's careers.

I think the synthesis is basically that Buckmaster and Alpoge had not quite solved the Navier-Stokes problem yet but thought they were really close, and had told friends as much, which is how the rumors got out. Now they're mad they got scooped. They aren't getting the money and recognition they thought they had locked down, and are engaging in a smear campaign.


I don't think your reading is correct. ML history is related with examples of models using hints from side channels.

If it would be possible to replicate OpenAIs success with a model cut-off earlier than the rumours, and without hinting from informed mathematicians, then we can consider it original. Otherwise it's indefensible.


By Buckmaster's own account, OpenAI (a) wanted him and Alpoge to publish their Euler result first and (b) wanted Buckmaster to author the Navier-Stokes paper, and (c) when he refused both of those, tried to get Alpoge to talk him into it.

Why would OpenAI go burn another $20 million trying to prove they didn't accidentally train on synthetic data derived from Buckmaster's codex sessions? They're desperately trying to give him credit anyway!


Are you certain of this? I would be inclined to believe you but it would be nice to know decisively.

But there is evidence, the blog post says: "While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models ."

In other words, yes, they had been using ChatGPT, and yes, ChatGPT could very well have trained on their data. Now that there is evidence, we need an investigation: yes or no, was it the case?


That is not an admission of malfeasance though? As I read it they don't know if anyone fed relevant private documents into the model under an account configured to permit training on user data.

If there's more to the story I'd be interested to hear it.


Of malfeasance no, but they could have easily plagiarized unintentionally. If you commit mansalughter, you still need to explain yourself, even if it was a complete unlucky accident.

So you're saying that they could have committed manslaughter, but acknowledge that we have no evidence that they did. So why should they need to explain themselves? Isn't is on the aggrieved party to bring evidence?

I don't believe people are denying that the model is impressive. The problem is that learning someone else is making progress on a topic using method X and then rushing to scoop them borders on academic misconduct. If, on top of this, their private conversations about X were used in the proof, I really don't see how its defensible...

The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype.

If I were at OpenAI, I'd naturally want to snipe that from them. I am completely unsurprised they formed a crack team to steal Anthropic's glory, and do so in just five days.


Between companies, direct malevolent competition is OK. Between academics, there are other rules to the game. When you go into a boxing match, you agree to get punched in the face.

All this to say, trust is important, and grounded in social convention. So I do agree with you, but also disagree.

Whenever this is OK or not really depends on how the breakthrough is contextualized, and how there people at play, here, agree to contextualize it.

In my view, in the blog post, there is much discussion about who will be publishing the paper. If instead it was just a blog post that said "oops, we beat you to it, our model is the best", it would have been different.


This is a common misconception, so its understandable that you have it. Generative models can both plagiarize and generalize. The question here is which of the two happened.

A needlessly condescending tone while failing to address the topic at hand. The person I replied to advanced the claim that training was sufficient to constitute plagiarism. You appear to be claiming that it is possible to generalize instead of plagiarize after training on something, so I take it that you must necessarily disagree with the original claim?

What I meant to say is that, in many cases, a generative model's output is not in fact steered by minor amounts by lots of training samples, but instead steered by a just few samples. Some outputs are influenced by many inputs, and some by very few, it really depends.

In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized. This is not the case, no, because generative models do not "copy" or "create", they do both at different times.

I did not agree or disagree with the original poster, I was explaining to you why I thought you disagreed with them. If you understand what I said above, then why do you disagree with them?

EDIT: I just saw your other post on "general inspiration" and I believe I read the situation exactly; you appear to believe that inputs used to train generative models get "lost in the parameter soup", but it is not always the case.


> In answer to a post suggesting that training on a datapoint could mean plagiarism, you said that this would imply that all outputs are plagiarized.

We read the original differently. As clearly stated in my previous reply to you, I interpret it as claiming that all outputs are necessarily plagiarizations of the training data. That is not my claim (as you wrongly stated) rather it is the claim I am responding to. I observe that it is absurd to object to a single action being a transgression on the basis of an argument which implies that all actions are inherently transgressions. Notice that nowhere do I take a position on whether or not the argument about all actions being transgressions is true or false.

> you appear to believe that ...

I do not, no. I have not taken a position of my own here. I've merely objected that the one I responded to does not make for a sensible line of argument in context. It seems that you (and many others) have read my objection to position A as support for position B and attempted to infer what I think from that.


I simply do not see how you can interpret "if they do not deny training on them, they can't deny plagiarism" as "all outputs are necessarily plagiarizations of the training data".

There is a difference between claiming an action is a transgression, and claiming it could be one.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: