Another researcher accuses OpenAI of training on conversations and then claiming a breakthrough

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/post/3mv4mt4ikss2d

13 points · 9 comments · view on lemmy.world

9 Comments

Malica@lemmy.zip · 6 pts · 2d

I side with the nerds, there's no low Altman won't sink to

reallykindasorta@slrpnk.net · 5 pts · 2d

I like that both these mathematicians have reported their info via Mastodon. Maybe they’ve learned something. https://mathstodon.xyz/@andreasthom/117240535270608201

ZeroHora@lemmy.ml · 2 pts · 2d (6 replies)

I'm confused. The conversations are between the researchers and gpt? Or gpt is training on stolen data between e-mail/messengers?

reallykindasorta@slrpnk.net · 4 pts · 2d (5 replies)

In both the recent cases the researchers were purposely using AI products

ZeroHora@lemmy.ml · 3 pts · 2d (4 replies)

So is not kind obvious that OpenAI will use all the data that they gave to them to train the AI?

reallykindasorta@slrpnk.net · 2 pts · 2d (3 replies)

I mean I would have expected them to guess that but mathematicians can be oblivious. In both cases it sounds like they talked directly with reps at the relevant companies asking them if the models were trained on their data and the companies gave shady replies though, so part of the story is also the companies trying to obfuscate so they could claim their model found the result.

I expect researchers will begin to think more critically about the tools they choose.

TragicNotCute@lemmy.world · 1 pts · 2d (2 replies)

I don’t know the full story, but it’s worth saying that OpenAI alleges they do NOT train models when using API calls. I’d expected researchers to know that and pick the right tool accordingly. Maybe what this researcher is alleging is that even that statement isn’t true.

At OpenAI, protecting user data is fundamental to our mission. We do not train our models on inputs and outputs through our API.

https://developers.openai.com/api/docs/concepts

reallykindasorta@slrpnk.net · 1 pts · 2d (1 reply)

In the Buckmaster case

OpenAI now says in the Buckmaster-Alpöge case that no specific user data was accessed, but adds that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

You might have more insight into how to interpret that with the policy you mention than I do (I’m assuming they’re not misquoting ofc).

TragicNotCute@lemmy.world · 3 pts · 2d

I think ultimately it depends on how the researcher was using their products. If you have a premium ChatGPT account and you’re using the web or codex attached to that account, they are training on that data.

If they purchased API credits and are using it with the API I linked docs to, they shouldn’t be consuming any of that data for training.

Both ways of using the service are near identical in terms of output they can create, but that small billing nuance carries a big impact.

I’ve not seen a technical write up of this dispute that clarifies that point though.