A second mathematician challenges OpenAI over whether ChatGPT interactions entered its training data.
Mathematician Andreas Thom accused OpenAI of dishonesty over its handling of training-data questions. Thom asked OpenAI researchers whether his prior ChatGPT interactions entered the company’s training data or remained available to its reasoning process. OpenAI’s answer did not satisfy him because it addressed only direct access to conversations, not whether those conversations became training data.
Tristan Buckmaster, a math professor at New York University, publicly questioned whether OpenAI’s models benefited from his use of OpenAI Codex. After OpenAI announced its non-sofic groups result, mathematicians criticized the company for failing to acknowledge contributions from Thom and Gábor Kun. OpenAI amended its writeup. Non-sofic groups are infinite mathematical structures that finite ones cannot approximate.
For builders and operators, the dispute highlights a governance gap around training data. Researchers say they cannot reverse-engineer OpenAI’s training pipeline to determine whether their work influenced a model. Only OpenAI holds the relevant data, so the company bears responsibility when it denies using user interactions. Enterprise teams face similar provenance questions as customers demand clearer records.
Thom and other mathematicians want OpenAI to provide evidence about whether their interactions with ChatGPT or Codex entered training data or influenced reasoning. OpenAI has not published a detailed account of how it handled the non-sofic groups result beyond amending its writeup. The next signal will be whether OpenAI answers Thom’s request, discloses more about data provenance, or faces challenges from researchers.
What matters
- Andreas Thom says OpenAI failed to explain whether his ChatGPT chats fed its models.
- Teams that train models on user interactions face new pressure to document data provenance.
- Watch whether OpenAI publishes training-data details or responds to Thom’s request for evidence.
Why it matters
Watch whether OpenAI publishes training-data details or responds to Thom’s request for evidence.
This GenAI News article was prepared in original wording using reporting and materials published by The Verge. Source reference: https://www.theverge.com/ai-artificial-intelligence/993263/where-does-openai-get-mathematics-training-data.
Drafted by the GenAI News review pipeline.
