Mathematicians are pressing OpenAI for greater transparency about whether researchers’ private interactions with its products may have indirectly contributed to mathematical results later announced by the company. The dispute centers on a distinction between direct access to identifiable conversations and the possible use of de-identified product data to improve models.
The Verge reports that mathematician Andreas Thom raised concerns after OpenAI announced a result involving non-sofic groups, an area connected to work by Thom and Gábor Kun. OpenAI acknowledged that its result relied heavily on their earlier work, and later amended its write-up after criticism that recent contributions had not been properly recognized.
Thom said the company demonstrated a detailed command of techniques he and colleagues had been exploring. He asked OpenAI researchers whether his ChatGPT interactions had entered training data or were available to the reasoning process. According to The Verge, Thom considered the response insufficient because it addressed direct access without conclusively explaining whether information from product use could have entered broader model-improvement pipelines.
His concern follows questions from New York University mathematics professor Tristan Buckmaster, who had been working with Anthropic researcher Levent Alpöge in a personal capacity. Their work became part of a separate controversy surrounding OpenAI’s claimed solution to the Navier–Stokes Millennium Prize problem. Any such solution would require independent mathematical verification before acceptance.
In discussing that work, OpenAI said its researchers and agents did not see the other researchers’ work before it became public and that no specific user data was accessed to solve the problem. The company also said it could not completely exclude the unlikely possibility that de-identified data derived from product usage had contributed to model improvements.
That qualification is at the heart of the current disagreement. Removing names or account identifiers may protect personal identity, but researchers argue that it does not necessarily remove the intellectual substance of an unpublished idea. Thom says only OpenAI has the records needed to establish whether relevant conversations entered its systems and has called for disclosure of datasets and settings that could resolve the issue.
No evidence cited in The Verge’s report establishes that OpenAI used Thom’s or Buckmaster’s private work to produce a mathematical result. The researchers’ complaint is instead that the company’s public assurances do not provide enough information for outsiders to rule that possibility out. OpenAI did not immediately respond to The Verge’s request for comment on Thom’s latest criticism.
The dispute illustrates a growing problem for research conducted with commercial AI tools. A chatbot can function as a notebook, collaborator or coding assistant, but its data controls may have consequences for unpublished work. If researchers cannot determine how their interactions are used, they may become less willing to test early ideas in these systems.
For OpenAI, resolving the concern will require more than celebrating strong results. Clear, auditable explanations of data handling, attribution and model-improvement practices could become essential to maintaining trust with the scientific communities whose work its systems analyze and extend.



