AI Slop, LLMs, Plagiarism, Decline of the Internet

Two recent stories about mathematicians using LLMs to assist with their work for months, then OpenAI using LLMs to solve what they were working on and publishing first, potentially by using their drafts and other work as training data:

If nonpublic research supplied by users improved a model and the provider then used that model to race those users to publication—without informed consent, disclosure, or credit—that would be ethically indefensible. De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.

I asked when the first prompt had been sent by them. This question was not answered directly by OpenAI for some time. Eventually it was agreed that it had been sent in the past few days, after information about our work had reached OpenAI.

I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer.

and

I said that if OpenAI released its result in the way proposed I would go public with what happened. The reply was, “Why would you ruin your career?” I replied that I am an academic, and asked why he thought going public would ruin my career. The reply was, “If you don’t want me to be nice, then I don’t have to be nice.”

Something similar can now happen with CF and a ton of other ideas too. LLMs have trained on public CF essays. A user could talk to an LLM about philosophy and reinvent a lot of CF, with the LLM being very helpful and supplying a lot of the ideas, without the person knowing they were plagiarizing. CF is already published but isn’t well known enough to stop someone else from potentially getting credit for lots of it. Also, I do have draft additions to CF that I’ve worked on and talked with LLMs about.