My community is under attack by commercial interests. Since last week, speedrun.com refuses to distribute our database of speedrun accomplishments under the agreed-upon Creative Commons license. This licensing dispute is a major breach of trust and an abdication of SRC's responsibilities to the community. Individual game's communities are all scrambling to find a new home for their leaderboards. SRC has the audacity to say its for our own good.
My other community is also under attack by commercial interests. OpenAI and Anthropic are telling their models to solve open mathematical problems, presumably to evaluate their models' capabilities. Sure, if that is productive then I won't judge that. What I do disapprove of, is that OpenAI is yeeting all their exhaust onto the internet.
When you prove a theorem, you are imparted a responsibility to the field. At minimum you need to explain what you did, including which steps required new ideas and which steps are already well-known. More substantively, you are expected to nurture the literature. If an important idea was poorly explained in its first incarnation, then you have to explain it better in your work. These demands are broadly accepted and are not controversial. Any time you meet the demands, the resulting paper is highly appreciated and valued.[1] The responsibility is yours because you have the opportunity to publish a paper on the subject.
OpenAI is doing the exact opposite. They do not care about their 'manuscripts'. They don't bother putting in appropriate attribution of ideas, their work contains errors, and they don't even bother with consistent presentation or formatting. This actively sets back the state of the literature.
Why are they doing this? Is there a benefit to their releasing this exhaust? Does science benefit? OpenAI says these results took 3 hours of LLM time on average with their internal model. Likely their public models can achieve the same outcomes when instructed by expert guidance. That's why they keep scooping their own customers.[2] A scooping that, I will add, is only possible because OpenAI outputs such sloppy work.
If an expert prompts a result, they take up the mantle of responsibility I describe above. OpenAI not only refuses to take their responsibility, but they prevent their own customers from doing so. And they have the audacity to say they do it for the good of science.
| [1] | One example of such a valued paper is Daniel and I's smoothed analysis paper. Yes we proved better running time bounds, but mostly using ideas that were in the literature already. The contribution of the paper was that it was nice to read, unlike the notoriously opaque literature that came before it. It would have been difficult to justify the time we spent writing this all up nicely if we hadn't also improved the result quantitatively, hence the responsibility. The paper remains my most visible piece of work, still accruing more citations per year than later follow-up work. |
| [2] | Here is one example from today of OpenAI customers who got scooped because they were spending their time writing up a nice paper instead of staking a flag on github dot com. |