AI Distillation Isn't Illegal. So Why Do OpenAI and Anthropic Call It Theft?
AI model distillation is a legal, widely used technique. Here's what actually crosses the line, and why OpenAI and Anthropic keep calling it theft anyway.

Every few months in 2026, a new AI model out of China has topped a benchmark chart that used to belong to OpenAI or Anthropic, and every few months, the same accusation has followed it. American labs have repeatedly said the model in question was trained by quietly siphoning answers out of their own products. The technique behind that accusation, AI model distillation, is not new, not secret, and not illegal. It is standard practice inside every major AI lab, including the ones doing the accusing. What actually separates a legitimate training method from an accusation of theft is much narrower than the headlines suggest. By mid-2026, three separate cases have each drawn that line in a slightly different place: OpenAI against DeepSeek; Anthropic against DeepSeek, Moonshot AI and MiniMax; and Anthropic against Alibaba’s Qwen.
What distillation actually is
Knowledge distillation was formalized in a 2015 paper by Geoffrey Hinton and colleagues at Google, Distilling the Knowledge in a Neural Network. The idea is simple once it’s unpacked: instead of training a smaller AI model from scratch, you train it to imitate the outputs of a larger, more capable model. The big model becomes a “teacher,” the small one a “student,” and the student learns faster and cheaper by studying the teacher’s answers instead of raw data. Every major lab, OpenAI, Anthropic, Google, Meta, does this constantly, distilling its own large flagship models down into smaller, cheaper versions it can sell at lower cost. There is nothing improper about the technique itself. Distillation as a method has never been the problem, and no one, including Anthropic and OpenAI, argues otherwise.
Where “legal” stops
The dividing line isn’t the technique itself, but where the training data for that student model comes from. If a company distills its own model, or trains on data it’s licensed to use, that’s ordinary engineering. The trouble starts when a company trains a rival model by automatically, and at scale, pulling huge volumes of answers out of a competitor’s open-weight model or paid API, using those answers as training data without permission. OpenAI, Anthropic, Mistral and xAI all explicitly ban this in the terms of service that come with every API account, the fine print almost nobody reads before clicking accept. That clause, not any copyright statute, is the actual legal hook behind every accusation described below. It makes the dispute a matter of broken contract terms, closer to a lease violation than a courtroom copyright fight, and that distinction matters more than it sounds like it should.
February 2026: Anthropic’s 24,000-account report
The best-documented case so far came first. In late February 2026, Anthropic published a detailed report accusing DeepSeek, Moonshot AI and MiniMax of running an industrial-scale distillation campaign against Claude, its own AI model. The company said it had identified roughly 24,000 fraudulent accounts responsible for more than 16 million exchanges with Claude, all aimed at harvesting training data. The traffic wasn’t evenly split. MiniMax accounted for the overwhelming majority, more than 13 million of those exchanges, according to Anthropic; Moonshot AI’s share was smaller, over 3.4 million exchanges. DeepSeek’s share was the smallest of the three and the most targeted, around 150,000 exchanges, which Anthropic said were focused specifically on reasoning and alignment, the safety tuning that keeps a model from answering requests it shouldn’t. In plain terms, Anthropic alleged DeepSeek wasn’t just copying Claude’s answers wholesale, it was probing how Claude’s guardrails worked in order to route around similar guardrails of its own. The report was covered within a day by Bloomberg, CNBC, TechCrunch and CNN. None of the three companies named has been sued over it.

The OpenAI-DeepSeek case, headed to Congress instead of court
A second, separate case runs alongside the Anthropic report and is often confused with it, because it involves the same Chinese lab. On February 12, 2026, OpenAI sent a memorandum to the U.S. House Select Committee on the Chinese Communist Party, arguing that DeepSeek uses increasingly sophisticated distillation methods to train its models. In that memo, OpenAI described programmatic code and what it called “obfuscated third-party routers,” software built to disguise where automated requests to American AI models were really coming from. OpenAI was not among the witnesses who testified at the committee’s subsequent April 16 hearing on China’s AI distillation campaign. This is an allegation made to lawmakers, not a lawsuit. OpenAI has not filed any legal action against DeepSeek over this claim. It is a political and regulatory move, aimed at shaping how Congress thinks about export controls and AI competition with China, and it should be read as exactly that rather than as a settled finding of wrongdoing.
June 2026: a bigger campaign, this time against Alibaba’s Qwen
Three months later, Anthropic went back to Congress with a new accusation, and this one dwarfed the February report. In June 2026, Anthropic told congressional committees, including the Senate Banking Committee, that it had identified an operation linked to Alibaba’s Qwen models running roughly 25,000 fraudulent accounts and 28.8 million exchanges with Claude between April and June 2026, the largest such campaign the company says it has documented to date. Anthropic said this traffic concentrated heavily on Claude’s coding and agentic reasoning abilities, the parts of the model most valuable to a competitor trying to catch up quickly. Tom’s Hardware reported the scale of the allegation in detail. Alibaba has denied the accusation. As with the February report, no lawsuit has been filed, and no court has evaluated the evidence behind either side’s claim.

July 2026: Kimi K3, and the same accusation from a new accuser
The pattern showed up again within weeks, this time voiced by the U.S. government rather than by a rival lab. In July 2026, Moonshot AI’s newest release, Kimi K3, one of the largest open-weight models ever published, drew a public accusation from the U.S. government. Treasury Secretary Scott Bessent warned on July 21 that sanctions could be on the table over suspected distillation from U.S. models. The next day, Michael Kratsios, the White House’s top science adviser, made it specific. Kratsios said on social media that Moonshot had distilled Anthropic’s Fable model, launched weeks earlier, to build Kimi K3. Moonshot was already one of the three companies named in Anthropic’s February report, responsible for over 3.4 million of the 16 million exchanges Anthropic tallied then, which makes this less a new incident than the same dispute resurfacing under a new accuser. This round also drew a sharper pushback than the earlier ones: independent AI researchers questioned whether the timeline even worked, since Fable had only been publicly available for about two weeks before Kimi K3 shipped. Nathan Lambert, a researcher at the Allen Institute for AI, put the broader skepticism plainly: distillation explains less and less of a model’s performance as the underlying models keep improving. Each time a Chinese lab ships a model that performs unexpectedly well, the same claim resurfaces, and each time it arrives as an accusation, not as a ruling from anyone with the authority to settle the question.
Why “theft” is doing more rhetorical work than legal work
Every one of these cases sits at the accusation stage. Anthropic has published detailed numbers twice. OpenAI has submitted testimony to Congress once. Neither Anthropic nor OpenAI has filed a lawsuit, and none of the four accused labs, DeepSeek, Moonshot AI, MiniMax or Alibaba, has been found to have done anything by a court. Alibaba has actively denied the claim against it. Calling any of this “theft” in the strict legal sense gets ahead of where the facts currently stand: theft, in law, usually requires establishing that something was property in the first place and that it was taken without right, and the legal status of an AI model’s own outputs is still unsettled in most courts, so nobody can yet say with certainty whether they can be owned the way a photograph or a song can. What these companies can point to with more confidence is a contract violation: an API terms-of-service clause that plainly bans automated bulk extraction of outputs, allegedly broken at industrial scale. That’s a real, potentially costly legal exposure on its own. It just isn’t the same thing as a verdict of theft, and treating an accusation as a proven fact skips the part of the process where evidence gets tested by someone other than the company making the claim.
The takeaway
The technique that keeps making headlines was never the actual dispute. Distillation is legal, decades old in spirit, and used by every lab named in these stories, including the ones filing the complaints. What separates ordinary engineering from an accusation is a single contractual line: whether the training data came from scraping a competitor’s API against its own rules. By mid-2026, that line has been crossed, according to the accusers, in three overlapping cases involving DeepSeek, Moonshot AI, MiniMax and Alibaba, each documented with more or less detail and none of them tested in court. Until one of these cases actually reaches a judge, “theft” describes how American labs want the story framed, not a fact anyone outside those companies has yet had the chance to verify.
Frequently asked questions
What is AI model distillation, and is it illegal?
AI model distillation is a training method where a smaller “student” model learns by imitating the outputs of a larger “teacher” model, a technique formalized by Google researchers in 2015. The method itself is completely legal and used internally by every major AI lab. What can become a legal problem is how the teacher’s outputs were obtained, not the training technique itself.
Why do OpenAI and Anthropic call it theft if no court has ruled on it?
Calling AI distillation theft is a framing choice, not a legal finding. Anthropic and OpenAI have documented what they say is automated, large-scale extraction of their models’ outputs in violation of their own terms of service. That’s a real accusation with real numbers behind it in Anthropic’s case, but neither company has filed a lawsuit, so no court has tested the claim.
What exactly did Anthropic accuse DeepSeek, Moonshot AI, and MiniMax of doing?
In February 2026, Anthropic reported roughly 24,000 fraudulent accounts and over 16 million exchanges with Claude used to gather training data: more than 13 million through MiniMax, over 3.4 million through Moonshot AI, and about 150,000 through DeepSeek, whose smaller share focused specifically on probing Claude’s safety behavior. Anthropic called it an industrial-scale distillation campaign; none of the three companies has been sued over it.
Is the Alibaba and Qwen case the biggest one so far?
Yes, Anthropic’s June 2026 accusation against operators linked to Alibaba’s Qwen models is the largest documented so far, citing roughly 25,000 fraudulent accounts and 28.8 million exchanges with Claude, more than the February report against DeepSeek, Moonshot AI and MiniMax. Alibaba has denied the claim, and no lawsuit has followed.
Could a case like this ever actually go to trial?
A trial is possible, but hasn’t happened yet in any of these cases. It would require one of the accusing labs to sue over the alleged terms-of-service violation, and a court would then have to decide questions current law hasn’t clearly settled, like whether an AI model’s outputs count as property that can be stolen. Until that happens, every case described here remains an accusation, not a verdict.