Kimi K3's Price War: What Two Weeks Actually Confirmed
Two weeks after Kimi K3 launched, here is what its pricing, GPU strain, and stock market fallout actually confirmed, and what did not hold up.

On July 16, 2026, Moonshot AI released Kimi K3, an open-weight AI model, meaning anyone can download its underlying files and run it themselves instead of only reaching it through a paid, closed service. The pitch was simple: performance close to the leading US models, at a fraction of the price. Two weeks later, that pitch has run into the real world in three separate ways. Moonshot’s own infrastructure buckled under demand within 48 hours of launch. Two of its Chinese competitors watched their stock prices fall by double digits the same week, though not entirely for the reason most coverage assumed. And a widely repeated claim about how cheaply Kimi K3 was trained turns out to describe a different model altogether. None of this makes Kimi K3 a bad product, but it does show what a low sticker price actually costs the company that sets it, and that story only becomes visible once the numbers from the same two weeks sit side by side. The launch-week hype itself has already been checked against Moonshot’s own benchmark claims; this piece stays on the price.
What Kimi K3 actually costs
AI companies charge for their models by the token, a small chunk of text the model reads or writes, priced per million tokens. Kimi K3’s official pricing, published directly on Moonshot’s own site, undercuts every major US lab on the cost of what the model writes back, which is usually the bigger expense for a heavy user.
| Model | Input ($/M tokens) | Output ($/M tokens) |
|---|---|---|
| Kimi K3 | 3.00 (0.30 cached) | 15.00 |
| DeepSeek V4 Flash | 0.14 (0.0028 cached) | 0.28 |
| Claude Opus 5 | 5.00 (0.50 cached) | 25.00 |
| Claude Fable 5 | 10.00 (1.00 cached) | 50.00 |
| GPT-5.6 Sol | 5.00 | 30.00 |
That $0.30 figure only applies when a prompt repeats content the model has already processed and cached; a fresh prompt costs ten times more, at $3.00 per million tokens. Even at the higher rate, Kimi K3 writes for a fraction of the price of Claude Fable 5, Anthropic’s flagship model, and less than half of GPT-5.6 Sol from OpenAI, while beating Anthropic’s smaller Claude Opus 5 on output cost too. It is nowhere near the cheapest model on the market, though: DeepSeek’s V4 Flash runs at a small fraction of Kimi K3’s price, for a smaller, faster model built for lighter jobs. Kimi K3 was never trying to win on price alone. It was trying to be the cheapest model at a size and capability that competes with the biggest labs’ flagship products, which is a harder claim to sustain, as the following two weeks showed.

Moonshot also sells Kimi K3 through a subscription product called Kimi Code, aimed at developers who want a flat monthly bill instead of a pay-per-token invoice. It comes in four tiers, from $19 to $199 a month, though the cheapest tier caps the context window, the amount of text the model can hold in one exchange, at 256,000 tokens. The full 1 million-token window Kimi K3 supports requires a pricier tier. Readers weighing Kimi K3 for actual coding work can find a direct coding test against Fable 5 for that specific use case.
The price Moonshot couldn’t actually deliver on day one
A low price only means something if the company behind it can serve it at scale, and Moonshot’s first two weeks show the gap between announcing a price and sustaining it. Around July 19, roughly 48 hours after launch, Moonshot paused new sign-ups for both the API and Kimi Code. In a statement Moonshot posted on X, reported by Yahoo Finance, PYMNTS, and Caixin Global, the company said demand for Kimi K3 had risen roughly sixfold in a matter of days, pushing its GPU capacity, the specialized chips that actually run the model, close to its limit. Moonshot’s own words were direct: “Kimi K3 has received far more love than we expected, and our GPUs are feeling it.”
Moonshot said it was testing two separate membership tiers to manage the load, a plan that lines up with the four-tier Kimi Code structure that appeared shortly after. Both products reopened for new signups within days, but the sequence still matters on its own terms. A frontier-scale model priced below the market rate drew more traffic than the infrastructure behind it could absorb on day one, and the bottleneck was capacity, not demand. Capacity is a much harder problem to fix overnight than a line on a pricing page.
The price war is hurting Chinese rivals more than it’s hurting Claude or GPT
The clearest sign that Kimi K3 rattled the market sits in the stock prices of Moonshot’s Chinese competitors, not its Western ones. In the week of July 17 to 20, shares in Zhipu, also known as Z.ai and listed in Hong Kong, fell as much as 30% intraday and closed down 19.56% on July 20 at HK$890.50, its worst single trading day since it went public in January, according to Quartz, Yahoo Finance, 36kr, and BigGo Finance. MiniMax Group, another Chinese AI lab, dropped up to 16% that same week and closed down 10.60%, a record low for the stock.
That second drop deserves a caveat the initial coverage mostly skipped. MiniMax’s decline coincided with the expiry of a lock-up period covering 46.44% of its shares, a scheduled event, set at the time of its stock market listing, that allows early investors to sell for the first time. It has nothing to do with Kimi K3. Reading the whole drop as competitive damage from Moonshot overstates the case for MiniMax specifically. Zhipu’s fall, which lacks a comparable structural explanation on the same day, is the cleaner signal of real competitive pressure between the two.

Neither Anthropic nor OpenAI showed anything close to this reaction, and the gap is not just Western investors shrugging off Chinese competition. By Artificial Analysis’s composite score, a benchmark that aggregates model performance across a range of tasks into a single number, Kimi K3 scores 57, ahead of Zhipu’s own flagship model GLM-5.2 at 51, and third worldwide, trailing only Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol. A closer look at how those benchmark scores hold up against the launch hype shows a real gap, not just a panic. Kimi K3 is not only cheaper than Zhipu’s own model, it also scores ahead of it, which is a harder thing for a direct rival to shrug off than a launch discount. For a company whose whole pitch is being the domestic alternative to the biggest labs, a rival that is both cheaper and measurably stronger on the same yardstick is not a headline to shrug off, and investors priced that in immediately rather than waiting for an earnings call to confirm it.
The training-cost number everyone’s repeating isn’t real
Some coverage of Kimi K3’s low price leaned on a second claim: that Moonshot trained a frontier-scale model unusually cheaply, echoing the efficiency story that circulated around Chinese AI labs in 2025. That claim rests on a mix-up between two different Moonshot models.
The only training-cost figure Moonshot has ever had attached to it publicly is $4.6 million, reported by CNBC, for Kimi K2 Thinking, a model the company released in November 2025, months before Kimi K3 existed. Even that figure is contested. Moonshot’s own chief executive, Yang Zhilin, has said the number “isn’t official” and that the real cost is hard to pin down precisely, according to Yicai Global. No comparable figure, official or otherwise, has been confirmed for Kimi K3 itself. Larger numbers that circulated in reactive coverage around the launch do not trace back to any primary source or reputable outlet.
The confusion is easy to make. Moonshot has shipped several models under similar names inside a year: Kimi K2 Thinking in late 2025, Kimi K2.7 as a separate, smaller open-source release, and now Kimi K3, a 2.8 trillion-parameter model that activates only a small slice of its internal components (16 out of 896, in a design called mixture-of-experts) for any single request, which is what keeps its running cost down despite its size. A training-cost figure that belongs to a different model, from a different month, does not describe what Kimi K3 cost to build. Treating it as if it does is how an unverified number turns into something people repeat as settled fact.
What it means for the rest of the AI industry
Two weeks is a short window, but long enough to separate what held up from what did not. Kimi K3’s launch price is real, confirmed directly from Moonshot’s own pricing page. The claim that it was also trained unusually cheaply is not confirmed for this model at all. And the price itself was not something Moonshot could serve at full scale on day one, a detail with more explanatory reach than this one company.
The pattern worth carrying forward is not that Chinese AI labs build cheaply. It is that a launch-week price is a marketing decision, not a guarantee of what it costs to run the service at scale, and the gap between the two shows up within days once real demand arrives. Every lab racing to undercut competitors on price, in China or anywhere else, is making the same bet Moonshot made: that its GPU capacity can be scaled up as fast as a pricing page can be published. For Moonshot, the bet held, but only after a temporary pause and a still-open question about managing demand at scale going forward.
The same test now sits in front of every lab that leans on a low sticker price to compete with Anthropic, OpenAI, and Google, from DeepSeek’s 2025 playbook to Zhipu’s own GLM models today. None of them get to skip the part where real demand tests whether that price was ever backed by real capacity.
The other lesson matters for readers as well as the industry. An unverified number, repeated confidently and often enough, becomes something people quote as settled fact, and that is exactly what happened here with Kimi K3’s training cost. Anyone comparing Kimi K3 against the wider field of open-weight models is better served starting from the pricing page itself, not the secondhand cost claims layered on top of it; a broader look at where Kimi K3 sits among today’s open-weight models is a useful next stop for that comparison.
Frequently asked questions
How much does Kimi K3 actually cost to use?
Kimi K3’s official API pricing is $0.30 per million input tokens on a cached prompt, $3.00 per million on a fresh one, and $15.00 per million output tokens, published directly on Moonshot’s own site. That undercuts Claude Fable 5 and GPT-5.6 Sol on output cost, though DeepSeek’s V4 Flash model remains far cheaper for lighter tasks.
Why did Moonshot pause new signups for Kimi K3?
Moonshot paused new API and Kimi Code signups roughly 48 hours after Kimi K3’s July 16 launch, after demand rose about sixfold and pushed the capacity of its GPUs, the chips that run the model, close to its limit. The company confirmed this directly in a public statement and reopened signups within days.
Did Kimi K3 really cause MiniMax’s stock to crash?
MiniMax’s stock fell 10.60% the same week Kimi K3 launched, but that decline coincided with the scheduled expiry of a lock-up period covering 46.44% of its shares, a structural event unrelated to competition. Zhipu, another Chinese AI lab, fell 19.56% the same week without a comparable explanation, making it the cleaner case of competitive pressure.
How much did Kimi K3 cost to train?
No confirmed training-cost figure exists for Kimi K3. The only number ever attached to a Moonshot model, $4.6 million, applies to Kimi K2 Thinking, an earlier and separate model, and Moonshot’s own chief executive has called even that figure unofficial.
Is Kimi K3 the same model as Kimi K2 Thinking or Kimi K2.7?
No. Kimi K3 is a separate, newer model launched in July 2026, a 2.8 trillion-parameter system that activates only a small fraction of its components per request. Kimi K2 Thinking is an earlier model from November 2025, and Kimi K2.7 is a distinct, smaller open-source release; the three are often confused in casual coverage of Moonshot’s lineup.