- The Deep View
- Posts
- The catch behind Kimi K3's benchmark leap
The catch behind Kimi K3's benchmark leap

Welcome back. AI’s duopoly is having long-term consequences by leaving most countries dependent on US and Chinese AI models and infrastructure. Meanwhile, creative resistance to AI may be softening as filmmakers and creators embrace tools that cut tedious work, lower costs, and expand what they can make, even as questions about originality and consent remain unresolved. And Kimi K3 is triggering claims that it has overtaken the leading models from Anthropic and OpenAI. While its leap forward is impressive, let's keep in mind that it's only in one benchmark for now, and the model could still largely be distilled from the latest US frontier models rather than being a research breakthrough. —Jason Hiner
1. Beware of grandiose Kimi K3 victory claims
2. How AI may be gradually winning over creatives
3. Most nations risk becoming AI dependents
PRODUCTS
The catch behind Kimi K3's benchmark leap
One of China's leading AI companies may have leapfrogged the industry's frontier labs.
Moonshot AI recently released Kimi K3, the latest iteration of its flagship model family, and one benchmark claims it has outflanked Anthropic, OpenAI, Google, and others.
Kimi K3 features 2.8 trillion parameters, a one-million-token context window, and a multimodal architecture that comprehends text, images, and video within a single model. The full open weights will be available on July 27, and for now the model is usable via Moonshot's API.
But Moonshot's release may represent a change in the tides:
According to Arena.ai, a benchmarking and evaluation platform for frontier models, Kimi K3 has taken the top spot in 6 of its 7 frontend domains, including brand and marketing, reference-based design, data and analytics, consumer product, simulations, and content creation tools.
The model surpassed the performance of the latest models from Anthropic, OpenAI, xAI, Meta and China's Z.ai in these categories. Kimi K3 only ranked second in the gaming domain, slotting just behind Anthropic's Claude Fable 5.
The score represents a 17-place jump from the company's previous model generation, Kimi-K2.6, according to Arena.
The results have sent shockwaves through the industry in recent days: Gavin Baker, managing partner and CIO of Atreides Management, said in a post on X that the release marks an "inflection point" for AI that could spell trouble for Anthropic and OpenAI by breaking up their dominance. Cisco CEO Jeetu Patel posted that improving models like these could lead to more competitive industry dynamics by creating a gross margin drop for frontier models. David Sacks, co-chair of President Trump's Council of Advisors on Science and Technology, said in a post that the release is "concerning" in light of calls to regulate AI in the US and that "this is how you lose the AI race."
It's important to note that Arena's benchmark results differ from those that Moonshot itself released. In its post announcing Kimi K3, the company said that the model sat just below Fable 5 in benchmarks for coding, agents and frontier software engineering, and only outranked Anthropic's model in benchmarks for codebase cleaning and long-horizon engineering.
It's also important to note that Moonshot is one of three Chinese labs, along with DeepSeek and MiniMax, accused of "industrial-scale" model distillation campaigns by Anthropic in February, taking part in more than 16 million exchanges with Claude through 24,000 fraudulent accounts to acquire and use its capabilities.
"Moonshot is marketing Kimi K3 with new architecture work unique to its lab, but that could be used in conjunction with distillation techniques as part of training a model," Brian Jackson, principal research director at Info-Tech Research Group, told The Deep View.

Though major industry voices are sounding the alarm on Moonshot's advancements, the natural tendency in such a fast-paced industry is to get excited by every shiny new toy that hits the market. However, Kimi K3 may not be evidence enough that China is pulling ahead. Though the nature of open-source technology broadly has always been to take the work of another and improve upon it, given the model distillation accusations by closed-source, proprietary labs, we need to question whether or not models from these major Chinese labs are powerful on their own accord, or simply powerful because they are piggybacking off of the research of the leading frontier AI labs.
TOGETHER WITH GRANOLA
The best AI notetaker is the one you barely notice
Most AI meeting tools are too visible: bots join your calls, notifications interrupt you, and afterwards you're left with a pile of transcripts.
Granola was built differently. It's an AI notepad that stays in the background so you can stay present in the meeting. You write down what matters, and Granola quietly captures the rest, turning the conversation into summaries, next steps, and context you can actually use.
You can even chat with your notes to write follow-up emails, prep for other meetings, or turn conversations into actions. Giving you, not your bot, a little boost.
CULTURE
How AI may be gradually winning over creatives
Creative professionals were among the first to speak out against AI. Now, that resistance may be fading.
George Lucas, creator of Star Wars and Lucasfilm, recently compared AI to the horse and buggy in an interview. He acknowledged the risks. Just as cars, born from that shift, eventually brought breakdowns, fuel dependency, and even weaponization in the form of tanks, he argued that this is simply the nature of progress.
"Artificial intelligence means it’s much easier for us to make movies,” said Lucas. "There’s nothing you can do about it. That’s progress. It’s the future."
He isn't the only creative to hold that opinion. A new Adobe report that surveyed 16,000 creators across the US, UK, France, Germany, South Korea, Japan, India and Australia found that 87% of respondents using creative AI reported it had accelerated the growth of their business or audience, while 75% described it as integrated or essential to how they work.
The use cases include moving faster in the ideation phase:
Around 93% of creators say creative AI helps them produce content faster, though 57% say their creative AI outputs typically require moderate or extensive editing before they're ready to share.
Even with the additional work needed, 35% say it gives them more freedom to experiment before pitching ideas, and 33% say it gives them the confidence to pursue more ambitious ideas and projects.
At Adobe Summit in April, Forest Key, VP of agentic AI for the creativity and productivity business at Adobe, who previously worked at Lucasfilm in the 1990s on Star Wars, told The Deep View that some of his most time-consuming tasks could now be done easily by AI.
"When I worked at Lucasfilm, we were tracing humans with these little splines and animated curves frame by frame," said Key. "40, 50, 70, 80 hours to do like a three-second shot. With Premiere and its AI power tools, you can do what would take hours and hours and hours literally in a minute."
He adds that applications of AI of that nature don't kill the creative process, but rather accelerate creativity. Meanwhile, on its latest earnings call, Netflix said that nearly 300 titles used AI for "highly complex sequences" that would otherwise not have been possible due to time or cost constraints, according to co-chief executive Ted Sarandos.

A lot of creative work involves tedious tasks that don't rely on creativity at all, making AI assistance a natural fit. However, the slippery slope is that relying on AI to help with tasks that involve even the slightest bit of creativity can disrupt the thing that makes art special: Because AI models are trained on existing work of other artists, using AI to assist in the actual act of creativity can diminish its uniqueness and originality. To further complicate the issue, these models often don't ask artists for permission before using their work as training data, leading to allegations and lawsuits against major AI firms like OpenAI, Suno, Midjourney and more. Though some models, such as Adobe Firefly, compensate all of the artists whose work is used to build the model, not every company abides by those same standards.
TOGETHER WITH OUTSKILL
Most People Use 5% of Claude. Master the Rest in One Weekend
You've probably seen this in your headlines over the past few days: the US government has forced Anthropic to pull its two most powerful models, Mythos and Fable, offline. But do you know why?
Because these models became so powerful that the US government deemed them a national security risk. But here's what people are gatekeeping from you: not using Claude to its full potential is a threat to your job security.
Claude ships new features almost every week:Skills, Connectors, Cowork, Vibe coding
So we found the perfect workshop for you, completely free, that condenses 800+ hours of research and real-world practice into a focused 16-hour curriculum. Introducing the 2-Day Claude AI Mastery Workshop: a live, end-to-end deep dive into Claude plus 10+ AI tools, LLMs and workflows.
It would be silly not to SIGN UP
- Saturday & Sunday
- 10 AM – 7 PM EST
MARKETS
Most nations risk becoming AI dependents
While the US and China go head-to-head on AI advances, other nations are struggling to grow their AI footprints.
A report from Forrester released last week found that technology sovereignty, or a country's ability to develop, run and secure new technologies independently of foreign influence, will grow slowly over the next five years. While the US and China have the highest tech sovereignty scores according to Forrester's index, sitting at 79% and 82% respectively, the other 12 countries assessed had an average score of 39%, relying heavily on imported AI and tech infrastructure.
In addition to the US and China, the countries surveyed in Forrester’s Global Sovereignty Forecast include Australia, Brazil, Canada, France, Germany, India, Italy, Japan, Mexico, South Korea, Spain, and the UK.
The index assesses several dimensions of tech sovereignty: government AI investment, cloud, workforce, AI model development, data center capacity and autonomy, chip production, software creation and rare earths processing.
On average, these countries are forecasted to see tech sovereignty rise by only 1% by 2030, from 39% to 40%, while the US and China are expected to maintain their leads.
The only dimension that shows promise in increased sovereignty is chip production, with South Korea, Japan and India seeing significant growth in independent production.
The fact that the US and China hold global dominance in AI and tech broadly presents a number of problems, Dario Maisto, principal analyst at Forrester, told The Deep View. For one, as organizations and nations seek more AI services, they will likely increase their dependence on foreign vendors. Feeding their data to these models could then "pose a problem for data appropriability and IP defensibility," said Maisto.
Furthermore, the US-China duopoly makes choosing an AI vendor more than just a matter of model capabilities. It is also a political decision, he said. It opens the problem of "tech diplomacy." He said, "By choosing one model or the other, for example, an organization might signal political closeness to one country or the other."
"Enterprises’ and nations’ choices will not be driven solely by price and performance of the AI technologies and solutions, but also by the possible retaliation measures of one country because they chose the technology of the other," Maisto added.

There are solutions for the problem of global dependence on two overwhelming powers, such as nations investing in homegrown talent and models and eventually restricting access to foreign models, Maisto told me. However, that scale of an investment could take years and may not even be economically sound. Because of this, nations are largely left with few options other than to feed into the US-China duopoly, treading carefully around the political weight of the models they choose. The situation echoes what we've seen in the private markets among AI companies themselves: a massive well of power held by a small pool of companies. Still, some companies seek to prevent this by creating more options, taking it upon themselves to feed into the open source AI market, including Nvidia's Nemotron family, Google's Gemma family, Thinking Machine's latest model Inkling, and the pool of open models from labs like DeepSeek, MiniMax and Moonshot. Though this at least creates some market fragmentation, these models primarily still come from the US and China.
LINKS

Meta in talks to lease compute power to Anthropic worth $10 billion
Z.ai on track to achieve $1 billion in annual recurring revenue
DeepMind's Demis Hassabis to lobby Washington for AI standards body
Nuclear startup Valar in talks to raise $1 billion at around $5 billion valuation
AWS customers receive $1.5 trillion bill in a global pricing glitch
SpaceX, Defense Department discuss data center capacity deal

Manus: Added skill to turn a single prompt into a professionally typeset PDF
NotebookLM: Is now renamed to Gemini Notebook but offers same capabilities
Google Vids: Now features Gemini Omni and personal avatars
ChatGPT desktop app: Was updated to reflect user-requested features
Gemini API: New cost controls for managed agents, including a free tier

Alpaca: Support AI analyst
Fanatics: AI engineer
Cargomatic: AI Engineer
Gamma: AI Engineer
POLL RESULTS
Do you think that EU AI regulation will have global ripple affects?
Yes (49%)
Somewhat (31%)
No (11%)
Other (9%)
The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“A flock of birds generally flies at the same elevation, as shown in this image. ”
|
“In [this image] the birds were not close and were flying in different directions. It didn't represent true murmurations.”
|


If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.












