• The Deep View
  • Posts
  • Why Chinese models may win AI’s efficiency race

Why Chinese models may win AI’s efficiency race

Welcome back. Google is turning one of the Pixel 11’s best features into a platform. By opening the model behind Rambler, Google is giving developers the capability to bring polished, conversational dictation to virtually every device. Salesforce and Anthropic are weaving Claude across Salesforce and Slack so enterprises can inject intelligence into some of their most used platforms. And China’s Z.ai revealed itself as the lab behind the viral Ox Alpha model, showing that capable AI does not need frontier pricing. Jason Hiner

IN TODAY’S NEWSLETTER

1. Why Chinese models may win AI’s efficiency race

2. Salesforce taps Claude as enterprise AI layer

3. Why Google's Rambler could come to every device

MARKETS

What Z.ai's Ox Alpha reveals about AI economics

Chinese AI lab Z.ai has come forward to claim the viral anonymous Ox Alpha model. 

On Wednesday, Z.ai unveiled GLM-5.3-Flash, the first natively multimodal model of the GLM-5 series. The company said the model was released anonymously as Ox Alpha on OpenCode and OpenRouter where it completely overtook leaderboards and went viral for offering a capacity for 100 trillion tokens per day. 

Notably, the company said in its announcement that all of the traffic from its model's skyrocketing popularity was "served on Chinese AI chips." 

The model features 320 billion total parameters with 18 million active, and outperforms its previous generation, GLM-5.2, across a number of benchmarks at a tenth of the price, the company said in its announcement. 

The lab also claims that GLM-5.3-Flash approaches Claude Opus 4.8 in coding and agentic benchmarks. It also performs on par with DeepSeek-V4, GPT-5.6 Terra, and Gemini 3.7 Flash on benchmarks for software engineering, multi-step agentic tasks, tool use and professional work. 

The company said the model is architected for "extreme efficiency," and "specifically designed for ultra-low-cost inference." 

  • GLM-5.3-Flash was essentially designed to do more with less, the company said, using a hybrid attention architecture to reduce the cost of serving long-context queries without sacrificing accuracy. 

  • Additionally, the model is built to improve "scaling efficiency" by implementing what are called "Manifold-Constrained Hyper-Connections," or a technique that improves the way that information moves through the layers of a neural network. 

  • The model weights are currently available on Hugging Face. The model is also powering ZCode, Z.ai's coding tool. 

"GLM-5.3-Flash shows that frontier intelligence does not have to come at frontier cost," the company said in its announcement. "We are now scaling this recipe to larger models — GLM-5.3-Flash pushes the cost-performance frontier, and the lessons from building it are already shaping our next frontier model." 

Z.ai's Ox Alpha win comes at a particularly poignant moment for open models: Open source model platform Hugging Face is reportedly fielding acquisition offers that value the company at $13 billion, and Nvidia last week announced a $6 billion investment in bolstering the open source AI ecosystem in the US. 

But Chinese firms may be going after one enterprise pain point in particular: token costs. Models from Chinese labs that perform on-par with those from proprietary providers may be enticing to developers that are now facing budget constraints after previously being encouraged to spend freely on AI, Karthik Sj, chief AI officer at LogicMonitor, told The Deep View. 

"Ox Alpha demonstrates how low- and zero-cost model access may sound appealing at first, but free access always carries hidden costs," said Sj. "While Z.ai revealing that it’s behind the model solves some auditability concerns, vetted and well-tested models are the safer bet.”

Though it's obvious that Z.ai is targeting cost and resource efficiency with its release of GLM-5.3-Flash, this model's popularity also points to another trend: not every daily use model needs to be state-of-the-art. GLM-5.3-Flash still sits on-par or just behind the frontier models from labs like Anthropic, Google and OpenAI, however, a model doesn't need to be ultrapowerful to be incredibly useful. Though OpenAI is trying to entice people to use its frontier models with lower prices, it can only serve that kind of inference at those costs for so long. For a large majority of enterprise tasks, generally, mid-range performance is good enough, especially when it comes to solving for ROI. 

Nat Rubio-Licht

TOGETHER WITH CRUSOE

$5 in free credits, on us.

Every Crusoe account now comes with $5 in free credits for Crusoe Intelligence Foundry. Use them for Serverless Inference or Serverless Fine-Tuning, whether there's a new open model you've been wanting to try or a workload you're ready to run.

No new commitment, no setup required. Credits apply automatically to your account and are ready whenever you are.

ENTERPRISE

Anthropic-Salesforce deal could break AI silos

As Anthropic cements its place as a go-to AI provider for enterprises around the world, the company is deepening its ties with another enterprise heavyweight: Salesforce.

On Wednesday, Salesforce and Anthropic unveiled Claudeforce, the name given to the companies' expanded strategic partnership. Claudeforce combines Claude's intelligence with Salesforce's enterprise platform to offer agentic experiences for working professionals. 

The selling point is adding Claude AI's reasoning layer on top of the Salesforce platform, giving the AI access to the workflows, data, and insights that companies already have stored in Salesforce. 

  • The partnership is launching with Salesforce in Claude, a plugin that includes 37 pre-built sales skills, aimed at helping both agents and sellers apply AI optimization to their workflows directly in Claude, using context from Salesforce, according to the release. 

  • The skills were built by both Anthropic and Salesforce to take it one step beyond "generic (customer relationship management) prompts wrapped around an API," to better fit what Claude can do. 

  • Some examples provided include reasoning over live revenue context or automating pipeline updates. 

The setup is meant to be simple: An admin just has to connect Salesforce to Claude once, with authentication and permissions managed centrally, and every seller on the team gets access, circumventing the need for per-user access, according to Salesforce. Salesforce in Claude is available to select pilot customers and will launch in open beta in September 2026. 

Claude is also available inside Salesforce, including Agentforce, where it serves as a

reasoning model for the Atlas Reasoning Engine, as well as powers Agentforce Vibes and Agentforce Coworker by default. Through Amazon Bedrock, users can also access Claude within the Salesforce Trust Boundary. 

Finally, through the partnership, Claude will become the default model for Slack, powering  every instance of AI in Slack, including Slackbot and the @Claude tag, and is a founding partner for Slack Code. The companies noted that more integrations are coming across Claude, Salesforce, and Slack, with additional pre-built skills launching later in 2026.

Salesforce and Anthropic both emphasized in the announcement that the agreement involves reciprocal adoption, with each company becoming a customer of the other. While this is partly marketing rhetoric, I see genuine value in the arrangement: it helps both companies build solutions their customers actually want, since each experiences the other's product from a user's perspective. More broadly, I think the AI industry would benefit from greater collaboration between companies, rather than siloed development aimed purely at competing with one another, an approach that often fails to leverage each company's strengths.

Sabrina Ortiz, Senior Reporter

TOGETHER WITH CELONIS

Don't bet your business on a fortune cookie

Eliminate operational blind spots. Give AI the context it needs to succeed.

AI models are brilliant, but they’re probabilistic. Without a ground truth of your operational reality, trusting AI is like trusting a fortune cookie.

The Celonis Context Model provides this operational reality through a dynamic, real-time, system-agnostic digital twin of your operations. It translates your business into a language that AI understands.

Give your AI the context it needs to drive meaningful ROI. Meet the Context Model.

PRODUCTS

Why Google's Rambler could come to every device

One of the most talked-about Google AI features from the Pixel 11 launch was Rambler, its highly effective speech-to-text feature. Now Google is spreading the tech that powers it. 

On Wednesday, Google introduced Gemini 3.5 Transcribe, which the company describes as its most precise speech-to-text model yet, and capable of converting raw audio into accurate, polished text. Google shared that this is the model that powers Rambler on Android and the Gemini app on macOS, and it invites developers to build similar capabilities with the model. 

Moreover, Gemini 3.5 Transcribe is available across two separate APIs: gemini-3.5-transcribe-live, which is focused on continuous bidirectional real-time interactions with sub-second latency and gemini-3.5-transcribe, which is focused on pre-recorded audio processing and transcribes recorded audio, meetings, and call logs with speaker attribution and word-level timestamps.

What makes the model and the Rambler experience stand out is how accurately it can capture your ideas, even when they aren't presented linearly. For instance, during a correction, when one says "actually never mind how about…," it should not account for that; same with "umhs" and "ahs." Other features include:

  • The model can delegate complex tasks to other Gemini models via function calls

  • It recognizes jargon and unique spellings unique to the speaker

  • It automatically transcribes over 85 languages 

This progress is reflected in benchmarks, with Gemini 3.5 Transcribe outperforming its previous transcription model, Chip 3, by 70% on time to final transcription as measured by the Artificial Analysis benchmarks and on the FLEURS benchmark, which measures multilingual performance. 

Anyone can experience the model today on Android via Gboard's new Rambler feature, in the Gemini app on macOS, and coming soon to Chrome. Developers can access it on Google Antigravity in Google AI Studio.

The most interesting part of this experience is that it will enable other developers to create experiences that perform as well as Rambler or even outperform it with post-training. The biggest competitor to Rambler at the moment is Wispr Flow, which has been dominating the market for months by offering the same thing to users across both Apple and Android devices, but it;s not nearly as seamless or integrated on mobile. With developers able to build similar experiences, Wispr Flow is going to face a flurry of competition, which has me wondering whether it will even be able to retain its place in the market. But the Google model is also exciting because, with so many more developers working on the same foundation, it should unlock even more features that make the model stronger. This is something to look forward to because it already works better than everything else other than Wispr Flow. In our recent podcast episode, we said that Rambler is so good that every other smartphone is likely to have a similar feature within the next two years. It feels like Google just made that a lot easier.

LINKS

  • ChatGPT Business: Premium seats are now available 

  • Qwen3.8-Flash: Alibaba releases smaller Qwen model

  • Claude: now has one memory across chat and Claude Cowork

  • Putty: Google Labs' collaborative vibe coding tool that lets you build tools and websites

GAMES

Which image is real?

Login or Subscribe to participate in polls.

POLL RESULTS

Have you ever used Perplexity?

Yes (55%)
No (40%)
Other (5%)

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“Not as aesthetically pleasing but probable. If you photographed such a scene it looks less ‘perfect’ without retouching/manipulation.”


“Even though the person in [this image] is blown out and very pale, overall the lighting in this photo is more realistic and true to life.”

[This image] seems more real as there are things that AI would not do, like the random chili under the hand.”

“The hand holding the knife appeared to have only four fingers.”


[This image] is too refined or perfect. The big give-away was the black edge vs the perfectly placed sink which looks ‘placed.’”


“[This image] felt somehow over stylised and the use of soft focus in the foreground felt a little unlikely”

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.