• The Deep View
  • Posts
  • Gemini 4 Argon muscles into the frontier AI race

Gemini 4 Argon muscles into the frontier AI race

Welcome back. OpenAI rolled out more than 20 products at DevDay, but the most telling decision was what it held back. Its new Dots agents are limited to premium plans, and the next Astra model stayed on the shelf for now. Meanwhile, CrowdStrike's Adam Meyers joined me on The Deep View Conversations podcast to explain why AI may finally tip cybersecurity's balance of power in a surprising way. And Google is back in the frontier race with Gemini 4 Argon, which tops rivals on several key benchmarks. But as the labs leapfrog each other and models commoditize, price is starting to matter as much as performance. —Jason Hiner

IN TODAY’S NEWSLETTER

1. Gemini 4 puts Google back in the frontier AI race

2. OpenAI's restraint was the real DevDay story

3. How AI changes cybersecurity's balance of power

PRODUCTS

Gemini 4 Argon muscles into the frontier AI race

As its two biggest rivals continue to one-up each other with new frontier models, Google is making its biggest move of 2026. 

On Wednesday, the search giant unveiled Gemini 4 Argon, its latest flagship frontier model. The company said that Argon offers frontier performance in a number of complex workflows, including in software engineering, knowledge work, cybersecurity, and domains such as legal and finance. 

Google will begin rolling out the models specifically to cyber defenders in the Fairwind program, its restricted-access AI cyber program that it launched in early September. The company is also currently going through the US government's voluntary pre-release model checks before opening up the model to the general public. 

Google noted that Argon surpassed its previous generation, Gemini 3.8 Flash Cyber, in cybersecurity tasks such as real world-vulnerability discovery and penetration testing. 

The company said that Argon sets a new state of the art score on DeepSWE v1.1, the benchmark testing performance in real-world long-horizon software engineering tasks, sweeping OpenAI's Astra and Anthropic's Claude Fable 5.1 and Opus 5.5. 

  • Google's Argon also outperforms these competitors in benchmarks for knowledge work, long-context tasks, and computer use. 

  • Specifically, for knowledge work, the company said that Argon leads on the Vals Index, which measures impact in finance, coding, legal, and tax work, and offers state-of-the-art performance in visual understanding tasks, such as chart, document and video analysis, measuring by the LVBench for multimodal understanding. 

  • However, Argon is still beat by Astra in FrontierSWE v2, which evaluates agents on complex, multi-hour technical tasks and Terminal-Bench Science for scientific workflows. Opus 5.5 also beats Argon on PostTrainBench for machine learning engineering, and Terminal-Bench 4.0 for tasks within sandboxed command-line terminal environments

In its announcement, Google noted several ways in which Argon is "fundamentally changing" its own workflows, including helping its quantum researchers optimize algorithms, improving memory efficiency, and handling large-scale codebase migrations. 

The model has an output limit of 1 million tokens, up from the previous 64,000 tokens. A Google representative told The Deep View that Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. That price puts it in line with both Anthropic's Claude Sonnet 5.5 and OpenAI's GPT-6.1 Sol, which run at the same cost of $2 per million input tokens and $10 per million output tokens. However, after the introductory period, the price of Argon will double to $4 per million input tokens and $20 per million output tokens. 

Google noted that the roll out will begin with paid API customers and Google AI Ultra subscribers.

Google could not have picked a more heated time to reenter the high-end frontier model race. In recent months, the company's contributions to the frontier landscape have largely focused on speed and efficiency. For instance, its early September release of Gemini 3.8 Flash cost a fraction of what its frontier competitors were charging at the time. But in recent weeks, with both Anthropic and OpenAI homing in on token efficiency as well, Google's Argon may not be able to compete on just state-of-the-art performance alone, especially as model labs continue to leapfrog each other in capability and efficiency week after week. What a model costs is quickly starting to matter as much as what it can do.

Nat Rubio-Licht

TOGETHER WITH TABS

Tabs + PwC: A practical framework for usage-based revenue in the AI era

Hybrid and usage-based models are changing how companies price, and how finance teams recognize revenue, forecast, and close the books.

In this on-demand session, Tabs co-founder Rebecca Schwartz and PwC partner Amit Dhir break down the real operational impact, with frameworks and examples you can apply now.

BIG TECH

OpenAI's restraint was the real DevDay story

OpenAI unveiled more than 20 new products at DevDay 2026, but the choices of what launched and to whom were just as notable as the announcements themselves.

The release of Dots always-on agents was the star of the show. The advanced agentic capabilities OpenAI touted would be enough on their own to draw attention: Dots can ingest all of your personal context to make proactive suggestions, and they can take multi-step actions using GPT-6 Astra, OpenAI's most advanced model—and what OpenAI also called its "most aligned" (safest) model. But part of the buzz comes from timing, since the launch follows the viral release of Meta Muse.

Dots and Muse share some similarities. At a high level, both are agentic AI assistants that can plug into nearly all of a user's data, and both are fronted by fuzzy, plush creatures (which I still find a bit strange). The most striking difference is who can access them. Meta's appeal is that Muse is easy to use and accessible to people at any level of AI experience. OpenAI, by contrast, is limiting Dots to its Pro plan, which starts at $100 per month and now goes up to $500, as well as its Business Premium and Enterprise plans.

OpenAI also decided to use GPT-6 Astra, which is not only its most capable but also its most aligned, or the best model at following directions and not going rogue. This comes with a trade-off: GPT-6 Astra requires lots of compute, so on a pragmatic level, limited access also keeps prices down, as OpenAI CEO Sam Altman told me when I asked about it in the press conference after the keynote. 

"We're starting out as a premium product—it uses a lot of compute and a high model—and I think there are very good reasons for that," said Altman. "It will help us understand and help people understand the expanse of what AI can do, but you should, of course, expect us to do a mass market version, bringing it to billions of people."

Limited access also keeps the number of people who could be impacted lower, and for whatever incidents do occur, OpenAI can collect feedback before releasing it to a wider subset of people, which is also a more responsible course of action. Another example of OpenAI showing restraint was it not releasing the next iteration of its most advanced model, Astra GPT-6.1 in efforts to, presumably, prioritize safety.

Using these agentic features requires handing off a lot of personal data and running the risk of the agents taking unauthorized action. As a result, although equal access to AI tools is important, especially to not widen the wealth gap further, it's crucial that the people accessing powerful AI tools have proper amounts of AI literacy, and, if a user is paying for a Pro plan, I think it is fair to assume they use AI often enough to justify that cost. Another thing that makes this strategy stand out is that OpenAI has spent the past few months loudly calling for AI safety measures, but posts and pledges only go so far. What the moment demands is action, and this kind of restraint is a sign that OpenAI is willing to back its words with action.

TOGETHER WITH NOOKS

Why sales is still AI's hardest test

In early 2025, Nooks' revenue agents scored 97% on a standard eval benchmark. However, in the field, reps still double-checked every agent output.

That gap drove new research from Nooks Labs, the company's applied AI team. Nooks CTO Nikhil Cheerla explains why sales agents break where coding agents don't: Feedback takes weeks, not seconds, key information is often hidden from the seller, and CRM data can be messy and conflicting.

The team's fix centered on the system around the model: a harder benchmark built from real deal data, a way to score a strategy before it runs, and 500+ examples of bad AI writing used as guardrails.

GOVERNANCE

How AI changes cybersecurity's balance of power

AI is accelerating cyberattacks. Could it also give defenders the upper hand?

In this episode of The Deep View Conversations, we sat down with Adam Meyers, CrowdStrike's head of counter adversary operations, at Fal.Con 2026 to explore how AI is changing the balance of power between attackers and defenders.

Meyers explains why he believes security teams can now match the speed and scale of their adversaries, how CrowdStrike is training offensive and defensive AI models against each other, and why the harness that guides an AI model can matter just as much as the model itself. He shares how changing that harness helped CrowdStrike reduce false positives in its vulnerability research without changing the underlying model, and how cyberdefenders are uniting to assist one another.

The conversation also goes inside the live disruption of the Sality botnet, explores how nation-state attackers are using agents that learn from their mistakes in seconds, and examines why powerful AI systems need clear boundaries as they pursue their goals.

Topics covered:

• Why AI could change the defender's dilemma
• How CrowdStrike and its partners disrupted a botnet that had operated for more than two decades
• Training red team and blue team models based on Nvidia's Nemotron models and CrowdStrike's security data
• Why AI harnesses are critical to reducing hallucinations and false positives
• How AI attacks are compressing the time organizations have to patch vulnerabilities
• What the Hugging Face incident raises about goal-seeking agents and guardrails
• Why local, open-weight models matter to both attackers and defenders

If you're trying to understand what AI means for cybersecurity, how to put agents to work safely, or where specialized models can make a difference, this conversation offers a view from someone tracking the adversaries firsthand.

Subscribe to Deep View Conversations for interviews with the leaders shaping the future of AI, business, and technology: tdv.transitor.fm

Jason Hiner, Editor-in-Chief

LINKS

  • Runway Ads: The AI video firm now can connect to a brand’s accounts on Meta, Google, and TikTok.

  • Manus Flex: Users can now bring their own API key to Manus's harness, tools, and execution environments.

  • Cohere Embed 5: Cohere's new state-of-the-art family of embeddings models.

  • fal Agent: The AI video firms' creative agent for image, video, audio, and 3D is now generally available.

GAMES

Which image is real?

Login or Subscribe to participate in polls.

POLL RESULTS

Do you plan to try OpenAI's "Dot" agents?

Yes (16%)
Maybe (30%)
No (50%)
Other (4%)

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“The out-of-focus background is what made me say [this image] was real.”


“The seed heads obscure the base of the dandelion as expected.”


“Short depth of field, very evocative of humans.”


“Out of focus background shows [this image] is real. AI would consider it unacceptable quality.”

“[This image] looked to sharp, clean and detailed.”


“I'm noticing a theme with some AI images: they all have a very shallow depth of field but everything else in frame is still super detailed somehow.”


“The center of the flower was simply too translucent, especially for having the sun behind it. It also felt like the 2 seeds sticking out were not realistic, like an effort to make the image "less than perfect" to fool the viewer.”

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.