• The Deep View
  • Posts
  • Google's lighter Gemini models reveal AI’s new race

Google's lighter Gemini models reveal AI’s new race

Welcome back. AI glasses are getting a workplace makeover. Halliday’s new second-generation glasses drop the camera and remake the experience around meetings, with live captions, translation, summaries, and other productivity boosts. Cisco, meanwhile, is betting that small language models (SLMs) can be stronger at cybersecurity. Its purpose-built Antares models aim to find software vulnerabilities faster, locally, and at a fraction of frontier-model costs. And Google’s latest Gemini models show where the broader AI race is heading: lower latency, fewer tokens, and cheaper agents. Jason Hiner

IN TODAY’S NEWSLETTER

1. Google's lighter Gemini models reveal AI’s new race

2. Cisco bets small models can solve AI's big problem

3. Halliday rebuilds smart glasses around meetings

PRODUCTS

3 new Gemini models want to reduce AI agent costs

As the race for token efficiency heats up, Google is making its move: three new models designed to help developers do more for less.

On Tuesday, Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Cyber. The new additions to the Flash lineup are meant to provide developers with lighter models that have lower cost and latency, but don't compromise efficiency. That's a balance that is especially necessary to build, deploy and scale AI agents. 

Each model has its own strengths. Gemini 3.6 Flash builds on Gemini 3.5 Flash, announced in May at Google I/O. The company incorporated developer and customer feedback to improve coding, knowledge, and multimodal performance, but all at a lower cost per token. According to Google, on the Artificial Analysis Index, 3.6 Flash uses up to 17% fewer output tokens than its predecessor, underscoring its efficiency. 

A quick look at the other models: 

  • 3.5 Flash-Lite: Google shares that it is its "fastest and most effective 3.5-class model," meant for tasks that require both low latency and high throughput. 

  • 3.5 Flash Cyber: Built on top of 3.5 Flash, it was fine-tuned to find and fix cybersecurity vulnerabilities at a lower price per token than its larger counterparts. 

Notably absent was Google's flagship Gemini 3.5 Pro, which has reportedly been delayed and is months behind schedule due to performance issues and not living up to the coding prowess of rivals Anthropic and OpenAI, according to the report. Google did mention in this release that Gemini 3.5 Pro is currently testing with partners and plans to make it broadly available "as soon as it’s ready."

Google also launched a preview of CodeMender, its security agent managed and hosted by Google, which can scan code to find vulnerabilities and fix them using the brand new 3.5 Flash Cyber. The company shares that CodeMender is already finding and fixing vulnerabilities in Google's internal codebases, including Chrome, Android, Cloud, Ads and YouTube. Gemini 3.5 Flash Cyber will be available exclusively via CodeMender to governments and trusted partners as part of a limited-access program.

Meanwhile, everyone can access Gemini 3.6 Flash and 3.5 Flash Lite in the Gemini app, and 3.5 Flash-Lite is also rolling out to Google Search. Developers can access them in the Gemini API via Google AI Studio and Android Studio. Enterprises can access them in the Gemini Enterprise Agent Platform. Gemini 3.6 Flash is also available in Google Antigravity and the Gemini Enterprise app.

AI agents, although prolific in enterprise environments, require a lot of processing to run in the background. As a result, demand is higher than ever for models that are highly capable while keeping token costs low, especially since most AI enthusiasts run multiple agents at the same time. Google is good at listening to developer requests and catering to the community, and this announcement is an example of that. Lastly, this shows the same trend I have been writing about nearly every day for the past two weeks: an industry committed to decreasing AI costs. The biggest challenge for Google is the continued delay of its flagship Pro model, which would allow its coding tools to match the runaway success of Claude Code and Codex among AI enthusiasts and builders. 

Sabrina Ortiz, Senior Reporter

TOGETHER WITH GENERAL ASSEMBLY

33% of leaders have already cut a role because of AI. Now they're worried about their own.

One in three leaders have eliminated or skipped opening a role because they believed AI could do the work. These same leaders making that call are growing less sure about their own future.

In 2024, 65% of leaders didn't think AI would replace them within ten years. That's down to 56% in 2025, and 34% now say there's at least a chance it happens to them directly.

GA's AI For Leaders training helps you lead through that shift with a plan, instead of reacting to it after the fact.

GOVERNANCE

Cisco bets small models can solve AI's big problem

With Anthropic and OpenAI both throwing their weight behind powerful (and expensive) cybersecurity capabilities, Cisco may have an affordable alternative. 

On Tuesday, Cisco unveiled Antares, a series of small language models that were built specifically for one time-consuming, expensive security task: Finding and rooting out known vulnerabilities in codebases. Because Antares is purpose-built, the models are specifically meant to take over the early stages of source code vulnerability triage, rather than replace expert judgment. 

The open-weight Antares series comes in two sizes, 350 million parameters and 1 billion parameters, and is currently available on Hugging Face. There's a heavier 3 billion parameter version coming soon. 

  • According to Cisco, benchmark testing shows that these models outperform both closed- and open-weight models on security tasks, including Google Gemini and Gemma, GPT-OSS and Alibaba's Qwen, and perform on par with OpenAI's GPT-5.5. 

  • Additionally, Antares can do this work at a fraction of the cost, sitting at 172 times cheaper than GPT-5.5 at the same tasks, and can run locally on-device and on your on-premise servers. 

"This is a great win for SLMs," Rob May, CEO of Neurometric, told The Deep View. "We've seen this time and time again across our customer base, that fine tuning a small model to a specific task will usually beat a frontier model on accuracy at 10% of the cost, and also run faster. But you have to choose the right workloads to make the fine tuning worthwhile, and it appears Cisco did a great job of that."

There are multiple reasons that Cisco decided to go small with Antares, Supriti Vijay, AI researcher at Cisco, told The Deep View. Along with dramatically undercutting large, general-purpose models on cost, these compact models can run locally and learn specific behavior needed to hunt down code vulnerabilities, she said, giving the user more control and limiting privacy concerns. Additionally, because these models are so lightweight, she said, inference can take "seconds rather than hours."

"The opportunity with SLMs is that they can handle the large volume of focused, repeatable work that makes frontier models too expensive to use everywhere," said Vijay. "That shows why specialized SLMs are likely to become an important part of everyday cybersecurity workflows." 

Cisco's models come at a time when defenders need to work faster than ever to find and fix weaknesses. And as AI labs continue to leapfrog one another by releasing increasingly powerful models, this job is only going to get harder, as we've seen with the recent releases of Mythos and GPT-5.6 that were so powerful the US government put on the brakes. Plus, while these massive models offer defense capabilities, the age-old adage in cybersecurity is that for every ten-foot wall there's a twelve-foot ladder. In other words, everything we create to defend ourselves can also be used against us. 

With limited resources, small models allow defenders to do more with less, changing the "economics of defense," said Vijay. 

"The biggest challenge in cybersecurity is often not a lack of knowledge, but a lack of time and resources to work through a growing backlog of code and vulnerabilities," said Vijay. "If organizations can cover common weaknesses faster and more consistently, attackers have to spend more effort finding the gaps that remain."

Cisco's small, purpose-built and inexpensive language models help cover a massive gap in the industry, and they point to an uncomfortable truth. AI has created a world in which cybersecurity defenders need all of the help they can get. And though AI can, and should, be used as a tool in the cyberdefenders' arsenal, new models from Anthropic and OpenAI are often sprawling in size, expensive, and simply not realistic to run every task within an organization. Cisco's models are just one example of the cost-saving power of small models, especially as enterprises finally start to bring down the hammer on tokenmaxxing. To put it simply: You don't need Thor's hammer to kill a fly. 

Nat Rubio-Licht

TOGETHER WITH CDATA

Secure MCP Architecture for AI Data Connectivity

AI agents need access to enterprise data, but unmanaged MCP servers introduce user and agent permission risks, credential sprawl, and zero audit visibility. CData's 2026 Security Best Practices Guide gives IT and security teams a concrete framework for deploying AI safely. 

Learn how identity-first access, role-and-attribute-based access controls, and comprehensive audit trails operate within Connect AI to keep you in control.

Download the free guide and learn more about security with CData Connect AI.

HARDWARE

Halliday rebuilds smart glasses around meetings

Halliday's first-generation glasses drew praise in demos for an optical display that required no in-lens projection, but a disappointing launch prompted a complete pivot for the successor model.

On Tuesday, the company announced its Halliday G2 AI smart glasses, which aim to offer AI assistance to working professionals, a pivot from focusing on a general consumer target audience in the first glasses unveiled last year at CES 2025. At launch, the core of Halliday G2 is found in Meeting Flow, which focuses on live-meeting capabilities such as live captions, translations, transcripts, summaries, and more. 

“AI has transformed how people write, research, plan and summarize, but the most important moments still happen live, in meetings, client conversations and the social interactions that shape everyday life,” said Carter Hou, COO and Co-Founder of Halliday, in the release. 

Halliday also pivoted away from last year's DigiWindow display, a tiny projector module hidden inside the top of the frame that beamed a small, green monochrome image directly into your right eye.

Instead, the G2 glasses have switched to waveguide in-lens displays that feature dual micro-LED projectors (up to 1,600 nits), which sit in an upper-field position to keep info glanceable without breaking immersion. In a brief hands-on demo, this placement stood out as different from competitors like the Even Realities G2, which display straight ahead, and instead echoed the original Halliday glasses' upward-glance design. 

Other specifications include: 

  • Microphones: 4-mic array that offers up to 2-meter voice recognition

  • Battery life: Halliday claims 12 hours of typical use 

  • Weight: Under 50 grams

  • Speakers: While the release doesn't include specifics, during my demo the audio was a crucial component of the hands-off experience

Notably, the glasses omit cameras, and the company claims it was done purposefully to make them suitable for work environments to avoid having coworkers fear being recorded and to comply with company security protocols. This is especially pertinent as Meta's AI glasses continue to face backlash over privacy concerns. 

Other workplace-focused AI meeting features include Thread Tracker, Decision Confirmation and Commitment Check, which are meant to help users stay with the conversation. The non-meeting features include real-time translation in over 45 languages, notifications, phone calls, music, teleprompter, and the Halliday AI Assistant. 

At the time of my demo, most of those capabilities were still in beta, so I am not going to comment on those features yet. However, I was able to try them on and can say that they felt light and comfortable. Yet, like most smart glasses, they appear visually chunkier, which makes them look like smart glasses. That could be a personal perception. You can decide from the photos. 

The glasses will retail for $599, with shipping expected in September. If interested, you can reserve the glasses today for a $10 deposit, securing a $100 discount on the final retail price.

It is interesting to see a hardware AI product follow the same path as many software companies by pivoting to focus on professionals and the enterprise. The decision makes sense, since workers invest in products that can boost their productivity and, as a result, are willing to spend money and commit to using the products. However, the challenge is that many non-enterprise products such as Even G2 can offer many of the same features, such as meeting summaries, teleprompters, notifications, translation, and more. Those also have the bonus of name recognition and interoperability with devices users already have. Plus, big names such as Samsung and Google are preparing to enter the market, making it even harder to compete. 

LINKS

  • Claude Cowork: Users can now record their screen while they're working to teach Claude a skill or task. 

  • Block Buzz: Block's new open source collaboration platform, offering a workspace for agents and users to work together. 

  • Substack: The newsletter platform has launched an AI detection feature in partnership with Pangram. 

  • Meta Storykit: Meta releases an AI iPhone app for personalized children's stories

GAMES

Which image is real?

Login or Subscribe to participate in polls.

POLL RESULTS

Do you think there should be stricter regulatory safeguards around the way minors use AI?

Yes (87%)
No (7%)
Other (6%)

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“Moles on woman's skin, her scalp and the way the water was distorted through the ring.”


“[This image] had subtle chaos, water droplets on the body and floaty, extra reflective motion in the water, the fingers were in tact with a struggling grip.”


“Condensation on what appears to be the inside of the tube would not be AI. “

“[This image] is too sharp, highlights too defined, water patterns too complex, colors too intense.”


“Unnatural positioning - no way the model [in this image] could be sitting on that tube at that angle.”


“Generally, it seems that AI reworks are adjusted more that they needed to be.”

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.