OpenAI's Hugging Face breach is a warning shot

Welcome back. AI may be writing more code, but software engineers aren't disappearing. Their work is shifting toward verification, making judgment calls and bigger strategic decisions. Meanwhile, Samsung and Google’s smart glasses look polished, practical, and privacy-friendly, but Gemini alone may not be enough to leapfrog Meta Ray-Bans. And of course, the most important story is OpenAI’s accidental breach of Hugging Face during a cybersecurity test. The company deserves credit for disclosing it, but the incident is a stark warning: Frontier models are becoming powerful enough to escape their boundaries, and existing safeguards aren't ready for how powerful the models are getting. Jason Hiner

IN TODAY’S NEWSLETTER

1. OpenAI's Hugging Face breach shifts safety debate

2. Samsung’s AI glasses need more than Gemini

3. Why AI still won't replace software engineers

GOVERNANCE

OpenAI's Hugging Face breach is a warning shot

Increasingly cyber-capable AI models are starting to show their teeth. 

On Tuesday, OpenAI claimed responsibility for a security breach of model platform Hugging Face, in which an agent, driven by a combination of OpenAI models, compromised the platform's infrastructure. The breach, which included the company's latest and most powerful model, GPT-5.6 Sol, and "an even more capable pre-release model," occurred while OpenAI was internally testing models on ExploitGym, a benchmark for cybersecurity capabilities. 

In a blog post, OpenAI said the incident occurred during an internal evaluation in which researchers were prompting the models to find and pursue "advanced exploitation" using complex attack paths. 

  • The evaluation aimed to help the company estimate the maximum cyber capabilities that its models had as a means of preventing them from "pursuing high-risk cyber activity." 

  • Though the benchmarks are performed in isolated environments, the models did their jobs too well, breaching containment by finding and chaining together vulnerabilities across OpenAI's research environment to access Hugging Face's infrastructure to search for solutions to the ExploitGym benchmark.

  • OpenAI has taken a number of actions in response, including implementing strict controls in infrastructure configuration, disclosing and patching the vulnerability that allowed for the incident, and improving protections around future evaluations. 

OpenAI has since received praise for coming forward about the event, including from employees at rival Anthropic: Jack Clark, a co-founder at Anthropic and former OpenAI staff member, said in a post on X that "there are many counter-incentives to publishing stuff like this, but by making it public we all get better info about safety at the frontier." 

Still, the incident marks one of the first major cybersecurity events as a result of massively powerful models going rogue, and it shouldn't come as a surprise. It could be the beginning of a trend that tech experts like AI godfather Yoshua Bengio have been warning about for years. Notably, even OpenAI said that this incident will likely not be singular, and that it expects exploitations like this to "become more commonplace with the proliferation of increasingly cyber-capable models."

It also highlights that enterprises and organizations may simply not be ready for models with power of this caliber. Barr Moses, co-founder and CEO of AI observability firm Monte Carlo, said that most organizations are overconfident in their ability to catch rogue AI agents. 

"The reason this risk exists is that most organizations don't yet know how to define "trust" for an agent," said Moses. "Trusting an agent isn't a one-time judgment, but rather an ongoing claim you can only back up if you have full visibility into its activity, decision-making, and underlying infrastructure."

It's noble that OpenAI owned up to its models being the root cause of this incident, and hopefully sets a precedent for other AI labs to continue taking accountability as more of these incidents occur. Still, this may also be a sign that the cutthroat AI race needs to slow down before more damage gets done. While Anthropic pitched an industry-wide unilateral pause on AI development earlier this summer when it published research on recursive self-improvement, its argument against hitting the brakes was that no other major industry player would agree to it. Additionally, many in the industry argue that slowing down would only cede the US's position as a dominant player in AI, allowing China to surpass it. However, many Chinese model companies have been accused of using model distillation of proprietary models from American AI labs as a means of propelling their own AI forward, with Moonshot's Kimi K3 being the most recent example. This may actually be an argument in favor of a slowdown. If the industry gets more methodical and thorough, not only does society have time to plan and calibrate for the ethical and safety consequences of these models' capabilities, but it also has more control over them going rogue or landing in the wrong hands.

Nat Rubio-Licht

TOGETHER WITH CRUSOE

Fine-tune open models. Keep your weights.

Open-weight models are catching up to proprietary models — and more teams are customizing them with their own data while retaining ownership of the fine-tuned weights.

Crusoe Serverless Fine-Tuning, now GA, runs in a tenant-isolated environment with zero data sharing: your data trains your models only.

Deploy in one click to Self-Serve Deployments for inference, or download raw weights in .safetensors format. You shouldn't have to choose between a managed experience and ownership of your model.

HARDWARE

Samsung’s AI glasses need more than Gemini

Samsung and Google's upcoming AI glasses could be what it takes to bring the category mainstream, and Unpacked gave us our clearest look yet at the product.

Ahead of the event, I went hands-on with Samsung's new AI glasses, and they look promising. The Warby Parker collaboration offers everyday styles that resemble the ones I wear every day, while the Gentle Monster collection is bolder and more fashion-forward, including the sunglasses-style frames that look a bit like the Meta Starfire Kylie Edition glasses launched in June. 

The Samsung glasses themselves are straightforward: they have one camera, an LED light to alert people when they are being recorded, open-ear speakers, and microphones. The use cases lean on capturing content, fast access to Gemini for everyday tasks, and hearing notifications from your most important apps. Powering them is the Snapdragon AR1 Gen1 Platform. 

Battery-wise, Google, Samsung, and Qualcomm teamed up to deliver what the company claims is 9 hours of active use on a charge, even if you're doing heavier things like streaming video or using Gemini Live. And the case can recharge them seven times. By comparison, Ray-Ban Meta Gen 2 has 8 hours of battery in general use, although much less for heavy use..

To the touch, the Samsung glasses feel pretty light, though the company hasn't released exact weight or specs yet. I wasn't able to try them on or test any of the features, so that will have to wait until they launch in the fall. 

Functionally, they operate similarly to the standard Meta Ray-Bans. They are audio glasses that lack in-lens displays. But Google and Samsung are more trusted by many users, and they have the advantage of interoperability with a much larger device and app ecosystem. Other than that, it's still hard to know how they will compare to Meta glasses performance-wise.

A big obstacle to entry in the smart glasses market is that the most popular and competent everyday smart glasses are made by Meta, and users have to contend with Meta's contentious reputation in privacy. For that reason alone, Samsung and Google's counterparts are already positioned to be more appealing to a broader segment of people. In addition, the fact that these models use Gemini, which has proven itself a much more capable assistant than Meta AI, is also likely to attract more customers. That said, hopefully, between now and the actual launch, Samsung will find other ways to differentiate these glasses from Meta. Samsung has deep expertise in hardware, and it'd be great to see them use that to drive the category forward rather than simply making their own version of something that's already been done.

TOGETHER WITH ATTIO

The next era of revenue runs on Attio.

Get agents working on every account, surfacing new opportunities, and handle the work that used to take your team days. With Attio’s new workflows engine, you can build agents for any GTM play you need. Just describe the goal, and Attio assembles the workflow for you to approve.

Then Ask Attio any question, from the weekly forecast to performance by rep, and get the answer in seconds. So you keep every deal moving.

Trusted by some of the world’s most successful startups, like Granola, Modal, and Railway, Attio is the agentic CRM that transforms how revenue work gets done.

WORKFORCE

Why AI still won't replace software engineers

Software engineering is changing fast as AI becomes increasingly capable of writing code. 

Michele Catasta, president and head of AI at Replit, an AI coding platform used by more than 50 million users to build software with natural language and agents, has a front row seat to that transformation.

The Deep View sat down with Catasta to discuss how AI is changing the role of the software engineer, why coding is becoming more strategic, and why entry-level talent may have an advantage in the job market. The conversation has been edited for length and clarity. 

Aaron Mok: In a recent A16z podcast interview, CEO Amjad Masad described Replit's Agent as "an automated software engineer" as capable as a "mid-level engineer at Meta or Google." That's a bold claim. From your conversations with customers, what are some of the most surprising ways you've seen them use the technology?

Michele Catasta: The thing that's been most surprising to me—even though I've been working on this for 10 years—is that I practically don't hear users say they can't build what they have in mind anymore. That wasn't the case a year ago. We've gone from people creating landing pages to building software that would've taken weeks or months earlier in our careers.

You're not going to take Replit's Agent, teleport it into Meta tomorrow, and replace all the software engineers. It's more that someone with those skills can suddenly build a new product or start a business on the side.

We're also seeing strong enterprise traction. Eighty-five percent of the Fortune 500 uses Replit. Before, if you had an idea for a product, you'd write a product requirements document, go through several meetings, and start designing prototypes. Now, you can show up to that first meeting with a functional prototype. Decision-makers can immediately decide whether it's worth building or whether it needs another iteration.

Mok: How is AI changing the role of the software engineer?

Catasta: I don't think anyone is doing less work today than before. If anything, this has been the most restless period in tech.

What's changing is the nature of the work. The cost of generating code is dropping. What's taking more time now is verifying whether it's correct. AI can generate a lot of code, but engineers still have to review it and decide whether to accept those changes.

Even if engineers spend less time typing code, they're spending more time debating trade-offs and working with designers and product managers. As the cost of generating code drops, there's more time to think. That's leading the industry to build better things.

Mok: There's growing concern that AI could reduce entry-level software engineering jobs. How do you see the career path for junior engineers changing?

Catasta: I know some large companies have stopped hiring below a certain level, and I think that's a huge missed opportunity.

For startups like ours, it's actually been fantastic because we're hiring exceptional graduates who grew up in the AI era. They started college around the time ChatGPT came out. They're AI natives. They've been using AI coding tools from the beginning, and no one is more ready for this revolution than them.

Mok: So is AI literacy becoming a competitive advantage?

Catasta: Absolutely. AI literacy has become table stakes.

The people we hire and the companies we sell to are AI-forward. They expect people to know how to work with these tools. As chaotic as this shift can be, it's also incredibly exciting because people can build much faster.

Six months ago, I would've looked at an ambitious project and told my team it would take a quarter. Now I can say, "It's going to take three weeks. Let's do it." It's extremely empowering to feel like we can move fast and make things happen.

Aaron Mok

LINKS

  • OpenAI Presence: A new tool that helps companies run agents more effectively on their data. 

  • Lucy 2.5: A real-time AI video and VFX editing tool by Decart.

  • Anthropic Economic Index: Users can now ask Claude specific questions about Anthropic's economic tracking system. 

  • ElevenMusic: ElevenLabs music generation tool now offers Vocals, allowing users to generate songs with their own voice or one from its library.

  • Salesforce: Senior Offensive Security Engineer (Red Team)

  • Reflection: Member of Technical Staff - Security Engineer

  • Neuralink: Security Engineer 

  • Meta: AI Research Scientist, CoreML - Monetization AI

GAMES

Which image is real?

Login or Subscribe to participate in polls.

A QUICK POLL BEFORE YOU GO

Do you think your organization is prepared to handle AI-enabled cyberattacks?

Login or Subscribe to participate in polls.

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“The bed's side, bottom is not perfect.”


“The plants are asymmetrical and lean, which is more realistic. The exposure is brighter than AI tends toward. The head depression on the pillow and the little pot's placement by the window are also rarer details that AI would not likely generate.”


“Light reflection on foot of bed frame.”

“In [this image], the mirror is reflecting a curve in the chair that is behind the mirror, and so could not be reflected. Also in the reflection, the carpet has a different pattern.”


“Was missing glare in [this image]. The hat on the chair looked too staged, the leaves on the plants too perfect.”


“Reflection in the mirror of the carpet pattern is where AI lost today.”

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.