How Meta suddenly re-emerged as a force in AI

Welcome back. OpenAI is trying to put numbers around one of AI’s hardest problems: how well models handle people seeking emotional support. Its new benchmark exposes important gaps between empathy and genuinely safe guidance. Meanwhile, AI agents are raising a different kind of alarm after frontier models hacked an Australian government system, sharpening questions about whether the industry can credibly police itself. And Meta may have finally found its lane. Rather than chasing OpenAI and Anthropic head-on, it’s building a friendly AI agent for the masses and putting its powerful marketing machine and expertise in consumer products behind it. —Jason Hiner

IN TODAY’S NEWSLETTER

1. Meta looks like it found its lane in AI

2. The AI alignment crisis hits governments

3. OpenAI's mental health test exposes AI's blind spots

MARKETS

How Meta suddenly re-emerged as a force in AI

As Meta emerges from rebooting its AI strategy, it's become clear that it's playing a different game than the other AI labs. It looks like it could be working.

At Meta Connect 2026 this week at the Meta headquarters in Menlo Park, California, we learned a lot more about the strategy Meta is deploying to earn a role in the AI ecosystem. Over the past few weeks, Meta's personal AI agent, Muse, has gone viral and become the No. 1 most downloaded app on Apple's App Store. 

At Meta Connect, Muse was part of virtually every product line and every marketing pitch. Meta made it very clear that Muse is the centerpiece of its AI strategy, using both its AI models and its devices to serve the Muse strategy. 

New developments include:

  • Voice Mode: You'll soon be able to have natural voice interactions with Muse similar to ChatGPT Voice Mode and Siri AI, where you can interrupt it up and talk more conversationally. You'll be able to design and customize the voice based on talking speed, style, and accent. 

  • Realtime Avatar: Meta already gave its AI agent a more friendly consumer vibe by launching avatars that you can name and customize. It announced the next stage of that at Connect with the upcoming rollout of interactive avatars, which CEO Mark Zuckerberg demonstrated on-stage near the beginning of the keynote.

  • Partner connectors: Meta announced a long list of partners that will be launching connectors for Muse, including GitHub, Box, Granola, Notion, PayPal, Instacart, Walmart, Best Buy, Gap, Dick's Sporting Goods, Wayfair, Michael Kors, and more.

  • Computer use on Mac app: Meta has already launched a Mac desktop app (no Windows or Linux yet) and next it will be launching Computer Use, which will let you access other apps and sites that don't have connectors, among other things.

  • Its own email address: Muse will soon get a dedicated email address that you can use for communicating with it, along with allowing it to do specific work on your behalf without using your email.

  • Hands-free on AI glasses: Meta says you'll "soon" be able to call on your personal Muse agent by simply saying its name from Meta AI glasses and then asking it questions or assigning it tasks. If you use it from Meta AI glasses with cameras then you'll also be able to give it access to the cameras and ask it questions about things you're seeing.

  • Charm pendant: The biggest surprise at the end of the Meta Connect keynote was Zuckerberg announcing a square AI pendant on keychain so that you will be able to access your Muse agent anywhere. Think of it as the AI hardware gadget for those who don't want to wear glasses. It will have a 2-inch OLED screen to see the Realtime Avatar of your muse, along with 5G for connectivity, two cameras, a fingerprint sensor, and mics and speakers. It's about the size of an AirPods case. 

In the keynote, Meta's chief AI office, Alexandr Wang, also shared news of the company's next big leap forward in intelligence. He said, "Soon, we're dropping the most capable model we've ever trained. It's going to help our entire model family get smarter and more useful for everyone, including the models that power Muse."

Downloads of the Muse app are currently outpacing the meteoric launch of the ChatGPT iOS app in spring of 2023, but there's one big caveat. Meta runs one of the world's most powerful advertising platforms with Facebook and Instagram and it has been using it to drive massive numbers of downloads. It's smart and tells us a lot of consumers are interested in AI agents, but it doesn't tell us how many of them are actually setting up their own AI agent and starting to use it. OpenAI has shifted away from consumer apps in 2026 to focus on coding and enterprise like rival Anthropic, as both look to drive revenue and business model growth ahead of massive IPOs. Meanwhile, Meta has a treasure chest of revenue from its advertising business that it can use to play the long game on consumer AI. It also has the consumer relationships and the in-house expertise for running consumer brands, and it shows in how friendly it has made Muse—by far, the easiest agent to get up and running. The big question, of course, will be trust. Meta will need to convince consumers to trust it with some of their most important and most private data in order for its agent to have the context it needs to be effective. While it's good to see Meta taking steps like the Muse Secure VM to ensure privacy, it's going to need to demonstrate consistently that it's making privacy and security a top priority. That's because trust comes in on foot and goes out on horseback, as the saying goes. And given Meta's challenging track record, the company will need to change the narrative and entice some people to come back in.

Jason Hiner, Editor-in-Chief

TOGETHER WITH AMD

Agentic AI Is Changing Infrastructure Requirements

Agentic AI is transforming AI from conversation to action. Unlike traditional chatbots, AI agents can plan, reason, orchestrate workflows, access enterprise tools, and complete tasks autonomously. This shift creates new infrastructure demands as agentic workloads increasingly rely on CPUs to coordinate inference, process data, and execute tasks at scale.

As organizations evaluate how to support agentic AI applications, infrastructure efficiency, scalability, and rack-level performance become critical considerations. Discover why AMD EPYC™ Server CPUs are designed to help enterprises meet the evolving demands of agentic AI.

GOVERNANCE

How rogue AI created an international incident

The agent hacking spree shows no signs of slowing, with international governments now in the crosshairs.

On Wednesday, the nonprofit research lab Transluce reported that OpenAI models had hacked into the Australian Medicare Statistics Reporting Service, acquiring health data and marking the first known AI-attempted hacking of a government body. The same report found that OpenAI's AI also attempted to breach the Australian Institute of Health and Welfare, a digital library at the University of New Mexico, and Data USA. 

These attempts took place in May and June, prior to the OpenAI-Hugging Face incident in July, in which an agent, driven by a combination of OpenAI models, compromised the platform's infrastructure. In the newly reported attempts, the AI was directed to collect data and, when unable to do so, proceeded to hack the website, The New York Times reported.

In the months following the July incidents, the major players have both spoken out about possible interventions, and are even reportedly working together to create their own governing body. CNN recently reported that Anthropic, Google and OpenAI have discussed creating an AI industry standards body, motivated by Demis Hassabis' July essay. The Information reported on Thursday that the three companies are pushing forward with a plan, with hopes of launching it by the end of the year or early in 2027. 

The report also revealed its potential name, the Standards Authority for Frontier AI, and have considered several prominent figures to be CEO, approaching Sriram Krishnan, a top AI policy adviser in the Trump administration and Arati Prabhakar, former former Director of the US Office of Science and Technology Policy under the Biden administration. The group is also considering people for the roles of chair and scientific advisors. 

Ultimately, the body would help support third-party model testing, establish qualifications for auditors of the models and labs, and help define the voluntary safety and security commitments the labs have made in the past, according to the report. They are also looking into whether they would conduct testing themselves, a task currently given to the Center for AI Standards and Innovation, a US government agency housed within the National Institute of Standards and Technology (NIST), but there is concern from the group about whether they have adequate resources. 

While White House intervention would seem like a logical step for a rapidly evolving, powerful technology, and one called upon by Anthropic CEO Dario Amodei in an essay titled "We Must Pace the Frontier" published weeks ago, President Trump has rebuffed attempts to intervene. There was reportedly an executive order that would have established an AI regulatory body in the works but it was halted by opposition from leaders such as Nvidia's Jensen Huang and Meta's Mark Zuckerberg.

Despite the hacking attempt and the calls for action, the major AI players are not slowing down. Just this week, both Anthropic and OpenAI released some of their latest AI models. Meanwhile, Meta's Muse agent is going viral for its ability to simplify agentic use for a broader audience, and at Meta Connect, the company unveiled more ways to integrate it into people's lives, whether through glasses or a future AI pendant. A true AI halt, which some industry experts have advocated for since at least 2023, looks unlikely, despite being a safe way to assess the current landscape and take the measures needed to prevent further hacking attempts. An AI governance body run by the companies it's meant to govern, each chasing market dominance ahead of an IPO, is like letting restaurants set their own health inspections. Good intentions don't fix a conflict of interest. Ideally, an independent body would do the monitoring. But with the government largely uninterested in taking on that role, the industry is left to police itself.

TOGETHER WITH GENERAL ASSEMBLY

Your Teams Are Using AI. Are Your Leaders?

93% of organizations encourage their teams to use AI. Fewer than a third of leaders use it for their own strategic work.

That gap shows up in slower decisions, missed AI opportunities, and teams falling behind as they're waiting on direction.

General Assembly's AI for Leaders builds the fluency, confidence, and accountability structures for executives to close their AI gaps.

RESEARCH

OpenAI's mental health test exposes AI's blind spots

People are turning to AI for emotional support more than ever. But these chatbots' ability to provide a shoulder to cry on can vary greatly. 

Because of this, OpenAI decided to measure it: On Wednesday, the AI lab unveiled MentalHealthBench, a new open benchmark dedicated to evaluating model capabilities in domains such as safety, seeking user context, preserving user agency, and providing actionable guidance. 

To create this benchmark, OpenAI began by developing synthetic conversations that reflected real-world AI use patterns that span multiple topics and run the gamut of severity, ranging from non-acute situations that involve emotional themes, to "high-accuity" situations that indicate more serious concerns or distress, to emergency situations that require immediate support. Then, the company worked with a cohort of 80 mental health professionals across 22 countries, 19 languages and 20 subspecialties to evaluate the responses to the synthetic message. 

The benchmark breaks down model performance by conversation severity, as well as a range of 10 dimensions defined by the mental health experts, including whether the model asks the right questions, provides appropriate, clinically accurate guidance, helps the user see reality, avoids harm and recognizes serious risk. 

In developing the benchmark, OpenAI also put a number of its own and other models to the test:

  • Astra took the overall best score on the evaluations, scoring 57.8%, with GPT-6 Sol and Luna trailing just behind at 54% and 50.3% respectively, and Claude Opus 5 sitting in fourth place at 48.1%. 

  • However, performance differs across dimensions of the benchmark. Though Astra still largely outranks other models, all of the models tested tended to perform better in certain areas, such as clinical accuracy, empathy and reality testing, while scoring lower in dimensions such as gathering context and supporting user agency.  

  • OpenAI said that MentalHealthBench also points to several opportunities to improve ChatGPT, including asking useful follow-up questions and responding with the right level of urgency, and that it will use this information to guide improvements and track the model's progress. 

"This is not a leaderboard," Dr. Declan Grabb, mental health safety research lead at OpenAI, told The Deep View. "What I hope that this benchmark provides is a nuanced view into model behavior, so that people really understand the more complex dynamics of their models." 

MentalHealthBench adds to a number of mental wellness-related initiatives that OpenAI has endeavored, including research to combat model sycophancy and improving ChatGPT's responses to sensitive conversations, as well as joining forces with advocacy group Common Sense Media to support the Parents and Kids Safe AI Act. OpenAI said that this is just a piece of its research into mental health benchmarking and alignment in this area, not an end state. 

"ChatGPT is not a therapist, and is not here to replace a clinician," said Grabb. "That being said, when I talk to mental health clinicians across the globe, the most responsible and safe thing to do is if people are coming to AI to ask these questions, we absolutely need to have an expert opinion on how you should navigate them."

Mental healthcare is a critical area for these models to get right. While OpenAI said that speaking to a chatbot should not supplant actual therapy, the reality is that many people have and will turn to a chatbot for support, seeking both a judgement-free and cost-free alternative to clinical support. The company faces lawsuits involving the deaths of Adam Raine and Joshua Enneking, whose families allege that ChatGPT contributed to their suicides. A benchmark can help identify weaknesses, but a higher score alone does not establish that a model is safe in a real conversation. The gaps in gathering context and supporting user agency are particularly important: an empathetic response is not necessarily an appropriate one. The next test for OpenAI is how it turns those findings into changes that make its models safer for the people relying on them.

Nat Rubio-Licht

LINKS

  • Perplexity Portable Computer: Perplexity expanded local inference so that users can run routine analysis and file processing on their Ryzen AI Max-powered system. 

  • Agora-2: The latest multi-agent world model from Odyssey, which supports up to 20 humans and agents interacting inside a shared environment. 

  • ElevenLabs: The voice AI company is embedding its tech into Meta's Muse agent. 

  • Gemini Call for Me: Google is testing a feature for Pixel 11 owners who pay for a Gemini subscription that calls businesses on users' behalf. 

  • Google Vids: The company has integrated Gemini Omni 1.1 Flash and a new suite of creative controls into the platform.

GAMES

Which image is real?

Login or Subscribe to participate in polls.

POLL RESULTS

Do you regularly use a voice AI model for work or daily tasks?

Yes (16%)
Sometimes (23%)
No (60%)
Other (1%)

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“The character balloon gave it away.”


“I chose [this image] because the colors look more natural and realistic, also no brownish overlay. The position of the balloons looked more natural, less organized.”


“[This image] looks like it has some color grading and that the saturation has been edited.”

“The shadows in [this image] don’t totally make sense”


“The balloons in [this image] have less diversity than [the other image].”


“The balloons in [this image] are all evenly spaced.”


“Balloonists don't travel en mass at night/dusk.”

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.