• The Deep View
  • Posts
  • Why Astra's opacity problem could force a pause

Why Astra's opacity problem could force a pause

Welcome back. OpenAI’s ChatGPT Work team joins our podcast to explain how agents and proactive AI will make prompting less central to the ways we use AI throughout the day. Google’s much-ballyhooed Rambler offers Wispr Flow-level speed and accuracy, but it's still missing the right tone by getting the punctuation wrong. And GPT-6 Astra raises the question of what happens when models become harder to monitor as they get smarter. With observability getting more challenging, OpenAI's chief scientist signals that it may be time to coordinate with other labs before control gets harder. Jason Hiner

IN TODAY’S NEWSLETTER

1. Why Astra could make an AI pause unavoidable

2. Google Rambler solved dictation, but not tone

3. What OpenAI is building for a post-prompt future

GOVERNANCE

Why Astra's opacity problem could force a pause

Last week, OpenAI's GPT-6 Astra made something clear: The future of AI is anything but clear. 

In the company's announcement of its most powerful model yet, it noted that Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s when tested explicitly on its ability to evade monitoring. The company attributed this to the model simply being smarter: It could solve problems in fewer steps and didn't need to write down every thought process on simpler tasks in order to think them through. 

In a briefing with the press last week, Jakub Pachocki, chief scientist at OpenAI, said that the company is working on ways to strengthen visibility and make the models "more verbose in their chain of thought." However, Pachocki said that lack of monitorability is "largely just a general consequence of increasing intelligence and a consequence of scaling."

"These more capable models can perform harder tasks using fewer language tokens or no language tokens, so we also see a big improvement in capability there, which also reduces our ability to monitor those easier tasks," said Pachocki. 

In the same briefing, OpenAI CEO Greg Brockman said that the capabilities of Astra mark a significant moment in the company's quest towards achieving artificial general intelligence, and that it's not unreasonable to think we are now in "the AGI era." 

"When we started OpenAI, we kind of thought that there was going to be this well-defined moment that everyone would recognize AGI," said Brockman. "It hasn't played out like that. It's a much more gray, fuzzy thing. But I think that if we fast-forward a couple of years, and we look back and say, 'when was it really that AGI was created?' I think it's going to be about this time, and I think it might be about this model." 

But having a more advanced model also heightens safety concerns. In an interview with The Deep View after the announcement, Pachocki reiterated, "We do see some tendency to kind of think less when it's told that it's being monitored, which is also a worrying trend." 

Extrapolating on that point, Arjun Jaggi, applied AI researcher, told The Deep View, "This isn't a theoretical risk. Earlier this year, when OpenAI's agents went rogue and attacked Hugging Face, investigators only understood what happened because they had chain-of-thought logs to read. That's how the tampering was caught. Take that visibility away and the next incident like it gets much harder to diagnose, possibly impossible to catch while it's happening."

OpenAI, Anthropic, and others have long talked about controllability, observability, and preparedness. However, the two rivals are also locked in a perpetual quest to one-up each other, creating powerful AI that can claim the crown of being state-of-the-art. But if we are already losing our ability to understand these models' inner thoughts, what's in store for us when they have 10x the capabilities that they do now? An inability to monitor the models risks being the first step towards an inability to control them. "I don't think anyone is prepared for a continued increase in machine intelligence at the current pace," Pachocki told The Deep View. "I think it's something we need to treat with extreme urgency, and we need to find ways to slow down AI development, to introduce safety gates, [and] to coordinate between labs but also between nations… [For] the preparedness framework, I think we have to evolve that to really also be about development, because currently it's very focused on deployment. In the future, I would like to involve third-party organizations more deeply into our development process." In June, the Anthropic Institute suggested that the "option to slow or temporarily pause frontier AI development" could give governments and AI labs time to align their processes around safety. With OpenAI signaling its willingness to pause, the ball is in Anthropic's court to make the next move. But if they do, then what about Meta, SpaceXAI, and the Chinese labs?

Nat Rubio-Licht
Jason Hiner, Editor-in-Chief

TOGETHER WITH RAMP

The unpredicted AI workforce

AI was supposed to replace workers - why hasn’t it? Many predicted it would.

Ramp's latest research suggests something entirely different (and almost nobody predicted).

As seen on Bloomberg, The Wall Street Journal, and The New York Times, join Ramp's lead economist, Ara Kharazian on Friday, September 11th at 11 am ET | 8 am PT, for a special 30-minute briefing on the real data regarding “The State of AI in the Workforce.”

CONSUMER

Google Rambler solved dictation, but not tone

For years, voice dictation's fatal flaw was that it struggled to accurately understand what you were saying. While AI has nearly eradicated that issue with products such as Wispr Flow and Google Rambler, a new problem has taken its place.

Just a couple of weeks ago, Google debuted its Pixel 11 lineup, and one of the standout AI features was the new Rambler voice dictation, which integrated into Google Keyboard (Gboard) and offered the same capabilities that Wispr Flow has been able to do for years: precisely understanding users' voice dictation, eliminating filler words, and following a user's train of thought. 

As someone who cannot work without Wispr Flow, I tested out Rambler for all my communication, both personal and work, as soon as I got my review units. I soon ran into a hang-up I had never noticed before: tone. 

To my boyfriend's credit, he was the first to bring it up, saying he preferred Wispr Flow because it adapted better to his expressions, such as automatically including exclamation points. In what would become foreshadowing, I told him to give Rambler some grace, since it, like Wispr Flow, claims to get to know you better over time and eventually implement that in its dictation.

However, despite my consistent use, Rambler has yet to adapt to my excited and bubbly tone, opting instead to send the driest texts. The catalyst for this review was when I was Slack messaging TDV editor Jason Hiner and had to clarify that I had used Rambler because the message I sent sounded absolutely nothing like me. It was much less friendly, more serious, and more to-the-point than I naturally speak or write in Slack messages. 

While it's still polite, the absence of tone indicators, such as punctuation, that you typically use can make people who know you think something is wrong.

For instance, my boyfriend sent a text via voice dictation using Rambler that was fine on its face, but the punctuation made it read as dry and dismissive. After a long day, I misread it that way and brought it up, sparking an issue that never would've happened without the dictation software's punctuation choices. There's actually research to back this reaction:

  • A 2016 Binghamton University study found that texts ending with a period were rated less sincere than those that did not. 

  • A 2025 follow-up study found that when periods were placed after every word, for instance, typing "What. Do. You. Need." instead of "what do you need," the perceived "mean" rating, or how mean of a tone the reader perceived in the texts, was higher, highlighting how readers interpret punctuation in a sentence as an intentional, meaningful act.

  • A 2018 The Atlantic article explored the phenomenon of people using so many exclamation points while texting and featured Gretchen McCulloch, a linguist who studies online communication, who said: The single exclamation mark is being used not as an intensity marker, but as a sincerity marker. If I end an email with ‘Thanks!,’ I’m not shouting or being particularly enthusiastic; I’m just trying to convey that I’m sincerely thankful, and I’m saying it with a bit of a social smile.

In theory, AI voice dictation features are huge game changers for productivity. If you say it out loud, it will transcribe it for you, saving you the time it takes to type, since most people can speak at least 2-3x faster than they can type. But transcribing what you say is only half the battle, as the consequences of not having that so-called social smile can be serious. For that reason, Wispr Flow remains the undisputed leader until Google fixes this in Rambler. However, this discussion does raise the question of how AI will transform text communication in the future, and how our perceptions of tone in text will evolve as a result. Oh, and to prove my point: I used Wispr Flow to write this entire article, and it understood my intent and intonation perfectly. 

TOGETHER WITH JUMPCLOUD

Your AI agent shouldn’t have to pretend it’s a person

Autonomous agents are doing real work, but many organizations still manage them through shared credentials, fake user accounts, or long-lived API keys.

That makes basic questions hard to answer: Who owns this agent? What should it access? What did it do? How do we shut it down?

Agentic IAM gives agents their own identity, lifecycle, ownership, and governed access, so IT can manage them like the distinct actors they are.

PRODUCTS

What OpenAI is building for a post-prompt future

AI is transitioning from just answering questions to doing valuable work. The next challenge is making agents more accessible and simple enough that the technical details fade into the background.

In this episode of The Deep View Conversations, we sit down with two members of OpenAI's ChatGPT Work team, Tara Seshan and Ty Geri, to dig into ChatGPT Work and what OpenAI is doing to make advanced agent capabilities useful to a lot more people. We also dig into some of the current challenges and how the team is approaching them.

Seshan and Geri explain how scheduled tasks and proactive assistance are changing the way people start their workdays, why AI lets teams move from debating ideas to testing prototypes, and how personalized software can turn one-off needs into purpose-built tools. They also discuss the challenge of token costs and model selection, why "super app" isn't the most useful framing for ChatGPT and Codex, and what it will take for agents to become more persistent, proactive, and connected.

The conversation also covers:
• How OpenAI is trying to bridge local and cloud workflows
• Why Tara and Ty start their days with agents instead of Slack
• Building personal apps and tools without traditional software overhead
• The tradeoff between model capability, cost, and user control
• More persistent agents and proactive personal assistance
• Connecting agents to email, calendars, enterprise systems and third-party tools
• Privacy, security and administrative controls for agentic work

If you’re figuring out where agents fit into your work or what has to improve before you trust them with more of it, then this conversation offers a practical look at how OpenAI is preparing for that transition.

Subscribe to Deep View Conversations for interviews with the leaders shaping the future of AI, business, and technology: tdv.transitor.fm

Jason Hiner, Editor-in-Chief

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.