• The Deep View
  • Posts
  • How OpenAI's Jalapeno chip surprised the industry

How OpenAI's Jalapeno chip surprised the industry

Welcome back. Perplexity and Nvidia are betting a new brand of local agents can cut token costs, speed up performance, and keep sensitive data private. Apple’s also thinking about local AI, and its new M6 Mac mini and M5 Ultra Mac Studio show how seriously Cupertino now is taking this opportunity. Meanwhile, OpenAI’s first custom inference chip, Jalapeño, beats Nvidia Blackwell on performance per watt in a SemiAnalysis benchmark. It's a striking result for silicon developed in under two years, with the help of OpenAI's latest models. Jason Hiner

IN TODAY’S NEWSLETTER

1. How OpenAI's Jalapeno chip surprised the industry

2. Apple's new M5, M6 Macs reveal local AI strategy

3. 3 reasons AI agents are moving to your desktop

HARDWARE

OpenAI Jalapeno is even bigger than the benchmarks

OpenAI has built its first AI chip, and the early returns on performance look impressive.

On Tuesday at the Hot Chips conference, OpenAI said that its custom inference chip, Jalapeño, is now up and running on working silicon and is already in testing. The frontier lab revealed the first third-party benchmarks and claims that it's accelerating performance-per-watt, the favorite measurement of efficiency in the AI industry. It's basically a way of showing how much AI work a chip can complete for every unit of power it consumes. 

On tests against various types of workloads, OpenAI claims that Jalapeño delivers 1.5 to 4 times the performance-per-watt compared to today's leading AI chips. 

OpenAI used SemiAnalysis's open-source InferenceX benchmark to test Jalapeño, using three models: GPT-OSS (OpenAI's open-source model), DeepSeek R1, and Kimi K2.5. In other words, it didn't just test Jalapeño on its own models fine-tuned to work on its own hardware. 

And while its own GPT-OSS model performed the best, the other models also performed remarkably well. The biggest surprise was that Jalapeño outperformed Nvidia's Blackwell, one of the current industry standards for AI chips. However, we don't yet know how it will perform against other inference accelerator leaders like Groq and Cerebres.

SemiAnalysis characterized it this way: "Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point in the curve. It excels not only in low-latency scenarios but also in high-throughput scenarios."

In a briefing with the media, I asked OpenAI's head of hardware, Richard Ho, where OpenAI would use this newfound power in its AI stack (because obviously it won't be running all workloads on its own chips anytime soon). "It's clear that the cost of serving low latency with this device is a lot lower than other low latency devices," said Ho, "and so there may be a good niche there where some of our very latency sensitive products such as Codex Ultra Fast Mode — even faster than the [fastest mode] we have now — may be a good option… We are expecting to ramp up the volume into 2027."

Other than the performance itself, the most impressive thing about Jalapeño may be how fast OpenAI was able to develop it and bring it into production. The project, which is a partnership with Broadcom, began as a concept in late 2024. The fact that it already has working silicon less than two years later is just as surprising as the benchmark performance. OpenAI credits its own AI models to helping the hardware team dramatically accelerate the process. The team reported that OpenAI's models aided in the optimization, programming, and development of the chips and are currently doing the same for the next generations that come after Jalapeño. While the speed is great and will help meet user expectations over time, the real win for OpenAI is in efficiency. If it can design and optimize its own chips then it can spend far less on compute and lower the long-term cost of intelligence to bring its most advanced models and features to more users.

Jason Hiner, Editor-in-Chief

TOGETHER WITH QUIQ

7 AI agents, one seamless journey

Most companies still think of AI as a chatbot for support questions. But 72% of buyers now expect a personalized experience at every touchpoint, and companies deploying agentic AI across the full customer journey are seeing up to a 30% increase in conversion rates, according to McKinsey research.

In Quiq’s new guide, “The Clock Is Ticking: 7 AI Agents Every Leader Needs,” we break down the seven essential agents, from the conversational ad that first engages a customer to the AI analyst that catches churn risk after every interaction, and where each one fits across the journey.

PRODUCTS

Apple's new M5, M6 Macs reveal local AI strategy

The Mac mini has helped fuel the AI agent revolution. Now Apple is ready to extend that momentum with its next-generation model.

On Tuesday, Apple unveiled its latest chips, the M6 in the new Mac mini and the M5 Ultra in the new Mac Studio, which the company is positioning as AI powerhouses. This launch follows months of Mac mini shortages earlier in 2026, when demand, driven largely by AI enthusiasts running OpenClaw and other agents as persistent, always-on machines, outstripped Apple's supply, compounded by a broader industry-wide shortage in memory and storage.

The Mac mini, the popular go-to for local and private AI inference on a budget, is now available with the all-new M6, which Apple says delivers up to 4x faster AI performance, 2x faster storage and graphics, and 40% faster CPU performance while retaining the same attractive, compact form factor it is known for. The M6 chipset features a 12-core CPU and 12-core GPU, two more cores than before in each, with the GPU adding Neural Accelerators to each core for the first time on Mac mini, powering many of these improvements. 

The M5 Pro version is meant to take it up a notch, with pro-level performance for more power-intensive applications such as video production and gaming. The Mac mini with M6 starts at $899, and the Mac mini with M5 Pro starts at $1,699, both $100 more than their predecessors. 

"With a more powerful CPU and graphics, Neural Accelerators in the GPU, and higher memory bandwidth, Mac mini with M6 delivers a whole new level of AI performance," said Johny Srouji, Apple’s chief hardware officer, in a blog post. 

For those who require even more power, Apple's Mac Studio also got 4.3x faster AI performance, according to Apple. Other upgrades include up to 2x faster storage, up to 1.8x faster graphics, and up to 1.3x faster CPU speed. The new M5 Ultra is Apple's most powerful silicon yet, featuring an 18-core CPU, up to 40-core GPU with Neural Accelerators built into each core, and up to 128GB of unified memory. 

The Mac Studio with M5 Max highlights include an 18-core CPU, a GPU with up to 40 cores, each with Neural Accelerators built into each core, and up to 128GB of unified memory. The Mac Studio with M5 Max starts at $2,499, and the Mac Studio with M5 Ultra starts at $5,499. 

Pre-orders for all the new Macs start today, and shipping starts on Sept 22, the same day the new Macs arrive in Apple stores.

Apple's Mac mini and Mac Studio have always been geared towards power users, often led by creative professionals, and that's typically the language Apple would use at launch to describe the product to the market. It is interesting how much has changed, with so much of the messaging geared towards AI and explaining the improvements made to better power those workflows. It is also interesting to see how the explosion of AI workflows has shifted demand for these products, as the Mac Studio blog post even includes a paragraph explaining how people could use Mac Studios with the M5 Ultra chip in clusters to take their local AI inference to the next level. This is a very purposeful AI launch ahead of what will likely be Apple's biggest iPhone launch in years this September, with Siri 2.0 finally going live, its new foldable phone expected to launch, and the possibility of its first AI hardware products with the AI AirPods with cameras and/or a desktop AI speaker.

TOGETHER WITH DESCOPE

The biggest MCP spec update since June 2025 just landed

Every team shipping MCP servers should pay attention to the July 2026 spec revision. Sessions are gone. DCR is deprecated in favor of CIMD. And there are six new authorization SEPs your server is now expected to handle.

This developer breakdown from Descope covers:

  • What changed in the transport layer, and why sessions were removed

  • The six new authorization SEPs explained

  • Enterprise-Managed Authorization (EMA) and what it means for consent sprawl

  • A migration checklist to run against your existing MCP server

PRODUCTS

Why AI agents are moving to your desktop

Perplexity is doing what it does best. Releasing something that feels like others could soon copy. 

But this time, Nvidia is joining in as a partner. 

On Tuesday, Perplexity announced Portable Computer, a version of its AI agent that can run open models locally on your own desktop hardware to do three things: 

  1. Save significant money on tokens

  2. Get faster performance

  3. Boost data privacy

That's the kind of stuff more and more people have been trying to do because of soaring token bills due to AI agents and because of wanting to run agents on their most sensitive data, which they'd prefer to run locally and not send to the latest frontier models in the cloud from Anthropic, OpenAI, Google, and others. 

Where Nvidia comes in is that Portable Computer will run on Nvidia's DGX Spark machine, which has cult favorite status among AI builders. It's a tiny box the size of Mac mini, but with the power of a Mac Studio for running AI. Spark has 128GB of unified memory and runs up to a petaflop of compute. Nvidia says it runs models up to 200B parameters and can fine-tune models up to 70B.

Perplexity's Personal Computer ran on a Mac mini at launch, so a big part of the announcement is Perplexity's AI agent now running on PC hardware powered by Nvidia GPUs. That means it's likely to come to other DGX Spark competitors like the ones from Dell, Lenovo, Asus, MSI, and others. Right now, the software only runs on Linux. But Perplexity is also working on a version that will run on Windows, as well as a version that will run on Nvidia's more powerful version of DGX Spark called DGX Station, which is the size of a full computer tower like a Mac Pro.

The other big part of the announcement, of course, is being able to run Perplexity's agent on open models that run locally on your machine. At launch, Portable Computer will run a post-trained version of Alibaba's Qwen 3.8 27B called PPLX 27B, with a version of Nvidia's Nemotron 3.5 Lightning coming to the device soon, according to Perplexity. You can also use a local inference server like Ollama or LM Studio to run any open model you'd like, but that starts to get a little more complicated.

The problem that Perplexity and Nvidia are trying to solve is that more and more people are trying their hand at AI agents and would like to save the token costs and get the privacy you can get by running these models locally, but it's a very involved process that requires time and technical expertise.

"Now that we're seeing people actually run frontier intelligence on their desk, the remaining bottleneck becomes that it's quite cumbersome to set up," said Nate Kupp, VP of infrastructure at Perplexity, in a briefing with the media. "We really focused on making this a straightforward experience where you can get up and running very quickly."

While DGX Spark and competitors cost $4,000-$5,000, they can save you so much in token costs that AI builders don't even blink at that price. Heavy AI agent users can spend up to twice that a month in token costs. And if Portable Computer catches on, and other AI companies offer something similar, then it could increase competition and drive down the price tag.

Perplexity's best attribute is arguably how fast it can execute. It's small and nimble and, again and again, the team has shown the ability to launch things quickly. Its AI agent was an idea that was hatched in about a month and launched around the same time OpenClaw went viral in early 2026. Perplexity's general purpose agent can now help you code your own software like Claude Code or OpenAI's Codex, but it can also help you carry out knowledge work tasks the same way Claude Cowork and ChatGPT Work can. The ability to run locally on open models is a super power. But even highly technical people have complained about how involved the setup is to get Nvidia's DGX Spark up and running, so this would be a win for both Nvidia and AI enthusiasts if Perplexity can make that process more streamlined. 

Jason Hiner, Editor-in-Chief

LINKS

GAMES

Which image is real?

Login or Subscribe to participate in polls.

A QUICK POLL BEFORE YOU GO

Have you ever used Perplexity?

Login or Subscribe to participate in polls.

The Deep View is written by Nat Rubio-Licht, Sabrina Ortiz, Jason Hiner, Faris Kojok and The Deep View crew. Please reply with any feedback.

Thanks for reading today’s edition of The Deep View! We’ll see you in the next one.

“Cracks and changes in the pavement looked real.”

“[The other image] looks too generic and lacks real motion where as [this image] feels alive.”


“The pavement repair just didn't seem like something AI would generate.”


“The use of wires attached to the hanging street lamp was more realistic.”

“The shadows in [this image] are wrong.”


“30 speed limit only one way.”


“Distance perspective is wrong.”


“License plate in [this image] has garbage symbols. Also, my first impression was that the building in the background didn't look real.”

If you want to get in front of an audience of 750,000+ developers, business leaders and tech enthusiasts, get in touch with us here.