OpenAI CFO Sarah Friar said at Goldman Sachs' Communacopia conference on September 8 that the company has already used its own frontier models in chip design, and that its in-house inference chip, Jalapeno, went from design to tape-out in just nine months. She also named chip design, life sciences, and financial services as three industries OpenAI is prioritizing, citing rising demand for AI built around specialized workflows in those fields.
Jalapeno is an inference chip co-developed by OpenAI and Broadcom, publicly announced in June, built to run OpenAI's own models at lower cost. According to public statements, early samples cost roughly half as much as a typical AI graphics processor. A shipping timeline has not been disclosed.
Nine months is fast for a custom chip. In the industry, taking a new inference ASIC from finalized architecture to tape-out typically takes a year and a half to two years. Friar attributed the compressed timeline to OpenAI's own models, but did not specify which stages they handled — layout and routing, verification, or earlier-stage architecture exploration.
Pitting Luna Against GLM-5.3
The pricing details were the more concrete part of her remarks. Friar said that after an 80% price cut, Luna's input price fell to $0.20 per million tokens and its output price to $1.20, and that usage has since grown roughly tenfold. She went further, pointing the comparison at a Chinese open-weight model:
"If you're deploying Luna and compare that to (Z.ai's) GLM 5.3, for example, on a cloud layer, we are cheaper."
The basis for that comparison deserves scrutiny. Friar compared Luna to "GLM-5.3 deployed on a cloud layer" — that is, the actual per-token price of the open-weight model after being hosted by a third-party cloud — which is a different figure from Zhipu's own API pricing. GLM-5.3's official output price at launch was $4.4 per million tokens. The same model can cost several times more or less depending on whether it runs through the vendor's own API or a third-party cloud, since the latter has to amortize GPU occupancy, memory, and concurrency utilization. OpenAI has not disclosed the full basis for this comparison, nor the throughput assumptions behind it.
Real Pressure on Chinese Open-Weight Models
This pitch is aimed squarely at enterprises that self-host. Over the past two years, companies at home and abroad have chosen open-weight models mainly for two reasons: keeping data in-house, and lower self-hosted cost. The first reason is unaffected. The second is now being squeezed by closed-source price cuts. With Luna's output price down to $1.20 — and a proprietary inference chip behind it promising further cuts — the math behind "self-hosting is cheaper" needs to be redone, especially for teams with lower concurrency and lower GPU utilization.
To be clear, Friar's figures represent the seller's own framing. No third party has yet published a per-token cost comparison between Luna and GLM-5.3 under the same throughput conditions, and Zhipu has not responded to the comparison.
Sources: Reuters, CocoLoop, Qz; multiple reports were cross-checked for Luna's input/output pricing and the scale of its price cut, Jalapeno's design timeline and sample cost ratio, and Friar's remarks in English.