# Jalapeño's first results show industry-leading speed and efficiency in AI inference

DevFeed: [Jalapeño's first results show industry-leading speed and efficiency in AI inference](<https://devfeed.tech/articles/jalapeno-s-first-results-show-industry-leading-speed-and-efficiency-in-ai-inference-6521.md>)

Original publisher: [Read original article](<https://openai.com/index/jalapeno-first-results>)

Published: 2026-08-25T07:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [gpt-oss](<https://devfeed.tech/topics/gpt-oss.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [performance](<https://devfeed.tech/tags/performance.md>), [software](<https://devfeed.tech/tags/software.md>), [speed](<https://devfeed.tech/tags/speed.md>), [systems](<https://devfeed.tech/tags/systems.md>)

## AI overview

OpenAI reports initial results for Jalapeño, its custom inference chip. The article says the chip delivers higher throughput, lower end-to-end latency, and greater AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T, based on tests using the InferenceX benchmark.

## Source excerpt

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.