# inference efficiency

Published articles for inference efficiency.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## How OpenAI Optimized AI Agents for Codex and ChatGPT Work at Scale

DevFeed: [How OpenAI Optimized AI Agents for Codex and ChatGPT Work at Scale](<https://devfeed.tech/articles/how-openai-optimized-ai-agents-for-scale-34931.md>)

Original publisher: [Read original article](<https://newsletter.eng-leadership.com/p/how-openai-optimized-ai-agents-for>)

Author: Gregor Ojstersek

Published: 2026-07-29T19:52:55Z

Content type: article

Language: en

Sources: [Engineering Leadership](<https://devfeed.tech/sources/engineering-leadership.md>)

Topics: [OpenAI](<https://devfeed.tech/topics/openai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [codex](<https://devfeed.tech/topics/codex.md>), [ChatGPT](<https://devfeed.tech/topics/chatgpt.md>), [API](<https://devfeed.tech/topics/api.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [api](<https://devfeed.tech/tags/api.md>), [chatgpt](<https://devfeed.tech/tags/chatgpt.md>), [codex](<https://devfeed.tech/tags/codex.md>), [harness](<https://devfeed.tech/tags/harness.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-efficiency](<https://devfeed.tech/tags/inference-efficiency.md>), [openai](<https://devfeed.tech/tags/openai.md>)

### AI overview

The article examines how OpenAI optimized the harness, API, and inference efficiency of AI agents supporting Codex and ChatGPT Work at scale.

### Source excerpt

Inside the Harness, API, and Inference efficiency optimizations that support the fast growth of Codex and ChatGPT Work.

## Making LLMs faster without sacrificing accuracy

DevFeed: [Making LLMs faster without sacrificing accuracy](<https://devfeed.tech/articles/making-llms-faster-without-sacrificing-accuracy-7603.md>)

Original publisher: [Read original article](<https://www.amazon.science/blog/making-llms-faster-without-sacrificing-accuracy>)

Author: Tao Yu; Youngsuk Park

Published: 2026-05-15T13:00:00Z

Content type: article

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [Language models](<https://devfeed.tech/topics/language-models.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [web applications](<https://devfeed.tech/topics/web-applications.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [chinchilla-scaling-law](<https://devfeed.tech/tags/chinchilla-scaling-law.md>), [conversational-ai](<https://devfeed.tech/tags/conversational-ai.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [grouped-query-attention](<https://devfeed.tech/tags/grouped-query-attention.md>), [hyperparameter-optimization](<https://devfeed.tech/tags/hyperparameter-optimization.md>), [iclr-2026](<https://devfeed.tech/tags/iclr-2026.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-efficiency](<https://devfeed.tech/tags/inference-efficiency.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [llm-optimization](<https://devfeed.tech/tags/llm-optimization.md>), [llms](<https://devfeed.tech/tags/llms.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [network-architectures](<https://devfeed.tech/tags/network-architectures.md>), [scaling-laws](<https://devfeed.tech/tags/scaling-laws.md>), [training](<https://devfeed.tech/tags/training.md>), [transformer-architecture](<https://devfeed.tech/tags/transformer-architecture.md>)

### AI overview

The article presents scaling laws that connect LLM architectural choices to the tradeoff between accuracy and efficiency. It describes how these choices can improve inference throughput without reducing accuracy.

### Source excerpt

A new scaling law that relates particular architectural choices to loss helps identify models that improve throughput by up to 47% with no loss of accuracy.

## Holotron-12B - High Throughput Computer Use Agent

DevFeed: [Holotron-12B - High Throughput Computer Use Agent](<https://devfeed.tech/articles/holotron-12b-high-throughput-computer-use-agent-7008.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Hcompany/holotron-12b>)

Author: Pierre-Louis Cedoz; Hamza Benchekroun; Aurélien Lac; Delfosse; Tony Wu; Antoine Bonnet; Kai Yuan; Aleix Cambray; Alexandra; Axel Moyal

Published: 2026-03-17T12:33:39Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-efficiency](<https://devfeed.tech/tags/inference-efficiency.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [model](<https://devfeed.tech/tags/model.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance](<https://devfeed.tech/tags/performance.md>), [production](<https://devfeed.tech/tags/production.md>), [release](<https://devfeed.tech/tags/release.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

Holotron-12B is a released multimodal computer-use policy model optimized for production throughput, long contexts, and agentic workloads. The article attributes its inference efficiency to a hybrid state-space-model and attention architecture, and reports over twice the throughput of Holo2-8B in a WebVoyager evaluation using vLLM on one H100 GPU.

### Source excerpt

We're thrilled to release Holotron-12B, a multimodal computer-use model from H Company. Post-trained from the open NVIDIA Nemotron-Nano-2 VL model on H Company's proprietary data mixture, Holotron-12B is the result of a close collaboration between our research labs to engineer a new type of model optimized primarily for scale and performance in production. H Company is part of the NVIDIA Inception Program. The model is now available on Hugging Face.

## T5Gemma: A new collection of encoder-decoder Gemma models

DevFeed: [T5Gemma: A new collection of encoder-decoder Gemma models](<https://devfeed.tech/articles/t5gemma-a-new-collection-of-encoder-decoder-gemma-models-6250.md>)

Original publisher: [Read original article](<https://deepmind.google/blog/t5gemma-a-new-collection-of-encoder-decoder-gemma-models/>)

Author: Biao Zhang; Paul Suganthan; Ben Hora

Published: 2025-10-25T18:14:00Z

Content type: article

Language: en

Sources: [Google DeepMind News](<https://devfeed.tech/sources/google-deepmind-news.md>)

Topics: [gemma](<https://devfeed.tech/topics/gemma.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [LLM Techniques](<https://devfeed.tech/topics/llm-techniques.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-efficiency](<https://devfeed.tech/tags/inference-efficiency.md>), [llms](<https://devfeed.tech/tags/llms.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [research](<https://devfeed.tech/tags/research.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

The article introduces T5Gemma, a collection of encoder-decoder large language models created by adapting pretrained decoder-only Gemma 2 models. It describes pretrained and instruction-tuned variants, flexible encoder-decoder configurations, and reported quality and inference-efficiency advantages across benchmarks such as SuperGLUE.

### Source excerpt

Introducing T5Gemma, a new collection of encoder-decoder LLMs.