# Inference

Published articles for Inference.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Introducing Amazon SageMaker HyperPod Inference Gateway

DevFeed: [Introducing Amazon SageMaker HyperPod Inference Gateway](<https://devfeed.tech/articles/introducing-amazon-sagemaker-hyperpod-inference-gateway-42780.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/machine-learning/introducing-amazon-sagemaker-hyperpod-inference-gateway/>)

Author: Vinay Arora

Published: 2026-09-18T13:08:34Z

Content type: release

Language: en

Sources: [Artificial Intelligence](<https://devfeed.tech/sources/artificial-intelligence.md>)

Topics: [Amazon SageMaker HyperPod](<https://devfeed.tech/topics/amazon-sagemaker-hyperpod.md>), [Model Routing](<https://devfeed.tech/topics/model-routing.md>), [model-serving](<https://devfeed.tech/topics/model-serving.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Amazon Elastic Kubernetes Service](<https://devfeed.tech/topics/amazon-elastic-kubernetes-service.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Prometheus](<https://devfeed.tech/topics/prometheus.md>)

Tags: [amazon-eks](<https://devfeed.tech/tags/amazon-eks.md>), [amazon-sagemaker-hyperpod](<https://devfeed.tech/tags/amazon-sagemaker-hyperpod.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [expert-400](<https://devfeed.tech/tags/expert-400.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [latency](<https://devfeed.tech/tags/latency.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [model-serving](<https://devfeed.tech/tags/model-serving.md>), [performance](<https://devfeed.tech/tags/performance.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals and model-serving metrics to route inference requests to suitable pods, aiming to reduce GPU waste and first-token latency without application changes.

### Source excerpt

Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals to send each inference request to the best-suited pod, cutting first-token latency by up to 82% with no changes to your model servers or client applications.

## Open-weight models now handle a majority of tokens on Vercel's AI Gateway. But Anthropic still takes 64% of the spend.

DevFeed: [Open-weight models now handle a majority of tokens on Vercel's AI Gateway. But Anthropic still takes 64% of the spend.](<https://devfeed.tech/articles/open-weight-models-now-handle-a-majority-of-tokens-on-vercel-s-ai-gateway-but-anthropic-still-takes-64-of-the-spend-42789.md>)

Original publisher: [Read original article](<https://thenewstack.io/open-weight-anthropic-spend/>)

Author: Paul Sawers

Published: 2026-09-18T12:42:58Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [gateway](<https://devfeed.tech/topics/gateway.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [api](<https://devfeed.tech/tags/api.md>), [developers](<https://devfeed.tech/tags/developers.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [production](<https://devfeed.tech/tags/production.md>), [providers](<https://devfeed.tech/tags/providers.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Vercel reports that open-weight models handled 56% of tokens routed through its AI Gateway in August, up from 7% in December 2025. Despite that volume, Anthropic accounted for 64% of spending, illustrating the difference between token usage and model costs.

### Source excerpt

The trend is clear: open-weight models are taking an increasingly large bite out of production AI usage. On Monday, The The post Open-weight models now handle a majority of tokens on Vercel's AI Gateway. But Anthropic still takes 64% of the spend. appeared first on The New Stack.

## GLM 5.3 FlashX now available on AI Gateway

DevFeed: [GLM 5.3 FlashX now available on AI Gateway](<https://devfeed.tech/articles/glm-5-3-flashx-now-available-on-ai-gateway-42759.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/glm-5-3-flashx-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-09-18T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [gateway](<https://devfeed.tech/topics/gateway.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [api](<https://devfeed.tech/tags/api.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [failover](<https://devfeed.tech/tags/failover.md>), [inference](<https://devfeed.tech/tags/inference.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [retries](<https://devfeed.tech/tags/retries.md>)

### AI overview

Vercel's AI Gateway now supports GLM 5.3 FlashX, a high-speed multimodal coding model from Z.ai. The release targets faster streamed responses for coding agents, tool loops, and interactive applications, with support for unified API access, usage and cost tracking, retries, failover, routing, and reporting.

### Source excerpt

GLM 5.3 FlashX is now available on AI Gateway. GLM 5.3 FlashX is a high-speed serving option for Z.ai's multimodal coding model, delivering inference at ~200 tokens per second for faster streamed responses. The higher serving speed is useful for coding agents, tool loops, and interactive applications where users wait on generated output. Use zai/glm-5.3-flashx across API formats and in coding agents: To use it in a coding agent, see the coding agents guide, then run vercel ai-gateway setup to create a key and configure your supported agents. Select zai/glm-5.3-flashx inside the agent. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try GLM-5.3-FlashX in the model playground, or view all language models available on AI Gateway. Read more

## Run open weight models on Amazon Bedrock in AWS European Sovereign Cloud

DevFeed: [Run open weight models on Amazon Bedrock in AWS European Sovereign Cloud](<https://devfeed.tech/articles/run-open-weight-models-on-amazon-bedrock-in-aws-european-sovereign-cloud-42093.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/security/run-open-weight-models-on-aws-bedrock-in-aws-european-sovereign-cloud/>)

Author: Marta Taggart

Published: 2026-09-17T21:19:20Z

Content type: release

Language: en

Sources: [AWS Security Blog](<https://devfeed.tech/sources/aws-security-blog.md>)

Topics: [Amazon Bedrock](<https://devfeed.tech/topics/amazon-bedrock.md>), [Amazon Web Services (AWS)](<https://devfeed.tech/topics/amazon-web-services-aws.md>), [gemma4](<https://devfeed.tech/topics/gemma4.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [digital sovereignty](<https://devfeed.tech/topics/digital-sovereignty.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [API](<https://devfeed.tech/topics/api.md>), [SDK](<https://devfeed.tech/topics/sdk.md>), [Access Control](<https://devfeed.tech/topics/access-control.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [access-control](<https://devfeed.tech/tags/access-control.md>), [amazon](<https://devfeed.tech/tags/amazon.md>), [amazon-bedrock](<https://devfeed.tech/tags/amazon-bedrock.md>), [api](<https://devfeed.tech/tags/api.md>), [apis](<https://devfeed.tech/tags/apis.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [aws](<https://devfeed.tech/tags/aws.md>), [digital-sovereignty](<https://devfeed.tech/tags/digital-sovereignty.md>), [europe](<https://devfeed.tech/tags/europe.md>), [gemma-4](<https://devfeed.tech/tags/gemma-4.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [inference](<https://devfeed.tech/tags/inference.md>), [least-privilege](<https://devfeed.tech/tags/least-privilege.md>), [security-blog](<https://devfeed.tech/tags/security-blog.md>), [security-identity-compliance](<https://devfeed.tech/tags/security-identity-compliance.md>)

### AI overview

AWS announces that Gemma 4 open-weight models are generally available through Amazon Bedrock in the AWS European Sovereign Cloud, allowing European organizations to run generative AI workloads with EU data residency and digital sovereignty controls. The article also describes the inference engine, OpenAI-compatible APIs, SDK compatibility, and access controls.

### Source excerpt

European organizations can run AI workloads on Amazon Web Services (AWS) while keeping data within the European Union (EU) and meeting regulatory requirements. You can now run generative AI workloads on open weight models on Amazon Bedrock in the AWS European Sovereign Cloud. We're excited to announce the general availability of the first open weight [...]

## Where does all the VRAM go during LLM inference?

DevFeed: [Where does all the VRAM go during LLM inference?](<https://devfeed.tech/articles/where-does-all-the-vram-go-during-llm-inference-42076.md>)

Original publisher: [Read original article](<https://blog.dailydoseofds.com/p/where-does-all-the-vram-go-during>)

Author: Avi Chawla

Published: 2026-09-17T19:30:12Z

Content type: article

Language: en

Sources: [Daily Dose of Data Science](<https://devfeed.tech/sources/daily-dose-of-data-science.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Decoding](<https://devfeed.tech/topics/decoding.md>)

Tags: [concurrency](<https://devfeed.tech/tags/concurrency.md>), [decoding](<https://devfeed.tech/tags/decoding.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model-architecture](<https://devfeed.tech/tags/model-architecture.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [vram](<https://devfeed.tech/tags/vram.md>)

### AI overview

This article explains how GPU memory is allocated during large language model inference. It distinguishes mostly fixed model-weight memory from dynamic memory used by the KV cache, temporary activations and workspace, and runtime overhead. Context length, batch size, concurrency and model architecture affect whether the workload fits, while quantization reduces weight memory but does not guarantee higher throughput.

### Source excerpt

...explained visually

## The guest journey, updated in real time: extending Airbnb's sequence recommender with Chronon

DevFeed: [The guest journey, updated in real time: extending Airbnb's sequence recommender with Chronon](<https://devfeed.tech/articles/the-guest-journey-updated-in-real-time-extending-airbnb-s-sequence-recommender-with-chronon-42165.md>)

Original publisher: [Read original article](<https://medium.com/airbnb-engineering/the-guest-journey-updated-in-real-time-extending-airbnbs-sequence-recommender-with-chronon-8f1582578553?source=rss----53c7c27702d5---4>)

Author: Pengyu Hou

Published: 2026-09-17T17:01:02Z

Content type: article

Language: en

Sources: [The Airbnb Tech Blog - Medium](<https://devfeed.tech/sources/the-airbnb-tech-blog-medium.md>)

Topics: [real-time](<https://devfeed.tech/topics/real-time.md>), [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [data](<https://devfeed.tech/topics/data.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [data](<https://devfeed.tech/tags/data.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [latency](<https://devfeed.tech/tags/latency.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [repo](<https://devfeed.tech/tags/repo.md>), [results](<https://devfeed.tech/tags/results.md>), [sequence](<https://devfeed.tech/tags/sequence.md>), [technology](<https://devfeed.tech/tags/technology.md>)

### AI overview

Airbnb describes extending its sequence-based recommender with Chronon's Push Mode and Near-real-time Model Transform capabilities. The changes update guest activity and model features more quickly, reducing the staleness of search-ranking inputs compared with the previous daily batch pipeline.

### Source excerpt

How two new Chronon capabilities, Push Mode and NRT Model Transform, allows us to provide more relevant search results instantly as a guest explores, rather than waiting for the next batch run. By: Pengyu Hou, Yuli Han, Daochen Zha, Haozhen Ding, Xin Liu, Sophie Wang, Pallavi Adusumilli, Sherry Li, Henry Saputra, Chun How Tan, Huiji Gao, Yan Zhang, Stephanie Moyerman, Yi Li, and Sanjeev Katariya A guest's interaction with Airbnb doesn't pause to wait for a nightly batch job. Someone might browse a dozen listings on a Tuesday afternoon, run a new search that evening, and expect the next search to reflect the recent activity; it's also to Airbnb's benefit for that to be the case. In our previous post, Personalizing Airbnb search by learning from the guest journey, we described how we built a Transformer-based sequence encoder that creates better, more personalized search rankings for a guest using the booking, review, and browsing data that is most relevant to them -- their own. That system ran as a daily batch job: each night it processed the previous day's activity and refreshed embeddings for guests who had something new to show for it. That design worked well, but it left a gap. Activity from earlier the same day wouldn't show up in the embedding until the following day's run, on top of the pipeline's own processing lag -- in practice, up to nearly two days of staleness. For a guest actively planning a trip, that meant the ranking model was often working from a slightly outdated picture of what they wanted, and the recent activities are often highly relevant to current search needs. This is a limitation that our original JourneyFormer research had already flagged as needing new serving infrastructure to solve. In this post, we describe how we closed that gap by adding two new capabilities to Chronon, Airbnb's feature platform: Near-real-time Model Transform and Push Mode. Chronon is an open source project, and these capabilities have been contributed back to our pub

## Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend

DevFeed: [Open-weight models take 56% of token volume, Astra doubles Fable 5.1 spend](<https://devfeed.tech/articles/open-weight-models-take-56-of-token-volume-astra-doubles-fable-5-1-spend-42109.md>)

Original publisher: [Read original article](<https://vercel.com/blog/ai-gateway-production-index-september-2026>)

Author: Eric Dodds

Published: 2026-09-17T07:00:00Z

Content type: article

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [gateway](<https://devfeed.tech/topics/gateway.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [gpt-6-astra](<https://devfeed.tech/topics/gpt-6-astra.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [fewer](<https://devfeed.tech/tags/fewer.md>), [gpt-6-astra](<https://devfeed.tech/tags/gpt-6-astra.md>), [growth](<https://devfeed.tech/tags/growth.md>), [inference](<https://devfeed.tech/tags/inference.md>), [median](<https://devfeed.tech/tags/median.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [production](<https://devfeed.tech/tags/production.md>), [tasks](<https://devfeed.tech/tags/tasks.md>), [volume](<https://devfeed.tech/tags/volume.md>)

### AI overview

Vercel's September 2026 AI Gateway Production Index reports that open-weight models handled 56% of gateway tokens in August, up from fewer than one in ten in December 2025. It also reports falling token prices, shifting model spend, and rapid adoption of OpenAI's GPT-6 Astra.

### Source excerpt

AI Gateway Production Index -- September 2026 Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from June, July, and August. September 2026 summary The September index reports on AI Gateway data collected through August 2026. Open-weight models ran the majority of gateway tokens for the first time, up from 7% in December to 56% in August. The average token costs less than half what it did five months ago. Price per token fell 23.2% in August, the third straight monthly drop, and the median team paid 7.6% less. Fable 5, Anthropic's most capable model, lost two-thirds of its share of gateway spend in one month. Opus 5, at half the price, tripled its share. Anthropic kept 64% of all spend. Gemini 3 Flash has lost 95% of its share of gateway tokens since May, and more than three-quarters of the volume it lost went to models from other labs. Special report: OpenAI launches Astra on September 3 GPT-6 Astra took a third of OpenAI's spend within two days of launch and twice Fable 5.1's share of gateway spend. Introduced two days apart at the same price, Astra took 7.7% of all gateway spend in its first twelve days, while Fable 5.1 took 3.7%. Open-weight models take a majority of token volume for the first time In August, open-weight models ran 56% of all tokens on AI Gateway, marking the first month they took the majority of volume. In December 2025, they processed fewer than one in ten tokens, and only eight months later, they ran more token volume than all closed-weight models combined. Though the frontier kept the majority of spend, open-weight dollar share is accelerating. As open-weight models become more capable, customers are moving more production workloads over to them. Growth in open-weight model adoption helped push the average price per token across the gatew

## Fujitsu ready to sell its custom 'Monaka' Arm chip, maybe to rival server-makers

DevFeed: [Fujitsu ready to sell its custom 'Monaka' Arm chip, maybe to rival server-makers](<https://devfeed.tech/articles/fujitsu-ready-to-sell-its-custom-monaka-arm-chip-maybe-to-rival-server-makers-41316.md>)

Original publisher: [Read original article](<https://www.theregister.com/systems/2026/09/17/fujitsu-ready-to-sell-its-custom-monaka-arm-chip-maybe-to-rival-server-makers/5297025>)

Author: Simon Sharwood

Published: 2026-09-17T06:20:33Z

Content type: news

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [Arm](<https://devfeed.tech/topics/arm.md>), [Server](<https://devfeed.tech/topics/server.md>)

Tags: [arm](<https://devfeed.tech/tags/arm.md>), [fujitsu](<https://devfeed.tech/tags/fujitsu.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [inference](<https://devfeed.tech/tags/inference.md>), [server](<https://devfeed.tech/tags/server.md>), [sovereign-cloud](<https://devfeed.tech/tags/sovereign-cloud.md>), [supercomputer](<https://devfeed.tech/tags/supercomputer.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

The article reports that Fujitsu is preparing to sell its custom Monaka Arm chip to other server makers, with potential interest from cloud and inference-focused customers.

### Source excerpt

Clouds, the sovereign-sensitive, and the inferencing-interested are also about to get sales calls

## Intern 2 offers an isolated environment and cloud inference for running personal bots

DevFeed: [Intern 2 offers an isolated environment and cloud inference for running personal bots](<https://devfeed.tech/articles/if-a-mac-mini-is-agentic-overkill-try-this-glowing-pyramid-that-runs-personal-bots-31537.md>)

Original publisher: [Read original article](<https://www.theregister.com/personal-tech/2026/09/17/if-a-mac-mini-is-agentic-overkill-try-this-glowing-pyramid-that-runs-personal-bots/5296964>)

Author: Thomas Claburn

Published: 2026-09-16T23:28:59Z

Content type: news

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [Bots](<https://devfeed.tech/topics/bots.md>), [Mac Mini](<https://devfeed.tech/topics/mac-mini.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-and-ml](<https://devfeed.tech/tags/ai-and-ml.md>), [automate](<https://devfeed.tech/tags/automate.md>), [autonomous-ai](<https://devfeed.tech/tags/autonomous-ai.md>), [bots](<https://devfeed.tech/tags/bots.md>), [edge-computing](<https://devfeed.tech/tags/edge-computing.md>), [inference](<https://devfeed.tech/tags/inference.md>), [intern-2](<https://devfeed.tech/tags/intern-2.md>), [mac-mini](<https://devfeed.tech/tags/mac-mini.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [personal-tech](<https://devfeed.tech/tags/personal-tech.md>)

### AI overview

Intern 2 provides an isolated environment and cloud inference for automating tasks with personal bots.

### Source excerpt

'Intern 2' offers an isolated environment and cloudy inference to automate your world

## NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

DevFeed: [NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut](<https://devfeed.tech/articles/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-31524.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/>)

Author: Zhihan Jiang

Published: 2026-09-16T15:00:48Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [NVIDIA Vera Rubin](<https://devfeed.tech/topics/nvidia-vera-rubin.md>), [Vera Rubin NVL72](<https://devfeed.tech/topics/vera-rubin-nvl72.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Dynamo](<https://devfeed.tech/topics/dynamo.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [TensorRT-LLM](<https://devfeed.tech/topics/tensorrt-llm.md>)

Tags: [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [dynamo](<https://devfeed.tech/tags/dynamo.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [mlperf](<https://devfeed.tech/tags/mlperf.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [software](<https://devfeed.tech/tags/software.md>), [tensorrt](<https://devfeed.tech/tags/tensorrt.md>), [tensorrt-llm](<https://devfeed.tech/tags/tensorrt-llm.md>), [vera-rubin-nvl72](<https://devfeed.tech/tags/vera-rubin-nvl72.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

NVIDIA reports MLPerf Inference v6.1 preview results for Vera Rubin NVL72 and GB300 NVL72 systems. Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x higher throughput on DeepSeek-R1, while a four-rack GB300 NVL72 submission achieved 99% scaling efficiency. The results used vLLM, NVIDIA Dynamo, and TensorRT-LLM.

### Source excerpt

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. [...]

## MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin's First Peer-Reviewed Numbers

DevFeed: [MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin's First Peer-Reviewed Numbers](<https://devfeed.tech/articles/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubin-s-first-peer-reviewed-numbers-31404.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/mlperf-inference-v6-1-5-7x-per-accelerator-gains-a-512-gpu-run-and-vera-rubins-first-peer-reviewed-numbers>)

Author: Harold Fritts

Published: 2026-09-16T15:00:00Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [agentic-coding](<https://devfeed.tech/topics/agentic-coding.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Vera Rubin NVL72](<https://devfeed.tech/topics/vera-rubin-nvl72.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Vera Rubin](<https://devfeed.tech/topics/vera-rubin.md>)

Tags: [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [gpt-oss](<https://devfeed.tech/tags/gpt-oss.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [numbers](<https://devfeed.tech/tags/numbers.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [qwen3](<https://devfeed.tech/tags/qwen3.md>), [rag](<https://devfeed.tech/tags/rag.md>)

### AI overview

MLCommons published MLPerf Inference v6.1 with record participation, two new inference tests, and peer-reviewed results for several newly covered accelerators. The release reports a 5.7x improvement in the best per-accelerator DeepSeek-R1 server result compared with v5.1.

### Source excerpt

MLCommons has published MLPerf Inference v6.1, and the round sets a participation record with 30 submitting organizations and 486 datacenter and edge results. Two new tests join the suite: an End-to-End RAG pipeline for the datacenter and an Edge Agentic Inference benchmark for single-user devices, and the results carry the first peer-reviewed numbers for NVIDIA's The post MLPerf Inference v6.1: 5.7x Per-Accelerator Gains, a 512-GPU Run, and Vera Rubin's First Peer-Reviewed Numbers appeared first on StorageReview.com.

## University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK

DevFeed: [University of Manchester Uses NVIDIA Earth-2 to Forecast Air Pollution Across the UK](<https://devfeed.tech/articles/university-of-manchester-uses-nvidia-earth-2-to-forecast-air-pollution-across-the-uk-30917.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/uk-air-pollution-research-earth-2/>)

Author: Isha Salian

Published: 2026-09-16T05:00:42Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Supercomputing](<https://devfeed.tech/topics/supercomputing.md>), [DGX Spark](<https://devfeed.tech/topics/dgx-spark.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [data](<https://devfeed.tech/topics/data.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [real-time](<https://devfeed.tech/topics/real-time.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-for-good](<https://devfeed.tech/tags/ai-for-good.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [climate](<https://devfeed.tech/tags/climate.md>), [compute](<https://devfeed.tech/tags/compute.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [government](<https://devfeed.tech/tags/government.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [science](<https://devfeed.tech/tags/science.md>), [supercomputing](<https://devfeed.tech/tags/supercomputing.md>), [training](<https://devfeed.tech/tags/training.md>), [uk](<https://devfeed.tech/tags/uk.md>)

### AI overview

The University of Manchester is working with NVIDIA to use Earth-2 generative AI models to forecast air pollution across the U.K. The team trained Earth-2 CorrDiff on chemistry-climate simulation data using Isambard-AI, added StormCast for time-dependent forecasts using air-quality observations, and demonstrated workflows on DGX Spark.

### Source excerpt

Air pollution is a serious public health risk, contributing to an estimated 30,000 deaths in the U.K. alone last year. Data-driven insights can help -- but computing air quality with traditional chemistry-based models is expensive, which limits how detailed they can be and how regularly they can be run. David Topping, a professor in the [...]

## CVITEK CV1842H-P-based edge AI camera module offers night vision and AI-ISP support (Crowdfunding)

DevFeed: [CVITEK CV1842H-P-based edge AI camera module offers night vision and AI-ISP support (Crowdfunding)](<https://devfeed.tech/articles/cvitek-cv1842h-p-based-edge-ai-camera-module-offers-night-vision-and-ai-isp-support-crowdfunding-27005.md>)

Original publisher: [Read original article](<https://www.cnx-software.com/2026/09/16/cvitek-cv1842h-p-based-edge-ai-camera-module-offers-night-vision-and-ai-isp-support/>)

Author: Debashis Das

Published: 2026-09-16T00:00:55Z

Content type: news

Language: en

Sources: [CNX Software - Embedded Systems News](<https://devfeed.tech/sources/cnx-software-embedded-systems-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Embedded Systems](<https://devfeed.tech/topics/embedded-systems.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [V](<https://devfeed.tech/topics/v.md>), [Arm](<https://devfeed.tech/topics/arm.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Ethernet](<https://devfeed.tech/topics/ethernet.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Toit](<https://devfeed.tech/topics/toit.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [arm](<https://devfeed.tech/tags/arm.md>), [artificial-intelligence-ai](<https://devfeed.tech/tags/artificial-intelligence-ai.md>), [camera](<https://devfeed.tech/tags/camera.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [debug](<https://devfeed.tech/tags/debug.md>), [edge-ai](<https://devfeed.tech/tags/edge-ai.md>), [embedded](<https://devfeed.tech/tags/embedded.md>), [ethernet](<https://devfeed.tech/tags/ethernet.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [kickstarter](<https://devfeed.tech/tags/kickstarter.md>), [linux](<https://devfeed.tech/tags/linux.md>), [module](<https://devfeed.tech/tags/module.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [risc-v](<https://devfeed.tech/tags/risc-v.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [rt-thread](<https://devfeed.tech/tags/rt-thread.md>), [soc](<https://devfeed.tech/tags/soc.md>), [sophgo](<https://devfeed.tech/tags/sophgo.md>), [tinyml](<https://devfeed.tech/tags/tinyml.md>), [usb](<https://devfeed.tech/tags/usb.md>), [video](<https://devfeed.tech/tags/video.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

The AIMORELOGY Ovis is an open-source modular AI vision camera built around the CVITEK CV1842H-P SoC. It provides full-color 1080p night vision, 1.5 TOPS edge AI inference, USB and Ethernet connectivity, and a dual-OS environment using Linux and RT-Thread.

### Source excerpt

The AIMORELOGY Ovis is an open-source AI vision camera module built around the CVITEK CV1842H-P SoC, with full-color 1080p night vision and 1.5 TOPS edge AI inference in a compact modular design. It is designed for drones, robotics, security systems, smart cameras, and custom embedded vision products. The camera features a compact stacked design, with a 20 x 20 mm Core board that includes the CVITEK CV1842H-P SoC, 2 Gbit NAND flash, USB, and UART debug pads. A Sensor board with the SC235HAI image sensor connects on top and also adds Ethernet and UART interfaces. An optional CVBS board goes between the Sensor board and the Ovis Core board. AIMORELOGY Ovis specifications: Ovis Core Board SoC - CVITEK CV1842H-P CPU - 1x Arm Cortex-A53 core @ 1.1 GHz, 1x RISC-V C906 core @ 800 MHz NPU - 1.5 TOPS @ INT8 with BF16 support ISP - AI-ISP with real-time 1080p [...] The post CVITEK CV1842H-P-based edge AI camera module offers night vision and AI-ISP support (Crowdfunding) appeared first on CNX Software - Embedded Systems News.

## Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation

DevFeed: [Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation](<https://devfeed.tech/articles/trajectory-as-the-teacher-few-step-discrete-flow-matching-via-energy-navigated-distillation-31491.md>)

Original publisher: [Read original article](<https://machinelearning.apple.com/research/trajectory-teacher-flow-matching>)

Published: 2026-09-16T00:00:00Z

Content type: article

Language: en

Sources: [Apple Machine Learning Research](<https://devfeed.tech/sources/apple-machine-learning-research.md>)

Topics: [text-generation](<https://devfeed.tech/topics/text-generation.md>), [Algorithms](<https://devfeed.tech/topics/algorithms.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Inference](<https://devfeed.tech/topics/inference.md>)

Tags: [inference](<https://devfeed.tech/tags/inference.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [perplexity](<https://devfeed.tech/tags/perplexity.md>), [research](<https://devfeed.tech/tags/research.md>), [text-generation](<https://devfeed.tech/tags/text-generation.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

The article introduces Trajectory-Shaped Discrete Flow Matching, a training method that guides intermediate trajectory decisions with an energy-based coherence measure. The authors argue that poor distillation trajectories, rather than insufficient student capacity, limit few-step generation. On a 170M-parameter language-modeling task, an 8-step student reportedly achieves lower perplexity than a 1,024-step teacher while reducing inference steps.

### Source excerpt

Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Distillation uses the multi-step trajectory to train a student to reproduce the process in a few steps. When the student underperforms, the usual explanation is insufficient capacity. We argue the opposite: the trajectory is the bottleneck, not the student. Each training trajectory is built through a chain of blind stochastic jumps with no evaluation of sequence quality; a single bad decision at an early midpoint propagates through subsequent steps...

## ML based ranking using Nrtsearch

DevFeed: [ML based ranking using Nrtsearch](<https://devfeed.tech/articles/ml-based-ranking-using-nrtsearch-31461.md>)

Original publisher: [Read original article](<https://engineeringblog.yelp.com/2026/09/ml-ranking-with-nrtsearch.html>)

Author: Mohammad Mohtasham (Software Engineer); Tao Yu (Software Engineer)

Published: 2026-09-16T00:00:00Z

Content type: article

Language: en

Sources: [Yelp](<https://devfeed.tech/sources/yelp.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [recommendation systems](<https://devfeed.tech/topics/recommendation-systems.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [bridge](<https://devfeed.tech/tags/bridge.md>), [inference](<https://devfeed.tech/tags/inference.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [overhead](<https://devfeed.tech/tags/overhead.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [service](<https://devfeed.tech/tags/service.md>)

### AI overview

Yelp's Nrtsearch Inference Plugin embeds machine-learning ranking directly in the search layer. The article explains the scoring workflow, including model configuration, feature extraction, candidate ranking, and application-specific business logic. It describes how co-locating feature storage and inference reduces network transfer, serialization overhead, and latency compared with a standalone inference service.

### Source excerpt

We've extended Nrtsearch with the Inference Plugin, which embeds ML-based ranking directly in the search layer -- eliminating the need for a standalone scoring service. We use Nrtsearch (read more information on the blog post), a Lucene-based open-source search engine built by Yelp, to power a variety of applications such as business search, reviews search, ad delivery and photo search. In this blog post, we give a high-level overview of the Machine Learning (ML) based scoring workflow in Nrtsearch. We'll show how ML models are configured and loaded, and how different applications use custom business logic to develop, test, and...

## Seagate and WD AI Storage Research Finds Enterprises Rank Storage Above Compute as the AI Bottleneck

DevFeed: [Seagate and WD AI Storage Research Finds Enterprises Rank Storage Above Compute as the AI Bottleneck](<https://devfeed.tech/articles/seagate-and-wd-ai-storage-research-finds-enterprises-rank-storage-above-compute-as-the-ai-bottleneck-26756.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/seagate-and-wd-ai-storage-research-finds-enterprises-rank-storage-above-compute-as-the-ai-bottleneck>)

Author: Lyle Smith

Published: 2026-09-15T17:23:54Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [idc](<https://devfeed.tech/topics/idc.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Retrieval-Augmented Generation](<https://devfeed.tech/topics/retrieval-augmented-generation.md>), [synthetic-data](<https://devfeed.tech/topics/synthetic-data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [compute](<https://devfeed.tech/tags/compute.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [genai](<https://devfeed.tech/tags/genai.md>), [hdd](<https://devfeed.tech/tags/hdd.md>), [idc](<https://devfeed.tech/tags/idc.md>), [inference](<https://devfeed.tech/tags/inference.md>), [reports](<https://devfeed.tech/tags/reports.md>), [research](<https://devfeed.tech/tags/research.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [storage](<https://devfeed.tech/tags/storage.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>)

### AI overview

Seagate and WD published separate studies indicating that AI is increasing enterprise storage requirements and extending data retention. Although their headline percentages differ because they asked different questions, both reports point to storage becoming a larger part of AI infrastructure planning alongside growing archive and retrieval needs.

### Source excerpt

Seagate and WD published separate AI storage studies within days of each other; the headline numbers: Seagate says 99% of enterprises expect AI to increase their storage requirements over the next three years, while WD's IDC research puts the comparable figure at 74%. Read the fine print, and both reports land in the same directional The post Seagate and WD AI Storage Research Finds Enterprises Rank Storage Above Compute as the AI Bottleneck appeared first on StorageReview.com.

## From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

DevFeed: [From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production](<https://devfeed.tech/articles/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production-26943.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/>)

Author: Vishal Ganeriwala

Published: 2026-09-15T16:55:59Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [NVIDIA DSX](<https://devfeed.tech/topics/nvidia-dsx.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [compute](<https://devfeed.tech/tags/compute.md>), [dsx](<https://devfeed.tech/tags/dsx.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [production](<https://devfeed.tech/tags/production.md>)

### AI overview

The article describes how Emerald AI's Conductor platform responds to utility demand signals by adjusting flexible data-center workloads while keeping high-priority AI inference running. It also reports that Lambda's validation found a fixed power budget could support 24% more token throughput when managed intelligently.

### Source excerpt

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram was watching on Zoom with about forty others -- his team at Emerald AI in their San Francisco conference room, engineers [...]

## AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

DevFeed: [AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories](<https://devfeed.tech/articles/ai-infra-summit-nvidia-vera-rubin-and-dsx-platform-advancements-showcase-energy-efficiencies-of-optimizing-tokens-per-watt-for-ai-factories-26942.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/>)

Author: NVIDIA Writers

Published: 2026-09-15T16:55:40Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [DSX](<https://devfeed.tech/topics/dsx.md>), [Vera Rubin](<https://devfeed.tech/topics/vera-rubin.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Conversational AI](<https://devfeed.tech/topics/conversational-ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [dsx](<https://devfeed.tech/tags/dsx.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infra](<https://devfeed.tech/tags/infra.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency-inference](<https://devfeed.tech/tags/low-latency-inference.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [nvidia-dsx](<https://devfeed.tech/tags/nvidia-dsx.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>)

### AI overview

NVIDIA's AI Infra Summit coverage describes collaborations and platform updates focused on improving AI factory efficiency. The article highlights Vera Rubin systems, DSX MaxLPS, Dynamo inference software, NVLink and networking technologies, including claims of up to 1.4x more tokens per megawatt through factory-wide power optimization.

### Source excerpt

Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech. Before a packed audience -- with more than 8,000 attendees this year, up from 3,500 last year -- [...]

## How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin

DevFeed: [How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin](<https://devfeed.tech/articles/how-nvidia-groq-3-lpx-deterministic-execution-drives-power-efficient-high-interactivity-inference-on-nvidia-vera-rubin-26913.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-nvidia-groq-3-lpx-deterministic-execution-drives-power-efficient-high-interactivity-inference-on-nvidia-vera-rubin/>)

Author: Tanya Lenz

Published: 2026-09-15T16:55:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Groq 3 LPX](<https://devfeed.tech/topics/groq-3-lpx.md>), [LPX](<https://devfeed.tech/topics/lpx.md>), [NVIDIA Vera Rubin](<https://devfeed.tech/topics/nvidia-vera-rubin.md>), [Vera Rubin NVL72](<https://devfeed.tech/topics/vera-rubin-nvl72.md>), [groq](<https://devfeed.tech/topics/groq.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Low Latency](<https://devfeed.tech/topics/low-latency.md>), [long-context](<https://devfeed.tech/topics/long-context.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [ai-platform](<https://devfeed.tech/tags/ai-platform.md>), [drive](<https://devfeed.tech/tags/drive.md>), [dsx](<https://devfeed.tech/tags/dsx.md>), [groq](<https://devfeed.tech/tags/groq.md>), [groq-3-lpx](<https://devfeed.tech/tags/groq-3-lpx.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [lpx](<https://devfeed.tech/tags/lpx.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [performance](<https://devfeed.tech/tags/performance.md>), [power-management](<https://devfeed.tech/tags/power-management.md>), [vera-rubin](<https://devfeed.tech/tags/vera-rubin.md>), [vera-rubin-nvl72](<https://devfeed.tech/tags/vera-rubin-nvl72.md>)

### AI overview

This NVIDIA developer article explains how Groq 3 LPX uses deterministic execution across 256 LPU chips to support low-latency inference on NVIDIA Vera Rubin. It describes compiler-scheduled execution and power-management techniques including Preemptive Power and Clock Period Synthesis.

### Source excerpt

Power is a defining constraint for AI factories. As AI workloads demand a full compute platform to serve them, each component of that platform must maximize...

## Build an AI-powered product tagging system with Amazon SageMaker serverless model customization

DevFeed: [Build an AI-powered product tagging system with Amazon SageMaker serverless model customization](<https://devfeed.tech/articles/build-an-ai-powered-product-tagging-system-with-amazon-sagemaker-serverless-model-customization-26940.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/machine-learning/build-an-ai-powered-product-tagging-system-with-amazon-sagemaker-serverless-model-customization/>)

Author: Linpo Guo

Published: 2026-09-15T16:11:36Z

Content type: tutorial

Language: en

Sources: [Artificial Intelligence](<https://devfeed.tech/sources/artificial-intelligence.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Amazon SageMaker AI](<https://devfeed.tech/topics/amazon-sagemaker-ai.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [rlvr](<https://devfeed.tech/topics/rlvr.md>), [grpo](<https://devfeed.tech/topics/grpo.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [SDK](<https://devfeed.tech/topics/sdk.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [amazon-sagemaker](<https://devfeed.tech/tags/amazon-sagemaker.md>), [amazon-sagemaker-ai](<https://devfeed.tech/tags/amazon-sagemaker-ai.md>), [aws](<https://devfeed.tech/tags/aws.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cost](<https://devfeed.tech/tags/cost.md>), [customization](<https://devfeed.tech/tags/customization.md>), [expert-400](<https://devfeed.tech/tags/expert-400.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference](<https://devfeed.tech/tags/inference.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [rlvr](<https://devfeed.tech/tags/rlvr.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>)

### AI overview

This walkthrough shows how to build a product tagging system by customizing Qwen3-8B with supervised fine-tuning and reinforcement learning with verifiable rewards on Amazon SageMaker serverless model customization. It then deploys the optimized model for asynchronous inference to enrich retail catalogs.

### Source excerpt

Manually tagging thousands of catalog products is slow and inconsistent. This walkthrough shows how to customize Qwen3-8B with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) on Amazon SageMaker serverless model customization, then deploy it for asynchronous inference to build a cost-efficient product tagging system.

## Measuring and Improving Consistency in Repeated Agent Runs

DevFeed: [Measuring and Improving Consistency in Repeated Agent Runs](<https://devfeed.tech/articles/your-agent-aced-the-task-will-it-do-it-again-26920.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ibm-research/altk-evolve-consistency>)

Author: Evelyn Duesterwald; Lilian Ngweta; Vatche Isahagian; Jayaram Radhakrishnan; Vinod Muthusamy; Gaodan Fang; Ashwath Vaithinathan Aravindan; Punleuk Oum; G Thomas; Merve Unuvar; Ayhan Sebin; Michał Ulewi

Published: 2026-09-15T16:00:44Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [consistency](<https://devfeed.tech/tags/consistency.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [reports](<https://devfeed.tech/tags/reports.md>), [standard](<https://devfeed.tech/tags/standard.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

This article presents the Consistency Analyzer, a diagnostic for finding decision points where an agent's behavior may change across repeated runs. It introduces consistency guidelines in ALTK-Evolve and reports that they reduced the consistency gap from 24.4 percentage points to 12.0 points without reducing average accuracy.

### Source excerpt

That is embarrassing onstage. In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request. For mission-critical work, such as reconciling a financial transaction or checking a contract for an obligation, that can be a showstopper. Most benchmarks hide this variability behind an average. On AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions.

## Axelera Europa Ships: 629 TOPS at 45W Per AIPU, in Validated Dell XE5 and Supermicro Servers

DevFeed: [Axelera Europa Ships: 629 TOPS at 45W Per AIPU, in Validated Dell XE5 and Supermicro Servers](<https://devfeed.tech/articles/axelera-europa-ships-629-tops-at-45w-per-aipu-in-validated-dell-xe5-and-supermicro-servers-26751.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/axelera-europa-ships-629-tops-at-45w-per-aipu-in-validated-dell-xe5-and-supermicro-servers>)

Author: Harold Fritts

Published: 2026-09-15T13:00:00Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Computer vision](<https://devfeed.tech/topics/computer-vision.md>), [servers](<https://devfeed.tech/topics/servers.md>), [dell](<https://devfeed.tech/topics/dell.md>), [RISC-V](<https://devfeed.tech/topics/riscv.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [dell](<https://devfeed.tech/tags/dell.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [inference](<https://devfeed.tech/tags/inference.md>), [risc-v](<https://devfeed.tech/tags/risc-v.md>), [servers](<https://devfeed.tech/tags/servers.md>)

### AI overview

Axelera AI is shipping Europa, a second-generation AI Processing Unit, in bare-chip and PCIe card configurations. The company says the 45W device delivers 629 TOPS and supports on-premises inference workloads including generative AI, vision-language models, and computer vision. The Edge 232p card is shipping in validated Dell XE5 and Supermicro 111AD systems.

### Source excerpt

Axelera AI is shipping Europa, the second-generation AI Processing Unit (AIPU) it has been previewing since last year, and it's launching with validated servers from Dell and Supermicro attached. The Eindhoven company's pitch is inference on infrastructure the customer controls: agentic systems, vision-language models, generative AI, and computer vision running in a standard rackmount server The post Axelera Europa Ships: 629 TOPS at 45W Per AIPU, in Validated Dell XE5 and Supermicro Servers appeared first on StorageReview.com.

## Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

DevFeed: [Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts](<https://devfeed.tech/articles/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts-17436.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts>)

Author: Harold Fritts

Published: 2026-09-14T16:23:21Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Multi-tenancy](<https://devfeed.tech/topics/multi-tenancy.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [TensorRT](<https://devfeed.tech/topics/tensorrt.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [cache](<https://devfeed.tech/tags/cache.md>), [concurrent](<https://devfeed.tech/tags/concurrent.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [latency](<https://devfeed.tech/tags/latency.md>), [nvme](<https://devfeed.tech/tags/nvme.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

Lightbits Labs is introducing Inferra, a KV cache orchestration engine for AI inference. It virtualizes GPU memory across DRAM and NVMe storage, preserving attention states for long-context and multi-session workloads. Lightbits claims up to 16 times more concurrent sessions, more than 100 times lower latency than recomputation, and context windows of up to 10 million tokens. Inferra supports vLLM, TensorRT, and SGLang and includes tiering, predictive prefetching, tenant isolation, and encrypted data transfer.

### Source excerpt

Lightbits Labs, the company that invented NVMe over TCP, is moving into inference software with Inferra, a KV cache orchestration engine that makes its public debut tomorrow, September 15, at the AI Infra Summit in Santa Clara. The software virtualizes GPU memory across DRAM and NVMe storage tiers and turns the KV cache into a The post Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts appeared first on StorageReview.com.

## On-Device AI Series (Part 5): LiteRT-LM

DevFeed: [On-Device AI Series (Part 5): LiteRT-LM](<https://devfeed.tech/articles/on-device-ai-series-part-5-litert-lm-22949.md>)

Original publisher: [Read original article](<https://proandroiddev.com/on-device-ai-series-part-5-litert-lm-d6c23b102094?source=rss----c72404660798---4>)

Author: Oğuzhan Aslan

Published: 2026-09-14T05:59:12Z

Content type: tutorial

Language: en

Sources: [ProAndroidDev - Medium](<https://devfeed.tech/sources/proandroiddev-medium.md>)

Topics: [LiteRT](<https://devfeed.tech/topics/litert.md>), [On-device AI](<https://devfeed.tech/topics/on-device-ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [android](<https://devfeed.tech/tags/android.md>), [android-development](<https://devfeed.tech/tags/android-development.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [litert](<https://devfeed.tech/tags/litert.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [on-device-ai](<https://devfeed.tech/tags/on-device-ai.md>), [programming](<https://devfeed.tech/tags/programming.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [software-development](<https://devfeed.tech/tags/software-development.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

This tutorial explains LiteRT-LM for running large language models on-device. It covers the Engine/Session API, streaming output, system prompts, tool calling, multimodal inputs, thinking mode, and CPU-versus-GPU benchmarking. The article also discusses tradeoffs involving privacy, network independence, latency, memory, sampling configuration, and model capability compared with cloud APIs.

### Source excerpt

Put your phone in airplane mode. Open the app, type a question, and watch the answer arrive one token at a time -- no spinner waiting on a network round-trip, no API key, no per-token bill, and nothing you typed ever leaving the device. LiteRT-LM removes the genuinely hard parts of running an LLM on-device -- KV-cache management, token streaming, backend selection -- but it doesn't remove your job so much as relocate it. What's left on your plate is a short, specific list: sizing a combined input+output token budget, owning your own sampling defaults, hand-building system prompts and tool calling out of raw text, and one native-library collision that presents as a SIGSEGV rather than a build error. Know those going in and the API itself is a clean three-step pattern. We'll get there in that order: Why you'd choose this runtime and what it costs you versus the cloud. The Engine/Session model you need to read the code at all. Real implementation samples -- streaming, system prompts and tool calling, multimodal inputs, thinking mode, and CPU-vs-GPU benchmarking. The anti-patterns to avoid. A developer-friendliness rating on the same rubric as Parts 1-4. Why Use LiteRT-LM? You reach for LiteRT-LM instead of hand-rolling generation on top of raw LiteRT when: You need multi-turn conversation, not single-shot inference -- session state and KV-cache bookkeeping are handled for you, and resetting a conversation is a session swap, not a model reload. You need streaming output -- token-by-token delivery for a responsive chat UI, instead of a blocking call that returns everything at once. You're choosing between CPU and GPU per device -- the explicit backend parameter turns that into a runtime decision instead of a build-time guess. You want a pre-converted model without doing your own PyTorch-to-LiteRT conversion work -- the Model Zoo covers Gemma, Qwen, Llama, and more out of the box. You're willing to own sampling -- the engine won't pick sane decoding defaults for you; that's on the

[Next page](<https://devfeed.tech/tags/inference.md?cursor=WyIyMDI2LTA5LTE0VDA1OjU5OjEyKzAwOjAwIiwgIjE0MGY1NTUyLWExOTItNDdmOS1iN2ZhLWYwNjY5OWM5MGM2YSJd>)