# AI Infrastructure

Published articles for AI Infrastructure.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Perplexity's AI agents helped build a database. They weren't allowed to run it.

DevFeed: [Perplexity's AI agents helped build a database. They weren't allowed to run it.](<https://devfeed.tech/articles/perplexity-s-ai-agents-helped-build-a-database-they-weren-t-allowed-to-run-it-31533.md>)

Original publisher: [Read original article](<https://thenewstack.io/perplexity-cobbledb-ai-database/>)

Author: Amanda Caswell

Published: 2026-09-16T21:51:15Z

Content type: article

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Database](<https://devfeed.tech/topics/database.md>), [DynamoDB](<https://devfeed.tech/topics/dynamodb.md>), [Rust](<https://devfeed.tech/topics/rust.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [rocksdb](<https://devfeed.tech/topics/rocksdb.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [api](<https://devfeed.tech/tags/api.md>), [coding](<https://devfeed.tech/tags/coding.md>), [coding-agents](<https://devfeed.tech/tags/coding-agents.md>), [database](<https://devfeed.tech/tags/database.md>), [databases](<https://devfeed.tech/tags/databases.md>), [dynamodb](<https://devfeed.tech/tags/dynamodb.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [perplexity](<https://devfeed.tech/tags/perplexity.md>), [rocksdb](<https://devfeed.tech/tags/rocksdb.md>), [rust](<https://devfeed.tech/tags/rust.md>), [s3](<https://devfeed.tech/tags/s3.md>)

### AI overview

Perplexity built CobbleDB, a Rust key-value store, after finding DynamoDB too costly and insufficiently controllable for its search workload. Coding agents helped develop it, but were not allowed to run it in production. Perplexity measured lower read latency and expects lower costs, with plans to open-source the database.

### Source excerpt

Perplexity decided it was paying too much for DynamoDB and wasn't getting the control it wanted over read performance. So The post Perplexity's AI agents helped build a database. They weren't allowed to run it. appeared first on The New Stack.

## Anthropic merges Claude Chat and Cowork into one interface

DevFeed: [Anthropic merges Claude Chat and Cowork into one interface](<https://devfeed.tech/articles/anthropic-bet-users-were-choosing-wrong-so-it-removed-the-choice-31531.md>)

Original publisher: [Read original article](<https://thenewstack.io/anthropic-claude-unified-interface/>)

Author: Amanda Caswell

Published: 2026-09-16T16:46:36Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [chat](<https://devfeed.tech/tags/chat.md>), [claude](<https://devfeed.tech/tags/claude.md>), [connectors](<https://devfeed.tech/tags/connectors.md>), [context](<https://devfeed.tech/tags/context.md>), [developer-tools](<https://devfeed.tech/tags/developer-tools.md>), [tools](<https://devfeed.tech/tags/tools.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

Anthropic is merging Claude Chat and Cowork into a single interface, allowing conversations to handle simple questions, multi-step projects, connected tools, and background execution. Claude Docs and Claude Slides are launching in beta on paid plans, while Claude Design is moving into conversations.

### Source excerpt

Using Claude for anything beyond a quick question has always started with a routing decision to use Chat or Cowork? The post Anthropic bet users were choosing wrong. So it removed the choice. appeared first on The New Stack.

## AI leaders propose embedded third-party evaluators for frontier AI safety

DevFeed: [AI leaders propose embedded third-party evaluators for frontier AI safety](<https://devfeed.tech/articles/ai-evaluator-the-most-important-ai-job-in-history-how-developers-might-fill-the-proposed-new-job-31530.md>)

Original publisher: [Read original article](<https://thenewstack.io/ai-embedded-evaluator-jobs/>)

Author: Adrian Bridgwater

Published: 2026-09-16T15:58:41Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Job](<https://devfeed.tech/topics/job.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Frontier AI](<https://devfeed.tech/topics/frontier-ai.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-strategy](<https://devfeed.tech/tags/ai-strategy.md>), [alignment](<https://devfeed.tech/tags/alignment.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [developers](<https://devfeed.tech/tags/developers.md>), [frontier-ai](<https://devfeed.tech/tags/frontier-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [job](<https://devfeed.tech/tags/job.md>), [meta](<https://devfeed.tech/tags/meta.md>), [openai](<https://devfeed.tech/tags/openai.md>), [safety](<https://devfeed.tech/tags/safety.md>), [tech-careers](<https://devfeed.tech/tags/tech-careers.md>)

### AI overview

The article examines Anthropic CEO Dario Amodei's proposal for frontier AI companies to provide embedded third-party evaluators with employee-like access. These evaluators would verify safety practices, report incidents, and assess AI models, training pipelines, and processes.

### Source excerpt

The pace of frontier AI model development spurred Anthropic CEO Dario Amodei to publish an essay last weekend, calling for The post AI evaluator: The most important AI job in history? How developers might fill the proposed new job appeared first on The New Stack.

## NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

DevFeed: [NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut](<https://devfeed.tech/articles/nvidia-vera-rubin-nvl72-delivers-leading-performance-in-mlperf-inference-v6-1-debut-31524.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/vera-rubin-nvl72-mlperf-inference/>)

Author: Zhihan Jiang

Published: 2026-09-16T15:00:48Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [NVIDIA Vera Rubin](<https://devfeed.tech/topics/nvidia-vera-rubin.md>), [Vera Rubin NVL72](<https://devfeed.tech/topics/vera-rubin-nvl72.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Dynamo](<https://devfeed.tech/topics/dynamo.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [TensorRT-LLM](<https://devfeed.tech/topics/tensorrt-llm.md>)

Tags: [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [dynamo](<https://devfeed.tech/tags/dynamo.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [mlperf](<https://devfeed.tech/tags/mlperf.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvl72](<https://devfeed.tech/tags/nvl72.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [software](<https://devfeed.tech/tags/software.md>), [tensorrt](<https://devfeed.tech/tags/tensorrt.md>), [tensorrt-llm](<https://devfeed.tech/tags/tensorrt-llm.md>), [vera-rubin-nvl72](<https://devfeed.tech/tags/vera-rubin-nvl72.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

NVIDIA reports MLPerf Inference v6.1 preview results for Vera Rubin NVL72 and GB300 NVL72 systems. Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x higher throughput on DeepSeek-R1, while a four-rack GB300 NVL72 submission achieved 99% scaling efficiency. The results used vLLM, NVIDIA Dynamo, and TensorRT-LLM.

### Source excerpt

System performance, efficient infrastructure scaling and continuous software optimization are key levers that determine AI inference economics. Higher system performance means more tokens generated, resulting in higher revenue. Efficient scaling means throughput grows proportionally as hardware gets added, requiring fewer resources to serve users at scale. Continuous optimization means generating more value from infrastructure investments. [...]

## Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers

DevFeed: [Emerald AI, Google and NVIDIA Launch Alliance to Advance Flexible AI Data Centers](<https://devfeed.tech/articles/emerald-ai-google-and-nvidia-launch-alliance-to-advance-flexible-ai-data-centers-30916.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/ai-energy-management-alliance/>)

Author: Josh Parker

Published: 2026-09-16T13:00:33Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [data centers](<https://devfeed.tech/topics/data-centers.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Google](<https://devfeed.tech/topics/google.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Requirements](<https://devfeed.tech/topics/requirements.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-data-centers](<https://devfeed.tech/tags/ai-data-centers.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [corporate](<https://devfeed.tech/tags/corporate.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [energy](<https://devfeed.tech/tags/energy.md>), [google](<https://devfeed.tech/tags/google.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [launch](<https://devfeed.tech/tags/launch.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance-metrics](<https://devfeed.tech/tags/performance-metrics.md>), [requirements](<https://devfeed.tech/tags/requirements.md>), [resource](<https://devfeed.tech/tags/resource.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

Emerald AI, Google and NVIDIA announced the AI Energy Management Alliance, a coalition focused on flexible AI data centers that can dynamically adjust electricity use in response to grid conditions. The article describes technology-neutral, performance-based requirements covering response speed, duration, predictability and emergency behavior.

### Source excerpt

AI factories are the infrastructure of the intelligence era. Scaling them responsibly will depend as much on innovation across the grid as inside the data center. Today, Emerald AI, Google and NVIDIA announced the launch of the AI Energy Management Alliance (AEMA), a first-of-its-kind coalition advancing data centers that can dynamically manage their electricity use [...]

## From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production

DevFeed: [From Megawatts to Tokens: How NVIDIA Maximizes AI Factory Production](<https://devfeed.tech/articles/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production-26943.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/from-megawatts-to-tokens-how-nvidia-maximizes-ai-factory-production/>)

Author: Vishal Ganeriwala

Published: 2026-09-15T16:55:59Z

Content type: article

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [NVIDIA DSX](<https://devfeed.tech/topics/nvidia-dsx.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [compute](<https://devfeed.tech/tags/compute.md>), [dsx](<https://devfeed.tech/tags/dsx.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [production](<https://devfeed.tech/tags/production.md>)

### AI overview

The article describes how Emerald AI's Conductor platform responds to utility demand signals by adjusting flexible data-center workloads while keeping high-priority AI inference running. It also reports that Lambda's validation found a fixed power budget could support 24% more token throughput when managed intelligently.

### Source excerpt

On a sweltering August evening in Silicon Valley, as the sun dropped and air conditioning loads spiked, Silicon Valley Power sent a signal to an AI factory to adjust its power consumption. Varun Sivaram was watching on Zoom with about forty others -- his team at Emerald AI in their San Francisco conference room, engineers [...]

## AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories

DevFeed: [AI Infra Summit: NVIDIA Vera Rubin and DSX Platform Advancements Showcase Energy Efficiencies of Optimizing Tokens Per Watt for AI Factories](<https://devfeed.tech/articles/ai-infra-summit-nvidia-vera-rubin-and-dsx-platform-advancements-showcase-energy-efficiencies-of-optimizing-tokens-per-watt-for-ai-factories-26942.md>)

Original publisher: [Read original article](<https://blogs.nvidia.com/blog/ai-infra-summit-vera-rubin-dsx-energy-efficiencies-tokens-per-watt-ai-factories/>)

Author: NVIDIA Writers

Published: 2026-09-15T16:55:40Z

Content type: news

Language: en

Sources: [NVIDIA Blog](<https://devfeed.tech/sources/nvidia-blog.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [DSX](<https://devfeed.tech/topics/dsx.md>), [Vera Rubin](<https://devfeed.tech/topics/vera-rubin.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Conversational AI](<https://devfeed.tech/topics/conversational-ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [dsx](<https://devfeed.tech/tags/dsx.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infra](<https://devfeed.tech/tags/infra.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency-inference](<https://devfeed.tech/tags/low-latency-inference.md>), [networking](<https://devfeed.tech/tags/networking.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvidia-blackwell](<https://devfeed.tech/tags/nvidia-blackwell.md>), [nvidia-dsx](<https://devfeed.tech/tags/nvidia-dsx.md>), [nvidia-vera](<https://devfeed.tech/tags/nvidia-vera.md>), [nvidia-vera-rubin](<https://devfeed.tech/tags/nvidia-vera-rubin.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>)

### AI overview

NVIDIA's AI Infra Summit coverage describes collaborations and platform updates focused on improving AI factory efficiency. The article highlights Vera Rubin systems, DSX MaxLPS, Dynamo inference software, NVLink and networking technologies, including claims of up to 1.4x more tokens per megawatt through factory-wide power optimization.

### Source excerpt

Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, Tuesday spoke on AI factory efficiency at the AI Infra Summit, the Santa Clara Convention Center event that has morphed into a Coachella of infrastructure tech. Before a packed audience -- with more than 8,000 attendees this year, up from 3,500 last year -- [...]

## How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories

DevFeed: [How NVIDIA NVLink 6 Delivers Multi-Layer Resiliency for AI Factories](<https://devfeed.tech/articles/how-nvidia-nvlink-6-delivers-multi-layer-resiliency-for-ai-factories-26914.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-nvidia-nvlink-6-delivers-multi-layer-resiliency-for-ai-factories/>)

Author: Elizabeth Goodman

Published: 2026-09-15T16:55:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [AI Factory](<https://devfeed.tech/topics/ai-factory.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-factory](<https://devfeed.tech/tags/ai-factory.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [industry](<https://devfeed.tech/tags/industry.md>), [networking](<https://devfeed.tech/tags/networking.md>), [networking-communications](<https://devfeed.tech/tags/networking-communications.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>), [resiliency](<https://devfeed.tech/tags/resiliency.md>), [vera-rubin](<https://devfeed.tech/tags/vera-rubin.md>)

### AI overview

The article describes how NVIDIA NVLink 6 supports resiliency in large-scale AI factories. It explains that Vera Rubin NVL72 connects 72 Rubin GPUs into a single scale-up domain and outlines a multilayer approach using lossless networking, error correction, retry, flow control, and error containment to support continuous training and inference operations.

### Source excerpt

For operators of large-scale AI factories, maximizing continuous output is essential for productivity. In massive-scale AI training, every GPU in the cluster...

## How Everpure proposes reducing GPU idle time by improving AI data access

DevFeed: [How Everpure proposes reducing GPU idle time by improving AI data access](<https://devfeed.tech/articles/how-everpure-plans-to-stop-ai-from-starving-without-data-26617.md>)

Original publisher: [Read original article](<https://www.theregister.com/ai-ml/2026/09/15/sponsored-how-everpure-plans-to-stop-ai-from-starving-without-data/5295812>)

Author: Chris Mellor

Published: 2026-09-15T08:00:00Z

Content type: article

Language: en

Sources: [www.theregister.com - Articles](<https://devfeed.tech/sources/www-theregister-com-articles.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-ml](<https://devfeed.tech/tags/ai-ml.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [rag](<https://devfeed.tech/tags/rag.md>), [sponsored](<https://devfeed.tech/tags/sponsored.md>)

### AI overview

This sponsored feature describes Everpure's approach to reducing GPU idle time in AI systems by improving access to large-scale insurance data. It discusses central metadata indexing, storage performance, self-describing data, and integration with Nvidia GPU infrastructure for AI agents and retrieval-augmented generation.

### Source excerpt

SPONSORED FEATURE: The vendor's AI solutions are dedicated to increasing GPU utilization and avoiding costly GPUs doing nothing while waiting for data

## Perplexity's new agent runs entirely on your GPU -- with one expensive catch

DevFeed: [Perplexity's new agent runs entirely on your GPU -- with one expensive catch](<https://devfeed.tech/articles/perplexity-s-new-agent-runs-entirely-on-your-gpu-with-one-expensive-catch-21600.md>)

Original publisher: [Read original article](<https://thenewstack.io/perplexity-portable-computer-windows/>)

Author: Amanda Caswell

Published: 2026-09-14T18:21:44Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [GPU](<https://devfeed.tech/topics/gpu.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Windows](<https://devfeed.tech/topics/windows.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [browser](<https://devfeed.tech/topics/browser.md>), [Ollama](<https://devfeed.tech/topics/ollama.md>)

Tags: [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [news](<https://devfeed.tech/tags/news.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [ollama](<https://devfeed.tech/tags/ollama.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [windows](<https://devfeed.tech/tags/windows.md>)

### AI overview

The article reports that Perplexity's Portable Computer, a local version of its Computer agent, is available in the Perplexity app for Windows on compatible Nvidia GeForce RTX and RTX PRO GPUs. It requires at least 24GB of VRAM and combines local models, orchestration, a browser, tool calling, and a proprietary SPACE sandbox. The article also discusses platform-specific engineering, external service connectors, and the boundary between local and cloud computing.

### Source excerpt

Running an LLM on your PC is easy enough, but putting an agent to work there is a different story. The post Perplexity's new agent runs entirely on your GPU -- with one expensive catch appeared first on The New Stack.

## Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November

DevFeed: [Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November](<https://devfeed.tech/articles/fujitsu-monaka-server-brings-2nm-144-core-cpus-to-air-cooled-ai-inference-on-sale-in-november-17435.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/fujitsu-monaka-server-brings-2nm-144-core-cpus-to-air-cooled-ai-inference-on-sale-in-november>)

Author: Lyle Smith

Published: 2026-09-14T18:03:44Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [data centers](<https://devfeed.tech/topics/data-centers.md>), [Confidential Computing](<https://devfeed.tech/topics/confidential-computing.md>), [Arm](<https://devfeed.tech/topics/arm.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [NVLink](<https://devfeed.tech/topics/nvlink.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [arm](<https://devfeed.tech/tags/arm.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [fujitsu](<https://devfeed.tech/tags/fujitsu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [nvlink](<https://devfeed.tech/tags/nvlink.md>)

### AI overview

Fujitsu is introducing MONAKA Servers built around its 2nm FUJITSU-MONAKA processor for AI inference in air-cooled data centers. The servers offer up to 144 CPU cores, matrix instructions, SVE2 vector processing, hardware-level confidential computing, and planned NVLink Fusion integration with NVIDIA GPUs. Fujitsu claims higher inference throughput and reduced cooling power consumption, but the article notes that supporting benchmark details are unavailable.

### Source excerpt

Fujitsu is bringing its 2nm FUJITSU-MONAKA processor to AI infrastructure with a new server family designed to run AI inference in air-cooled data centers without requiring specialized liquid cooling. The MONAKA Server is designed, developed, and manufactured in Japan, with component and manufacturing traceability for sovereign AI deployments. The first MONAKA Servers will come in The post Fujitsu MONAKA Server Brings 2nm 144-Core CPUs to Air-Cooled AI Inference, On Sale in November appeared first on StorageReview.com.

## Chinese AI models dominate OpenRouter's US token consumption. It can now guarantee that traffic stays entirely in the US.

DevFeed: [Chinese AI models dominate OpenRouter's US token consumption. It can now guarantee that traffic stays entirely in the US.](<https://devfeed.tech/articles/chinese-ai-models-dominate-openrouter-s-us-token-consumption-it-can-now-guarantee-that-traffic-stays-entirely-in-the-us-21599.md>)

Original publisher: [Read original article](<https://thenewstack.io/openrouter-us-region-routing/>)

Author: Paul Sawers

Published: 2026-09-14T13:59:33Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [Security](<https://devfeed.tech/topics/security.md>), [data](<https://devfeed.tech/topics/data.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [availability](<https://devfeed.tech/tags/availability.md>), [data](<https://devfeed.tech/tags/data.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [models](<https://devfeed.tech/tags/models.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [openai](<https://devfeed.tech/tags/openai.md>), [routing](<https://devfeed.tech/tags/routing.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

OpenRouter has launched US in-region routing for business and enterprise customers. Requests sent through its US endpoint are decrypted, processed, and served entirely inside the United States, or rejected if that cannot be guaranteed. The feature addresses concerns about data location when businesses use open-weight models, including models developed in China.

### Source excerpt

Everyone knows the open-weight model pitch by now: companies can download the weights, customize them, run them on infrastructure of The post Chinese AI models dominate OpenRouter's US token consumption. It can now guarantee that traffic stays entirely in the US. appeared first on The New Stack.

## OpenSearch Wins Analytics & Data Intelligence Solutions Category in the SiliconANGLE TechForward Awards

DevFeed: [OpenSearch Wins Analytics & Data Intelligence Solutions Category in the SiliconANGLE TechForward Awards](<https://devfeed.tech/articles/opensearch-wins-analytics-data-intelligence-solutions-category-in-the-siliconangle-techforward-awards-17450.md>)

Original publisher: [Read original article](<https://opensearch.org/announcements/opensearch-wins-analytics-data-intelligence-solutions-category-in-the-siliconangle-techforward-awards/>)

Author: Kristi Piechnik

Published: 2026-09-14T12:00:14Z

Content type: news

Language: en

Sources: [OpenSearch](<https://devfeed.tech/sources/opensearch.md>)

Topics: [Open Source](<https://devfeed.tech/topics/open-source.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [observability](<https://devfeed.tech/topics/observability.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Security](<https://devfeed.tech/topics/security.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [real-time](<https://devfeed.tech/topics/real-time.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [awards](<https://devfeed.tech/tags/awards.md>), [data](<https://devfeed.tech/tags/data.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [opensearch](<https://devfeed.tech/tags/opensearch.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [recognition](<https://devfeed.tech/tags/recognition.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [retrieval-augmented-generation-rag](<https://devfeed.tech/tags/retrieval-augmented-generation-rag.md>), [search](<https://devfeed.tech/tags/search.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

OpenSearch won the Analytics & Data Intelligence Solutions category in SiliconANGLE Media's 2026 TechForward Awards. The recognition highlights its open source, vendor-neutral platform for enterprise search, observability, security analytics, vector databases, and agentic AI workloads.

### Source excerpt

Recognition validates open source momentum, architectural consolidation, and enterprise scale as the project marks five years of community growth The post OpenSearch Wins Analytics & Data Intelligence Solutions Category in the SiliconANGLE TechForward Awards appeared first on OpenSearch.

## Using Exact-Match Response Caching to Reduce LLM Costs

DevFeed: [Using Exact-Match Response Caching to Reduce LLM Costs](<https://devfeed.tech/articles/why-an-old-caching-trick-is-your-secret-to-lower-llm-costs-17399.md>)

Original publisher: [Read original article](<https://thenewstack.io/llm-response-caching-costs/>)

Author: Abhilash Rao Mesala

Published: 2026-09-14T11:00:00Z

Content type: article

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Caching](<https://devfeed.tech/topics/caching.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [caching](<https://devfeed.tech/tags/caching.md>), [contributed](<https://devfeed.tech/tags/contributed.md>), [cost](<https://devfeed.tech/tags/cost.md>), [finops](<https://devfeed.tech/tags/finops.md>), [generation](<https://devfeed.tech/tags/generation.md>), [hash](<https://devfeed.tech/tags/hash.md>), [llm](<https://devfeed.tech/tags/llm.md>), [token](<https://devfeed.tech/tags/token.md>)

### AI overview

The article explains how to reduce LLM costs by fingerprinting requests, context, model settings, and underlying data to create exact-match cache keys. Valid cached responses can be reused without calling the model. It distinguishes response caching from provider prompt caching, where only eligible prompt computation is reused.

### Source excerpt

An LLM can answer the same question a thousand times and charge you each time. Before paying for another answer, The post Why an old caching trick is your secret to lower LLM costs appeared first on The New Stack.

## Temporal raises $550M at a $12.55B valuation as demand grows for reliable AI infrastructure

DevFeed: [Temporal raises $550M at a $12.55B valuation as demand grows for reliable AI infrastructure](<https://devfeed.tech/articles/temporal-raises-550m-at-a-12-55b-valuation-as-demand-grows-for-reliable-ai-infrastructure-36026.md>)

Original publisher: [Read original article](<https://temporal.io/blog/temporal-raises-usd550m-series-e-at-usd12-55b-valuation-ai>)

Author: Allanah Hughes

Published: 2026-09-14T00:00:00Z

Content type: release

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [reliability](<https://devfeed.tech/topics/reliability.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [funding](<https://devfeed.tech/tags/funding.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [outage](<https://devfeed.tech/tags/outage.md>), [reliability](<https://devfeed.tech/tags/reliability.md>), [series](<https://devfeed.tech/tags/series.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

Temporal announces a $550 million Series E funding round at a $12.55 billion valuation. The company says the funding will support reliable infrastructure for long-running AI agents and applications, including orchestration and recovery across systems.

### Source excerpt

AI is raising the bar for reliability. See why Temporal's $550M Series E, backed by Lightspeed and others, is built to meet that demand.

## Chip Huyen explains how to cut inference costs without new hardware

DevFeed: [Chip Huyen explains how to cut inference costs without new hardware](<https://devfeed.tech/articles/chip-huyen-explains-how-to-cut-inference-costs-without-new-hardware-10830.md>)

Original publisher: [Read original article](<https://thenewstack.io/pg-99-conf-2026-inference-costs/>)

Author: Tim Koopmans

Published: 2026-09-13T15:00:00Z

Content type: article

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Frontier Model](<https://devfeed.tech/topics/frontier-model.md>), [AI Engineering](<https://devfeed.tech/topics/ai-engineering.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [math](<https://devfeed.tech/topics/math.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [frontier-model](<https://devfeed.tech/tags/frontier-model.md>), [inference](<https://devfeed.tech/tags/inference.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [post-contributed](<https://devfeed.tech/tags/post-contributed.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [scylladb](<https://devfeed.tech/tags/scylladb.md>), [sponsor-scylladb](<https://devfeed.tech/tags/sponsor-scylladb.md>), [sponsored](<https://devfeed.tech/tags/sponsored.md>), [sponsored-post-contributed](<https://devfeed.tech/tags/sponsored-post-contributed.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

Chip Huyen explains why inference costs can outweigh one-time frontier-model training costs and outlines ways to optimize inference without new hardware. The article emphasizes latency metrics such as time to first token, time per output token, end-to-end latency, and goodput, especially for reasoning models.

### Source excerpt

Last October, the P99 conference -- the online gathering for developers focused on high-performance, low-latency applications -- featured a cracking The post Chip Huyen explains how to cut inference costs without new hardware appeared first on The New Stack.

## "Machine translation is still broken for most of the world's languages": Cohere builds non-reasoning for a reason

DevFeed: ["Machine translation is still broken for most of the world's languages": Cohere builds non-reasoning for a reason](<https://devfeed.tech/articles/machine-translation-is-still-broken-for-most-of-the-world-s-languages-cohere-builds-non-reasoning-for-a-reason-10829.md>)

Original publisher: [Read original article](<https://thenewstack.io/cohere-north-translate-sovereignty/>)

Author: Adrian Bridgwater

Published: 2026-09-13T14:21:46Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [cohere](<https://devfeed.tech/topics/cohere.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [gemma4](<https://devfeed.tech/topics/gemma4.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [aya](<https://devfeed.tech/tags/aya.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [cohere](<https://devfeed.tech/tags/cohere.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google](<https://devfeed.tech/tags/google.md>), [inference](<https://devfeed.tech/tags/inference.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [model](<https://devfeed.tech/tags/model.md>), [open](<https://devfeed.tech/tags/open.md>), [qwen](<https://devfeed.tech/tags/qwen.md>)

### AI overview

Cohere's North Small Translate is an open-weight mixture-of-experts machine translation model covering 50 languages. The article discusses its non-reasoning design, sovereign AI positioning, deployment options, efficiency claims, and reported WMT26 benchmark comparisons.

### Source excerpt

Enterprise AI company Cohere announced North Small Translate last week, a mixture-of-experts (MOE) open-weight machine translation model that works across The post "Machine translation is still broken for most of the world's languages": Cohere builds non-reasoning for a reason appeared first on The New Stack.

## HeyGen x Google Cloud: Bringing Avatar IV to TPUs

DevFeed: [HeyGen x Google Cloud: Bringing Avatar IV to TPUs](<https://devfeed.tech/articles/heygen-x-google-cloud-bringing-avatar-iv-to-tpus-4211.md>)

Original publisher: [Read original article](<https://developers.googleblog.com/heygen-x-google-cloud-bringing-avatar-iv-to-tpus/>)

Published: 2026-09-12T11:04:33.891311Z

Content type: article

Language: en

Sources: [Google Developers Blog](<https://devfeed.tech/sources/google-developers-blog.md>)

Topics: [Google](<https://devfeed.tech/topics/google.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>), [Compiler](<https://devfeed.tech/topics/compiler.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [api](<https://devfeed.tech/tags/api.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [code](<https://devfeed.tech/tags/code.md>), [compiler](<https://devfeed.tech/tags/compiler.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [generation](<https://devfeed.tech/tags/generation.md>), [google](<https://devfeed.tech/tags/google.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [model](<https://devfeed.tech/tags/model.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [time](<https://devfeed.tech/tags/time.md>), [tpu](<https://devfeed.tech/tags/tpu.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [video](<https://devfeed.tech/tags/video.md>)

### AI overview

HeyGen and Google Cloud describe porting the 18B+ parameter Avatar IV talking-head video generation pipeline to an eight-chip Trillium TPU host. Using torchax, JAX, XLA, FSDP sharding, Ulysses sequence parallelism, and custom Pallas kernels, the team improved performance by 1.86x for real-time chunked streaming while preserving output quality through strict quality gates.

### Source excerpt

HeyGen ported their 18B+ parameter Avatar IV video generation model to Google Cloud's Trillium (v6e) TPUs via torchax and XLA, utilizing FSDP and Ulysses sequence parallelism across an eight-chip mesh. To achieve a 1.86x speedup for real-time streaming, the engineering team pipelined exposed all-to-all collectives, aligned sparse attention block sizes to eliminate mask padding, and bypassed softmax serial dependencies using a precomputed Cauchy-Schwarz upper bound. These custom Pallas kernel and compiler optimizations were deployed only after passing rigorous two-tier quality gates to guarantee byte-identical or mathematically equivalent pixel outputs.

## OpenAI's researchers burned $7,000 a day on AI agents -- now it's opening the floodgates

DevFeed: [OpenAI's researchers burned $7,000 a day on AI agents -- now it's opening the floodgates](<https://devfeed.tech/articles/openai-s-researchers-burned-7-000-a-day-on-ai-agents-now-it-s-opening-the-floodgates-8483.md>)

Original publisher: [Read original article](<https://thenewstack.io/openai-agents-api-compute/>)

Author: Amanda Caswell

Published: 2026-09-11T21:27:42Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [long-context](<https://devfeed.tech/topics/long-context.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [api](<https://devfeed.tech/tags/api.md>), [cloud-services](<https://devfeed.tech/tags/cloud-services.md>), [codex](<https://devfeed.tech/tags/codex.md>), [compute](<https://devfeed.tech/tags/compute.md>), [developers](<https://devfeed.tech/tags/developers.md>), [inference](<https://devfeed.tech/tags/inference.md>), [openai](<https://devfeed.tech/tags/openai.md>), [research](<https://devfeed.tech/tags/research.md>), [sandbox](<https://devfeed.tech/tags/sandbox.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

OpenAI's public-beta Agents API lets developers run long-lived agents with managed job state, context compression, optional tools, parallel subagents, and execution in OpenAI's sandbox or developer-controlled infrastructure. The article highlights the resulting inference and compute costs, citing internal research-agent usage figures.

### Source excerpt

OpenAI rolled out its Agents API in public beta Thursday, opening the backend behind Codex to developers looking to run The post OpenAI's researchers burned $7,000 a day on AI agents -- now it's opening the floodgates appeared first on The New Stack.

## Cohere's new translation model is open weights -- but not for commercial use

DevFeed: [Cohere's new translation model is open weights -- but not for commercial use](<https://devfeed.tech/articles/cohere-s-new-translation-model-is-open-weights-but-not-for-commercial-use-8474.md>)

Original publisher: [Read original article](<https://thenewstack.io/cohere-translation-commercial-licensing/>)

Author: Meredith Shubel

Published: 2026-09-11T17:50:11Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [moe](<https://devfeed.tech/topics/moe.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [ai-strategy](<https://devfeed.tech/tags/ai-strategy.md>), [api](<https://devfeed.tech/tags/api.md>), [cohere](<https://devfeed.tech/tags/cohere.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [production](<https://devfeed.tech/tags/production.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

Cohere released North Small Translate 1.0 as open weights under CC BY-NC 4.0, allowing download, evaluation, and study but requiring a commercial agreement for production use. Commercial deployment requires a license and use of Cohere's managed Model Vault platform.

### Source excerpt

This week, Cohere released North Small Translate 1.0 under a CC BY-NC 4.0 license: the weights are there to download, The post Cohere's new translation model is open weights -- but not for commercial use appeared first on The New Stack.

## Kubernetes v1.37 brings 67 enhancements. Which matter for operators?

DevFeed: [Kubernetes v1.37 brings 67 enhancements. Which matter for operators?](<https://devfeed.tech/articles/kubernetes-v1-37-brings-67-enhancements-which-matter-for-operators-8480.md>)

Original publisher: [Read original article](<https://thenewstack.io/kubecon-kubernetes-updates-security/>)

Author: Bill Doerrfeld

Published: 2026-09-11T17:40:04Z

Content type: news

Language: en

Sources: [Kubernetes Overview, News and Trends | The New Stack](<https://devfeed.tech/sources/kubernetes-overview-news-and-trends-the-new-stack.md>), [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Infrastructure as code](<https://devfeed.tech/topics/infrastructure-as-code.md>), [migration](<https://devfeed.tech/topics/migration.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Authorization](<https://devfeed.tech/topics/authorization.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [cloud-native-ecosystem](<https://devfeed.tech/tags/cloud-native-ecosystem.md>), [hpe](<https://devfeed.tech/tags/hpe.md>), [infrastructure-as-code](<https://devfeed.tech/tags/infrastructure-as-code.md>), [kubecon-cloudnativecon-na-2026](<https://devfeed.tech/tags/kubecon-cloudnativecon-na-2026.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [migration](<https://devfeed.tech/tags/migration.md>), [post](<https://devfeed.tech/tags/post.md>), [road-to-kubecon](<https://devfeed.tech/tags/road-to-kubecon.md>), [sponsor-hpe](<https://devfeed.tech/tags/sponsor-hpe.md>), [sponsored-post](<https://devfeed.tech/tags/sponsored-post.md>), [terraform](<https://devfeed.tech/tags/terraform.md>)

### AI overview

A Kubernetes roundup covers v1.37 enhancements, HPE Morpheus and Terraform provider changes, migration tooling, and recent CNCF project graduations including Kubeflow.

### Source excerpt

Welcome to the first edition of Road to KubeCon, where we'll track the world of Kubernetes as we approach KubeCon The post Kubernetes v1.37 brings 67 enhancements. Which matter for operators? appeared first on The New Stack.

## Palantir and NVIDIA Deploy a Sovereign Nemotron Supply Chain Stack, Starting With the 1.3 Million Parts in Every Vera Rubin Rack

DevFeed: [Palantir and NVIDIA Deploy a Sovereign Nemotron Supply Chain Stack, Starting With the 1.3 Million Parts in Every Vera Rubin Rack](<https://devfeed.tech/articles/palantir-and-nvidia-deploy-a-sovereign-nemotron-supply-chain-stack-starting-with-the-1-3-million-parts-in-every-vera-rubin-rack-12372.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/palantir-and-nvidia-deploy-a-sovereign-nemotron-supply-chain-stack-starting-with-the-1-3-million-parts-in-every-vera-rubin-rack>)

Author: Harold Fritts

Published: 2026-09-10T20:56:11Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Vera Rubin](<https://devfeed.tech/topics/vera-rubin.md>), [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [cuOpt](<https://devfeed.tech/topics/cuopt.md>), [Complex Systems](<https://devfeed.tech/topics/complex-systems.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [complex-systems](<https://devfeed.tech/tags/complex-systems.md>), [cuopt](<https://devfeed.tech/tags/cuopt.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [models](<https://devfeed.tech/tags/models.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open](<https://devfeed.tech/tags/open.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [supply-chain](<https://devfeed.tech/tags/supply-chain.md>), [systems](<https://devfeed.tech/tags/systems.md>), [vera-rubin](<https://devfeed.tech/tags/vera-rubin.md>)

### AI overview

Palantir and NVIDIA have deployed a sovereign AI stack for supply chain operations, initially using NVIDIA's own Vera Rubin supply chain as the first customer. The system combines Nemotron open models with Palantir Foundry and AIP, NVIDIA NeMo Data Libraries, and cuOpt to support materials allocation, scenario planning, optimization, and risk detection while keeping final decisions with supply chain experts.

### Source excerpt

Palantir and NVIDIA have built a sovereign AI stack for supply chain operations and are running it first inside NVIDIA's own supply chain, the one that has to line up 1.3 million parts for every Vera Rubin rack. The stack brings NVIDIA Nemotron open models into Palantir Foundry and its Artificial Intelligence Platform (AIP), grounded The post Palantir and NVIDIA Deploy a Sovereign Nemotron Supply Chain Stack, Starting With the 1.3 Million Parts in Every Vera Rubin Rack appeared first on StorageReview.com.

## Backblaze B2 and WEKA NeuralMesh Validated as a Two-Tier AI Storage Pipeline, With Snap-to-Object Checkpoints in B2

DevFeed: [Backblaze B2 and WEKA NeuralMesh Validated as a Two-Tier AI Storage Pipeline, With Snap-to-Object Checkpoints in B2](<https://devfeed.tech/articles/backblaze-b2-and-weka-neuralmesh-validated-as-a-two-tier-ai-storage-pipeline-with-snap-to-object-checkpoints-in-b2-12360.md>)

Original publisher: [Read original article](<https://www.storagereview.com/news/backblaze-b2-and-weka-neuralmesh-validated-as-a-two-tier-ai-storage-pipeline-with-snap-to-object-checkpoints-landing-in-b2>)

Author: Harold Fritts

Published: 2026-09-10T20:40:57Z

Content type: news

Language: en

Sources: [StorageReview.com](<https://devfeed.tech/sources/storagereview-com.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [accelerators](<https://devfeed.tech/tags/accelerators.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [integration](<https://devfeed.tech/tags/integration.md>), [performance](<https://devfeed.tech/tags/performance.md>), [snapshots](<https://devfeed.tech/tags/snapshots.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

Backblaze and WEKA validated a two-tier AI storage pipeline that uses WEKA NeuralMesh as the high-performance tier for GPU workloads and Backblaze B2 Cloud Storage as the capacity tier. Raw data, training sets, media, source files, checkpoints, and other assets can move between the tiers according to access needs. NeuralMesh's Snap-to-Object feature was also tested with B2 for storing consistent filesystem snapshots and supporting recovery.

### Source excerpt

Backblaze and WEKA have validated their two platforms together for AI pipelines, pairing WEKA NeuralMesh as the performance tier that feeds GPUs with Backblaze B2 Cloud Storage as the capacity tier that holds everything else. The integration, sizing, tuning, and testing are already done, so an AI infrastructure team can deploy a proven two-tier layout The post Backblaze B2 and WEKA NeuralMesh Validated as a Two-Tier AI Storage Pipeline, With Snap-to-Object Checkpoints in B2 appeared first on StorageReview.com.

## Powering the AI era: How wave energy can complement a 24/7 energy mix

DevFeed: [Powering the AI era: How wave energy can complement a 24/7 energy mix](<https://devfeed.tech/articles/powering-the-ai-era-how-wave-energy-can-complement-a-24-7-energy-mix-10939.md>)

Original publisher: [Read original article](<https://blogs.cisco.com/our-corporate-purpose/powering-the-ai-era-how-wave-energy-can-complement-a-24-7-energy-mix>)

Author: Elias Habbar-Baylac

Published: 2026-09-10T15:20:22Z

Content type: article

Language: en

Sources: [Cisco Blogs](<https://devfeed.tech/sources/cisco-blogs.md>)

Topics: [AI Infrastructure](<https://devfeed.tech/topics/ai-infrastructure.md>), [data centers](<https://devfeed.tech/topics/data-centers.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [chief-sustainability-office](<https://devfeed.tech/tags/chief-sustainability-office.md>), [cisco-purpose](<https://devfeed.tech/tags/cisco-purpose.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [energy](<https://devfeed.tech/tags/energy.md>), [environmental-sustainability](<https://devfeed.tech/tags/environmental-sustainability.md>), [generation](<https://devfeed.tech/tags/generation.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [megawatt](<https://devfeed.tech/tags/megawatt.md>), [our-corporate-purpose](<https://devfeed.tech/tags/our-corporate-purpose.md>), [sustainability](<https://devfeed.tech/tags/sustainability.md>)

### AI overview

The article explains how wave energy could complement solar, wind, and batteries in meeting the continuous electricity needs of AI-era data centers. It discusses a modeled 100 MW flat-load data center and energy portfolios evaluated for cost, reliability, emissions, and round-the-clock availability.

### Source excerpt

CorPower Ocean, a Cisco Investments portfolio company, has explored how wave energy technology could complement other sources for 24/7 energy needs.

[Next page](<https://devfeed.tech/tags/ai-infrastructure.md?cursor=WyIyMDI2LTA5LTEwVDE1OjIwOjIyKzAwOjAwIiwgIjUyNTBiYmY4LTllZTctNDU2OS05ZmNiLTRhNGYzZjY5ZWFiZSJd>)