# kueue

Published articles for kueue.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Microsoft Open-Sources TauGrid to Simplify AI Workload Management on Kubernetes

DevFeed: [Microsoft Open-Sources TauGrid to Simplify AI Workload Management on Kubernetes](<https://devfeed.tech/articles/microsoft-open-sources-taugrid-to-simplify-ai-workload-management-on-kubernetes-31518.md>)

Original publisher: [Read original article](<https://www.infoq.com/news/2026/09/microsoft-taugrid-open-source/>)

Author: Sergio De Simone

Published: 2026-09-16T18:00:00Z

Content type: news

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [kueue](<https://devfeed.tech/topics/kueue.md>), [Go](<https://devfeed.tech/topics/go.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [azure](<https://devfeed.tech/tags/azure.md>), [development](<https://devfeed.tech/tags/development.md>), [devops](<https://devfeed.tech/tags/devops.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kueue](<https://devfeed.tech/tags/kueue.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [microsoft-taugrid-open-source](<https://devfeed.tech/tags/microsoft-taugrid-open-source.md>), [news](<https://devfeed.tech/tags/news.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>)

### AI overview

Microsoft has open-sourced TauGrid, a cloud-native platform for managing, scheduling, and monitoring AI workloads on GPU-enabled Kubernetes clusters. It combines workload submission, Kueue-based queuing, KubeRay orchestration, GPU-node monitoring, and observability, while planned capabilities remain on its roadmap.

### Source excerpt

Microsoft has open-sourced TauGrid, a cloud-native platform designed to manage, schedule, and monitor AI workloads on GPU-enabled Kubernetes clusters. By Sergio De Simone

## Monitor TAS and gang scheduling for AI training in Kubernetes

DevFeed: [Monitor TAS and gang scheduling for AI training in Kubernetes](<https://devfeed.tech/articles/monitor-tas-and-gang-scheduling-for-ai-training-in-kubernetes-26969.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/monitor-tas-and-gang-scheduling-for-ai-training-in-kubernetes/>)

Author: David Lentz; Kathy Lin

Published: 2026-09-15T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [kueue](<https://devfeed.tech/topics/kueue.md>), [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [datadog](<https://devfeed.tech/topics/datadog.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [batch](<https://devfeed.tech/tags/batch.md>), [containers](<https://devfeed.tech/tags/containers.md>), [datadog](<https://devfeed.tech/tags/datadog.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-monitoring](<https://devfeed.tech/tags/gpu-monitoring.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kueue](<https://devfeed.tech/tags/kueue.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [scheduler](<https://devfeed.tech/tags/scheduler.md>)

### AI overview

This article explains why Kubernetes scheduling is insufficient for distributed AI training workloads and how topology-aware scheduling and gang scheduling address hardware placement and simultaneous startup requirements. It discusses implementing these capabilities with Kueue and the Coscheduling plugin, and monitoring and troubleshooting them with Datadog GPU Monitoring.

### Source excerpt

Learn how Datadog helps you correlate Kueue, Coscheduling, GPU, and training framework signals to validate gang scheduling and topology-aware scheduling.

## China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes

DevFeed: [China Merchants Bank Wins CNCF End User Case Study Contest for Unifying AI Training and Inference on Kubernetes](<https://devfeed.tech/articles/china-merchants-bank-wins-cncf-end-user-case-study-contest-for-unifying-ai-training-and-inference-on-kubernetes-4594.md>)

Original publisher: [Read original article](<https://www.cncf.io/announcements/2026/09/07/china-merchants-bank-wins-cncf-end-user-case-study-contest-for-unifying-ai-training-and-inference-on-kubernetes/>)

Author: Haley White

Published: 2026-09-08T01:54:31Z

Content type: news

Language: en

Sources: [Cloud Native Computing Foundation](<https://devfeed.tech/sources/cloud-native-computing-foundation.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Platform](<https://devfeed.tech/topics/ai-platform.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [kueue](<https://devfeed.tech/topics/kueue.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Cloud Native Ecosystem](<https://devfeed.tech/topics/cloud-native-ecosystem.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-infrastructure](<https://devfeed.tech/tags/ai-infrastructure.md>), [ai-training](<https://devfeed.tech/tags/ai-training.md>), [announcements](<https://devfeed.tech/tags/announcements.md>), [case-study](<https://devfeed.tech/tags/case-study.md>), [china](<https://devfeed.tech/tags/china.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cost](<https://devfeed.tech/tags/cost.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference](<https://devfeed.tech/tags/inference.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kueue](<https://devfeed.tech/tags/kueue.md>), [lora](<https://devfeed.tech/tags/lora.md>)

### AI overview

China Merchants Bank won a CNCF case-study contest for a Kubernetes-based AI platform that shares nearly 10,000 accelerator cards across training, fine-tuning, and online inference. The bank reports increased average accelerator utilization and lower inference costs.

### Source excerpt

New cloud native platform lifted average accelerator compute utilization from 35% to more than 60% and cut inference cost per 1 million tokens by more than 60% Key Highlights SHANGHAI, China - KubeCon + CloudNativeCon +...

## How Netflix Simplified Batch Compute with Kueue

DevFeed: [How Netflix Simplified Batch Compute with Kueue](<https://devfeed.tech/articles/how-netflix-simplified-batch-compute-with-kueue-139.md>)

Original publisher: [Read original article](<https://netflixtechblog.com/how-netflix-simplified-batch-compute-with-kueue-87860682629c?source=rss----2615bd06b42e---4>)

Author: Netflix Technology Blog

Published: 2026-06-22T21:35:01Z

Content type: article

Language: en

Sources: [Netflix](<https://devfeed.tech/sources/netflix.md>), [Netflix TechBlog - Medium](<https://devfeed.tech/sources/netflix-techblog-medium.md>)

Topics: [kueue](<https://devfeed.tech/topics/kueue.md>), [Netflix](<https://devfeed.tech/topics/netflix.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [migration](<https://devfeed.tech/topics/migration.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [App](<https://devfeed.tech/topics/app.md>)

Tags: [applications](<https://devfeed.tech/tags/applications.md>), [aws](<https://devfeed.tech/tags/aws.md>), [batch](<https://devfeed.tech/tags/batch.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [compute](<https://devfeed.tech/tags/compute.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kueue](<https://devfeed.tech/tags/kueue.md>), [migration](<https://devfeed.tech/tags/migration.md>), [netflix](<https://devfeed.tech/tags/netflix.md>), [platform](<https://devfeed.tech/tags/platform.md>), [resource](<https://devfeed.tech/tags/resource.md>), [work](<https://devfeed.tech/tags/work.md>)

### AI overview

Netflix describes its transition toward a more Kubernetes-native compute infrastructure and its adoption of Kueue, a cloud-native job queueing system for batch workloads. The article explains how Kueue replaced custom queuing and scheduling logic in Compute Managed Batch, supported the migration of millions of batch jobs, and enabled tenant-based capacity management and workload execution through Titus.

### Source excerpt

By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha Chandrashekar As a part of the journey to transition Netflix's compute infrastructure to be more Kubernetes-native, we have leaned into incorporating components from the Kubernetes ecosystem into our container platform Titus. One example of this is our use of Kueue, a cloud-native job queueing system for batch workloads, which has largely replaced the custom queuing and scheduling logic in our homegrown managed batch solution Compute Managed Batch (CMB). In this post, we'll give an overview of what motivated the migration, how we migrated millions of batch jobs to use Kueue, and what Kueue allows us to offer as a Compute platform. Brief Overview of CMB and Titus CMB is a managed batch solution that allows users and applications to execute and manage workloads that run to completion. Using a tenant hierarchy, workloads are managed and queued with ordered execution through priorities, and capacity is managed on a per-tenant basis. Workloads that are submitted to CMB are then run on Titus. The features of Titus relevant to CMB are workload federation across multiple cells (Kubernetes clusters) and federated capacity reservations. This means CMB can talk to a single Titus endpoint to get/submit workloads and update capacity reservations without having to worry about the underlying cell/cluster topology. CMB Tenant Hierarchy Tenants provide a grouping mechanism for jobs submitted on behalf of certain organizations, platforms, or applications. Users can create and organize tenants however best suits their organization or use case. For example, an organization may use a single tenant across several applications or a complex hierarchical structure that matches its team and application ownership structure. Tenants are associated with a capacity configuration. The capacity configuration defines the amount of compute capacity available to the tenant and provides certain guarantees around isolation from other te

## Kubernetes v1.36: Mutable Pod Resources for Suspended Jobs (beta)

DevFeed: [Kubernetes v1.36: Mutable Pod Resources for Suspended Jobs (beta)](<https://devfeed.tech/articles/kubernetes-v1-36-mutable-pod-resources-for-suspended-jobs-beta-4540.md>)

Original publisher: [Read original article](<https://kubernetes.io/blog/2026/04/27/kubernetes-v1-36-mutable-pod-resources-for-suspended-jobs/>)

Author: Kevin Hannon

Published: 2026-04-27T18:35:00Z

Content type: article

Language: en

Sources: [Kubernetes Blog](<https://devfeed.tech/sources/kubernetes-blog.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [jobs](<https://devfeed.tech/topics/jobs.md>), [kueue](<https://devfeed.tech/topics/kueue.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [batch](<https://devfeed.tech/tags/batch.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [kueue](<https://devfeed.tech/tags/kueue.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>)

### AI overview

Kubernetes v1.36 promotes mutable CPU, memory, GPU, and extended-resource requests and limits for suspended Jobs to beta. Queue controllers such as Kueue can adjust resources before a Job resumes, avoiding deletion and recreation when cluster capacity or workload requirements change.

### Source excerpt

Kubernetes v1.36 promotes the ability to modify container resource requests and limits in the pod template of a suspended Job to beta. First introduced as alpha in v1.35, this feature allows queue controllers and cluster administrators to adjust CPU, memory, GPU, and extended resource specifications on a Job while it is suspended, before it starts or resumes running. Why mutable pod resources for suspended Jobs? Batch and machine learning workloads often have resource requirements that are not precisely known at Job creation time. The optimal resource allocation depends on current cluster capacity, queue priorities, and the availability of specialized hardware like GPUs. Before this feature, resource requirements in a Job's pod template were immutable once set. If a queue controller like Kueue determined that a suspended Job should run with different resources, the only option was to delete and recreate the Job, losing any associated metadata, status, or history. This feature also provides a way to let a specific Job instance for a CronJob progress slowly with reduced resources, rather than outright failing to run if the cluster is heavily loaded. Consider a machine learning training Job initially requesting 4 GPUs: apiVersion: batch/v1 kind: Job metadata: name: training-job-example-abcd123 labels: app.kubernetes.io/name: trainer spec: suspend: true template: metadata: annotations: kubernetes.io/description: "ML training, ID abcd123" spec: containers: - name: trainer image: example-registry.example.com/training:2026-04-23T150405.678 resources: requests: cpu: "8" memory: "32Gi" example-hardware-vendor.com/gpu: "4" limits: cpu: "8" memory: "32Gi" example-hardware-vendor.com/gpu: "4" restartPolicy: Never A queue controller managing cluster resources might determine that only 2 GPUs are available. With this feature, the controller can update the Job's resource requests before resuming it: apiVersion: batch/v1 kind: Job metadata: name: training-job-example-abcd123 labels