# AWS Parallel Computing Service

Published articles for AWS Parallel Computing Service.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Part 1: Managing Large-Scale LLM Training with AWS ParallelCluster

DevFeed: [Part 1: Managing Large-Scale LLM Training with AWS ParallelCluster](<https://devfeed.tech/articles/part-1-managing-large-scale-llm-training-with-aws-parallelcluster-50154.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/hpc/part-1-managing-large-scale-llm-training-with-aws-parallelcluster/>)

Author: Seokjae Jang

Published: 2026-08-25T22:33:54Z

Content type: article

Language: en

Sources: [AWS HPC Blog](<https://devfeed.tech/sources/aws-hpc-blog.md>)

Topics: [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Computing](<https://devfeed.tech/topics/computing.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [scaling](<https://devfeed.tech/topics/scaling.md>), [YAML](<https://devfeed.tech/topics/yaml.md>)

Tags: [ai-research](<https://devfeed.tech/tags/ai-research.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-parallel-computing-service](<https://devfeed.tech/tags/aws-parallel-computing-service.md>), [aws-parallelcluster](<https://devfeed.tech/tags/aws-parallelcluster.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [compute](<https://devfeed.tech/tags/compute.md>), [customer-solutions](<https://devfeed.tech/tags/customer-solutions.md>), [elastic-fabric-adapter](<https://devfeed.tech/tags/elastic-fabric-adapter.md>), [experience-based-acceleration](<https://devfeed.tech/tags/experience-based-acceleration.md>), [government](<https://devfeed.tech/tags/government.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-infrastructure](<https://devfeed.tech/tags/gpu-infrastructure.md>), [high-performance](<https://devfeed.tech/tags/high-performance.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llm-training](<https://devfeed.tech/tags/llm-training.md>), [slurm](<https://devfeed.tech/tags/slurm.md>), [teams](<https://devfeed.tech/tags/teams.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>)

### AI overview

This post shares operational lessons from building a large-scale GPU cluster for LLM training by using AWS ParallelCluster and FSx for Lustre. It covers resource stability, cluster configuration, storage and networking integration, Slurm scheduling, and practical strategies for managing long-running training workloads.

### Source excerpt

This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. Introduction The Korean Government announced a national AI initiative to provide high-performance GPU infrastructure for Korea's national AI research teams. AWS was selected as a supplier of GPU resources [...]

## Resilient HPC and ML on AWS: Running Tightly Coupled Workloads on Spot Instances

DevFeed: [Resilient HPC and ML on AWS: Running Tightly Coupled Workloads on Spot Instances](<https://devfeed.tech/articles/resilient-hpc-and-ml-on-aws-running-tightly-coupled-workloads-on-spot-instances-50156.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/hpc/resilient-hpc-and-ml-on-aws-running-tightly-coupled-workloads-on-spot-instances/>)

Author: Santosh Kumar

Published: 2026-08-05T00:00:55Z

Content type: tutorial

Language: en

Sources: [AWS HPC Blog](<https://devfeed.tech/sources/aws-hpc-blog.md>)

Topics: [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Computational Fluid Dynamics (CFD)](<https://devfeed.tech/topics/cfd.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [GROMACS](<https://devfeed.tech/topics/gromacs.md>), [Molecular Dynamics](<https://devfeed.tech/topics/molecular-dynamics.md>), [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>)

Tags: [amazon-ec2](<https://devfeed.tech/tags/amazon-ec2.md>), [amazon-elastic-file-system-efs](<https://devfeed.tech/tags/amazon-elastic-file-system-efs.md>), [amazon-fsx-for-lustre](<https://devfeed.tech/tags/amazon-fsx-for-lustre.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-fault-injection-service-fis](<https://devfeed.tech/tags/aws-fault-injection-service-fis.md>), [aws-parallel-computing-service](<https://devfeed.tech/tags/aws-parallel-computing-service.md>), [aws-parallelcluster](<https://devfeed.tech/tags/aws-parallelcluster.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [cfd](<https://devfeed.tech/tags/cfd.md>), [computational-fluid-dynamics](<https://devfeed.tech/tags/computational-fluid-dynamics.md>), [compute](<https://devfeed.tech/tags/compute.md>), [computing](<https://devfeed.tech/tags/computing.md>), [cost](<https://devfeed.tech/tags/cost.md>), [developer-tools](<https://devfeed.tech/tags/developer-tools.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [gromacs](<https://devfeed.tech/tags/gromacs.md>), [high-performance](<https://devfeed.tech/tags/high-performance.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [ml](<https://devfeed.tech/tags/ml.md>), [research-computing](<https://devfeed.tech/tags/research-computing.md>), [scientific-computing](<https://devfeed.tech/tags/scientific-computing.md>), [simulations](<https://devfeed.tech/tags/simulations.md>), [slurm](<https://devfeed.tech/tags/slurm.md>), [storage](<https://devfeed.tech/tags/storage.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>)

### AI overview

This AWS post presents an architecture for running tightly coupled, long-running HPC and machine learning workloads on Amazon EC2 Spot Instances. It addresses interruption risk through checkpointing and demonstrates the approach with CFD and GROMACS workloads.

### Source excerpt

This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. This post was contributed by Santosh Kumar, Bhagyaraju Kasina, Dr. Sandeep Sovani and Dr. Max Starr Researchers and engineering teams running High Performance Computing (HPC) jobs face a constant [...]

## Monitoring AWS Parallel Computing Service

DevFeed: [Monitoring AWS Parallel Computing Service](<https://devfeed.tech/articles/monitoring-aws-parallel-computing-service-50159.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/hpc/the-complete-picture-unified-monitoring-for-aws-parallel-computing-service/>)

Author: Ronald Hudson

Published: 2026-06-02T16:29:06Z

Content type: tutorial

Language: en

Sources: [AWS HPC Blog](<https://devfeed.tech/sources/aws-hpc-blog.md>)

Topics: [Computing](<https://devfeed.tech/topics/computing.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [High-Performance Computing](<https://devfeed.tech/topics/high-performance-computing.md>), [Grafana](<https://devfeed.tech/topics/grafana.md>), [Grafana Cloud Metrics](<https://devfeed.tech/topics/grafana-cloud-metrics.md>), [Amazon Managed Service for Prometheus](<https://devfeed.tech/topics/amazon-managed-service-for-prometheus.md>), [parallel](<https://devfeed.tech/topics/parallel.md>), [Amazon CloudWatch Logs](<https://devfeed.tech/topics/amazon-cloudwatch-logs.md>), [OpenTelemetry](<https://devfeed.tech/topics/opentelemetry.md>)

Tags: [amazon-cloudwatch-logs](<https://devfeed.tech/tags/amazon-cloudwatch-logs.md>), [amazon-managed-grafana](<https://devfeed.tech/tags/amazon-managed-grafana.md>), [amazon-managed-service-for-prometheus](<https://devfeed.tech/tags/amazon-managed-service-for-prometheus.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-distro-for-opentelemetry](<https://devfeed.tech/tags/aws-distro-for-opentelemetry.md>), [aws-parallel-computing-service](<https://devfeed.tech/tags/aws-parallel-computing-service.md>), [cluster](<https://devfeed.tech/tags/cluster.md>), [collector](<https://devfeed.tech/tags/collector.md>), [compute](<https://devfeed.tech/tags/compute.md>), [computing](<https://devfeed.tech/tags/computing.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [elastic-fabric-adapter](<https://devfeed.tech/tags/elastic-fabric-adapter.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [grafana](<https://devfeed.tech/tags/grafana.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

This AWS post presents a unified observability solution for AWS Parallel Computing Service environments. It combines Amazon Managed Grafana, Amazon Managed Service for Prometheus, CloudWatch Logs, exporters, and AWS Distro for OpenTelemetry to provide dashboards for cluster resources, job performance, and diagnostic data.

### Source excerpt

This post was contributed by Ronald Hudson and Nate Haynes High Performance Computing (HPC) on AWS demands precise monitoring, like the racing telemetry used by Formula 1 teams to deliver results. Like race engineers tracking car performance, AWS Parallel Computing Service (AWS PCS) administrators must monitor computing metrics in real-time. This vigilance is critical because [...]

## Accelerating HPC Deployment with AWS Parallel Computing Service and Kiro CLI

DevFeed: [Accelerating HPC Deployment with AWS Parallel Computing Service and Kiro CLI](<https://devfeed.tech/articles/accelerating-hpc-deployment-with-aws-parallel-computing-service-and-kiro-cli-50143.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/hpc/accelerating-hpc-deployment-with-aws-parallel-computing-service-and-kiro-cli/>)

Author: Markus Adhiwiyogo

Published: 2026-03-26T18:06:41Z

Content type: tutorial

Language: en

Sources: [AWS HPC Blog](<https://devfeed.tech/sources/aws-hpc-blog.md>)

Topics: [Computing](<https://devfeed.tech/topics/computing.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Kiro](<https://devfeed.tech/topics/kiro.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Provisioning](<https://devfeed.tech/topics/provisioning.md>), [configuration-management](<https://devfeed.tech/topics/configuration-management.md>), [networking](<https://devfeed.tech/topics/networking.md>), [scaling](<https://devfeed.tech/topics/scaling.md>), [scheduling](<https://devfeed.tech/topics/scheduling.md>), [cost-optimization](<https://devfeed.tech/topics/cost-optimization.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>)

Tags: [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-parallel-computing-service](<https://devfeed.tech/tags/aws-parallel-computing-service.md>), [aws-parallelcluster](<https://devfeed.tech/tags/aws-parallelcluster.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cluster](<https://devfeed.tech/tags/cluster.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [compute](<https://devfeed.tech/tags/compute.md>), [compute-resources](<https://devfeed.tech/tags/compute-resources.md>), [configuration-management](<https://devfeed.tech/tags/configuration-management.md>), [cost-optimization](<https://devfeed.tech/tags/cost-optimization.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [integration](<https://devfeed.tech/tags/integration.md>), [intermediate-200](<https://devfeed.tech/tags/intermediate-200.md>), [kiro](<https://devfeed.tech/tags/kiro.md>), [networking](<https://devfeed.tech/tags/networking.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [slurm](<https://devfeed.tech/tags/slurm.md>), [teams](<https://devfeed.tech/tags/teams.md>), [technical-how-to](<https://devfeed.tech/tags/technical-how-to.md>)

### AI overview

This tutorial explains how Kiro CLI custom agents can automate the deployment and configuration of AWS Parallel Computing Service clusters for cloud-based HPC workloads. It covers infrastructure provisioning, networking, storage, Slurm scheduling, monitoring, and cost optimization.

### Source excerpt

This post references AWS ParallelCluster. Check out AWS Parallel Computing Service (AWS PCS), our new managed Slurm service for running HPC and AI workloads on AWS. Research teams moving from on-premises HPC environments often struggle with the complexity of cloud deployment. Traditional approaches require deep expertise in AWS networking, storage architectures, and Slurm configuration management. [...]

## How Aionics accelerates chemical formulation and discovery with AWS Parallel Computing Service

DevFeed: [How Aionics accelerates chemical formulation and discovery with AWS Parallel Computing Service](<https://devfeed.tech/articles/how-aionics-accelerates-chemical-formulation-and-discovery-with-aws-parallel-computing-service-50149.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/hpc/how-aionics-accelerates-chemical-formulation-and-discovery-with-aws-parallel-computing-service/>)

Author: Moe Elshazly

Published: 2025-11-26T15:35:53Z

Content type: article

Language: en

Sources: [AWS HPC Blog](<https://devfeed.tech/sources/aws-hpc-blog.md>)

Topics: [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Panic Playdate SDK](<https://devfeed.tech/topics/playdate-sdk.md>), [Simulation](<https://devfeed.tech/topics/simulation.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Molecular Dynamics](<https://devfeed.tech/topics/molecular-dynamics.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [DSQL](<https://devfeed.tech/topics/dsql.md>)

Tags: [accelerated](<https://devfeed.tech/tags/accelerated.md>), [aerospace](<https://devfeed.tech/tags/aerospace.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-parallel-computing-service](<https://devfeed.tech/tags/aws-parallel-computing-service.md>), [batteries](<https://devfeed.tech/tags/batteries.md>), [computing](<https://devfeed.tech/tags/computing.md>), [elastic-fabric-adapter](<https://devfeed.tech/tags/elastic-fabric-adapter.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-accelerated](<https://devfeed.tech/tags/gpu-accelerated.md>), [high-performance](<https://devfeed.tech/tags/high-performance.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [machine](<https://devfeed.tech/tags/machine.md>), [models](<https://devfeed.tech/tags/models.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [research-computing](<https://devfeed.tech/tags/research-computing.md>), [scientific-computing](<https://devfeed.tech/tags/scientific-computing.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [simulations](<https://devfeed.tech/tags/simulations.md>), [slurm](<https://devfeed.tech/tags/slurm.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

This AWS post explains how Aionics uses AWS Parallel Computing Service to build flexible high-performance computing infrastructure for computational chemistry. It describes plane-wave DFT simulations with MPI parallelization and GPU-accelerated molecular dynamics using machine-learned interatomic potentials to support battery electrolyte discovery.

### Source excerpt

This post was contributed by Mohamed K. Elshazly, PhD, Kareem Abdol-Hamid, Sam Bydlon, PhD, Aarabhi Achanta, and Mark Azadpour The decarbonization of our modern economy depends on solving a defining scientific challenge: developing batteries that are both safe and high performing. From electrical grids to vehicles and aviation, these energy storage devices must provide power [...]

## How Daiichi Sankyo modernized drug discovery using AWS Parallel Computing Service

DevFeed: [How Daiichi Sankyo modernized drug discovery using AWS Parallel Computing Service](<https://devfeed.tech/articles/how-daiichi-sankyo-modernized-drug-discovery-using-aws-parallel-computing-service-50150.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/hpc/how-daiichi-sankyo-modernized-drug-discovery-using-aws-parallel-computing-service/>)

Author: Ryo Kunimoto

Published: 2025-11-19T14:34:36Z

Content type: article

Language: en

Sources: [AWS HPC Blog](<https://devfeed.tech/sources/aws-hpc-blog.md>)

Topics: [Computing](<https://devfeed.tech/topics/computing.md>), [parallel](<https://devfeed.tech/topics/parallel.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>), [networking](<https://devfeed.tech/topics/networking.md>)

Tags: [amazon-ec2](<https://devfeed.tech/tags/amazon-ec2.md>), [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-parallel-computing-service](<https://devfeed.tech/tags/aws-parallel-computing-service.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [computing](<https://devfeed.tech/tags/computing.md>), [configuration-management](<https://devfeed.tech/tags/configuration-management.md>), [high-performance-computing](<https://devfeed.tech/tags/high-performance-computing.md>), [hpc](<https://devfeed.tech/tags/hpc.md>), [life-sciences](<https://devfeed.tech/tags/life-sciences.md>), [parallel](<https://devfeed.tech/tags/parallel.md>), [research-computing](<https://devfeed.tech/tags/research-computing.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [scientific-computing](<https://devfeed.tech/tags/scientific-computing.md>), [simulations](<https://devfeed.tech/tags/simulations.md>), [slurm](<https://devfeed.tech/tags/slurm.md>)

### AI overview

This AWS blog describes how Daiichi Sankyo modernized its drug-discovery research environment by adopting AWS Parallel Computing Service for managed HPC cluster operations. It discusses the architecture, usability, operational benefits, and use of Slurm with Amazon EC2 and AWS storage services.

### Source excerpt

This blog was co-authored by Takehiro Nakajima and Mark Azadpour from AWS and Rintaro Yamada, Rei Kajitani and Ryo Kunimoto from Daiichi Sankyo In recent years, the informatics field of drug discovery has seen a rapid increase in workloads requiring large-scale parallel computing, such as genome analysis, structure prediction, and drug design. Daiichi Sankyo has [...]