# ec2

Published articles for ec2.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Amazon Linux 2027 Enters Public Preview with SELinux Enforcing by Default

DevFeed: [Amazon Linux 2027 Enters Public Preview with SELinux Enforcing by Default](<https://devfeed.tech/articles/amazon-linux-2027-enters-public-preview-with-selinux-enforcing-by-default-26598.md>)

Original publisher: [Read original article](<https://www.infoq.com/news/2026/09/amazon-linux-2027-preview/>)

Author: Steef-Jan Wiggers

Published: 2026-09-15T03:40:00Z

Content type: news

Language: en

Sources: [InfoQ](<https://devfeed.tech/sources/infoq.md>)

Topics: [Linux](<https://devfeed.tech/topics/linux.md>), [SELinux](<https://devfeed.tech/topics/selinux.md>), [migration](<https://devfeed.tech/topics/migration.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Kernel](<https://devfeed.tech/topics/kernel.md>), [amazon](<https://devfeed.tech/topics/amazon.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [amazon-linux-2027-preview](<https://devfeed.tech/tags/amazon-linux-2027-preview.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [architecture-design](<https://devfeed.tech/tags/architecture-design.md>), [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [development](<https://devfeed.tech/tags/development.md>), [devops](<https://devfeed.tech/tags/devops.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kernel](<https://devfeed.tech/tags/kernel.md>), [linux](<https://devfeed.tech/tags/linux.md>), [migration](<https://devfeed.tech/tags/migration.md>), [news](<https://devfeed.tech/tags/news.md>), [operating-systems](<https://devfeed.tech/tags/operating-systems.md>), [preview](<https://devfeed.tech/tags/preview.md>), [selinux](<https://devfeed.tech/tags/selinux.md>), [x86-64](<https://devfeed.tech/tags/x86-64.md>)

### AI overview

AWS has released Amazon Linux 2027 in public preview with SELinux enforcing by default. The change may require application and migration work because policies that only logged violations on AL2023 can block applications on AL2027. The release has no announced AL2023 end-of-support date, general-availability date, or in-place migration path.

### Source excerpt

AWS has released Amazon Linux 2027 in public preview, built on the AL2023 baseline with kernel 7.1 and SELinux in enforcing mode by default. Applications that pass on AL2023's permissive mode may fail under enforcing. The announcement gives no AL2023 end-of-support date, no GA date, and no in-place migration path. By Steef-Jan Wiggers

## How AI Cost Management Agents Connect Cloud Spending to Engineering Workflows

DevFeed: [How AI Cost Management Agents Connect Cloud Spending to Engineering Workflows](<https://devfeed.tech/articles/cloud-costs-are-ignored-how-ai-cost-agents-help-engineers-13499.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/why-engineers-ignore-cloud-costs-and-how-ai-cost-management-agents-fix-it>)

Author: Kelsey Rosen

Published: 2026-09-11T00:00:00Z

Content type: opinion

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [cloud cost management](<https://devfeed.tech/topics/cloud-cost-management.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-cost-management](<https://devfeed.tech/tags/cloud-cost-management.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [inference](<https://devfeed.tech/tags/inference.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [llm](<https://devfeed.tech/tags/llm.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>)

### AI overview

The article argues that engineers overlook cloud costs because spending data arrives too late and lacks the context needed for action. It explains how AI cost management agents can surface AI and cloud infrastructure spend in real time, flag waste, right-size resources, and enforce budget policies within engineering workflows.

### Source excerpt

Engineers ignore cloud costs because of broken feedback loops, not apathy. Learn what AI cost management is, why AEO matters more than ever. | Blog

## Replication Lag on AWS FSx: The Hidden EC2 Single-Flow Bandwidth Limit

DevFeed: [Replication Lag on AWS FSx: The Hidden EC2 Single-Flow Bandwidth Limit](<https://devfeed.tech/articles/replication-lag-on-aws-fsx-the-hidden-ec2-single-flow-bandwidth-limit-14112.md>)

Original publisher: [Read original article](<https://www.percona.com/blog/replication-lag-on-aws-fsx-the-hidden-ec2-single-flow-bandwidth-limit/>)

Author: Pablo Svampa

Published: 2026-08-25T21:00:16Z

Content type: article

Language: en

Sources: [Blog - Percona](<https://devfeed.tech/sources/blog-percona.md>)

Topics: [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Network](<https://devfeed.tech/topics/network.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [database](<https://devfeed.tech/tags/database.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [hardware-and-storage](<https://devfeed.tech/tags/hardware-and-storage.md>), [insight-for-dbas](<https://devfeed.tech/tags/insight-for-dbas.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [network](<https://devfeed.tech/tags/network.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [replication](<https://devfeed.tech/tags/replication.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

This Percona case study examines PostgreSQL replication lag on an EC2 instance using Amazon FSx over NFS. Although AWS reported healthy storage with substantial provisioned capacity remaining, vmstat showed persistent I/O blocking and wait, indicating a bottleneck somewhere across the storage and network path, including an EC2 single-flow bandwidth limit.

### Source excerpt

A recent case in our Percona Support team started with a familiar complaint. A PostgreSQL standby lagging behind its primary. Although the problem was simple, it brought a specific flavor that's worth sharing. The customer had already reached out to AWS Support about the storage layer behind the database, an Amazon FSx filesystem mounted over ... Continued The post Replication Lag on AWS FSx: The Hidden EC2 Single-Flow Bandwidth Limit appeared first on Percona.

## Infrastructure Control Plane | Day 2 Operations & Drift

DevFeed: [Infrastructure Control Plane | Day 2 Operations & Drift](<https://devfeed.tech/articles/infrastructure-control-plane-day-2-operations-drift-13429.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/infrastructure-breaks-after-deployment-why-day-2-operations-demand-a-control-plane>)

Author: Nicole Morgan

Published: 2026-08-06T00:00:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Infrastructure as code](<https://devfeed.tech/topics/infrastructure-as-code.md>), [Software-defined networking](<https://devfeed.tech/topics/sdn.md>), [Provisioning](<https://devfeed.tech/topics/provisioning.md>), [Ansible](<https://devfeed.tech/topics/ansible.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [AWS CloudFormation](<https://devfeed.tech/topics/aws-cloudformation.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [ansible](<https://devfeed.tech/tags/ansible.md>), [aws](<https://devfeed.tech/tags/aws.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [cloudformation](<https://devfeed.tech/tags/cloudformation.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [incident](<https://devfeed.tech/tags/incident.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [infrastructure-as-code](<https://devfeed.tech/tags/infrastructure-as-code.md>), [provisioning](<https://devfeed.tech/tags/provisioning.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

This article explains that infrastructure problems often emerge during day 2 operations, after initial deployment. It describes how console changes, incidents, and isolated Terraform, Ansible, and CI/CD workflows create infrastructure drift, and argues that control planes can enforce governance, detect divergence, and automate remediation.

### Source excerpt

Learn why infrastructure breaks after deployment and how control planes enforce governance, detect drift, and automate remediation across Terraform, Ansible, an | Blog

## Cloud Cost Visibility at Scale: Best Practices

DevFeed: [Cloud Cost Visibility at Scale: Best Practices](<https://devfeed.tech/articles/cloud-cost-visibility-at-scale-best-practices-13497.md>)

Original publisher: [Read original article](<https://www.harness.io/blog/why-cloud-cost-visibility-at-scale-fails-and-how-to-fix-it>)

Author: Kelsey Rosen

Published: 2026-08-04T00:00:00Z

Content type: article

Language: en

Sources: [Harness Blog](<https://devfeed.tech/sources/harness-blog.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [finops](<https://devfeed.tech/topics/finops.md>), [cloud cost management](<https://devfeed.tech/topics/cloud-cost-management.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Microservice](<https://devfeed.tech/topics/microservice.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-cost-management](<https://devfeed.tech/tags/cloud-cost-management.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [finops](<https://devfeed.tech/tags/finops.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [microservices](<https://devfeed.tech/tags/microservices.md>), [s3](<https://devfeed.tech/tags/s3.md>)

### AI overview

The article explains why cloud cost visibility and allocation practices break down as organizations scale across accounts, teams, services, and workloads. It argues that periodic reporting, manual reviews, and static tagging models become ineffective, and highlights FinOps governance and automation as ways to improve control and trust in cost data.

### Source excerpt

Discover why cloud cost visibility at scale breaks down and how to regain control. Learn proven FinOps strategies. Explore now. | Blog

## AWS Lambda MicroVMs Are Positioned Between Lambda and EC2

DevFeed: [AWS Lambda MicroVMs Are Positioned Between Lambda and EC2](<https://devfeed.tech/articles/aws-lambda-microvms-what-new-devilry-38704.md>)

Original publisher: [Read original article](<https://dataengineeringcentral.substack.com/p/aws-lambda-microvms-what-new-devilry>)

Author: Daniel Beach

Published: 2026-07-06T12:24:57Z

Content type: opinion

Language: en

Sources: [Data Engineering Central](<https://devfeed.tech/sources/data-engineering-central.md>)

Topics: [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [lambda](<https://devfeed.tech/tags/lambda.md>)

### AI overview

The article offers commentary on AWS Lambda MicroVMs, describing them as a compute option between AWS Lambda and EC2 instances. It considers whether they address broader demand or a narrower niche, but the supplied text does not provide a definitive conclusion.

### Source excerpt

example with Polars and CSVs

## Tuning Android build nodes for maximum throughput

DevFeed: [Tuning Android build nodes for maximum throughput](<https://devfeed.tech/articles/tuning-android-build-nodes-for-maximum-throughput-20464.md>)

Original publisher: [Read original article](<https://eng.wealthfront.com/2026/06/25/tuning-android-build-nodes-for-maximum-throughput/>)

Author: Chris Mathew

Published: 2026-06-25T15:29:45Z

Content type: tutorial

Language: en

Sources: [Wealthfront](<https://devfeed.tech/sources/wealthfront.md>)

Topics: [Gradle](<https://devfeed.tech/topics/gradle.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Android](<https://devfeed.tech/topics/android.md>), [ci](<https://devfeed.tech/topics/ci.md>), [Processes](<https://devfeed.tech/topics/processes.md>)

Tags: [android](<https://devfeed.tech/tags/android.md>), [caching](<https://devfeed.tech/tags/caching.md>), [ci](<https://devfeed.tech/tags/ci.md>), [continuous-integration](<https://devfeed.tech/tags/continuous-integration.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [gradle](<https://devfeed.tech/tags/gradle.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [performance](<https://devfeed.tech/tags/performance.md>), [wealthfront-engineering](<https://devfeed.tech/tags/wealthfront-engineering.md>)

### AI overview

This article explains how Wealthfront's Android team tuned Gradle for build workloads on Amazon m7i.8xlarge EC2 instances. It covers Gradle daemon configuration, JVM memory settings, parallel execution, build caching, and the choice to use single-use daemons for CI builds.

### Source excerpt

Introduction Our Android team uses Gradle as our build tool of choice. Gradle offers lots of options for tuning its resource consumption, giving engineers an opportunity to optimize performance for running tasks on well-known hardware. In this post, we'll explore how the team tuned our Gradle setup for Amazon's m7i.8xlarge EC2 instances. General concepts Before... Read more

## From Single Instance to Split-Brain: A Database Scaling Journey

DevFeed: [From Single Instance to Split-Brain: A Database Scaling Journey](<https://devfeed.tech/articles/from-single-instance-to-split-brain-a-database-scaling-journey-22540.md>)

Original publisher: [Read original article](<https://medium.com/walmartglobaltech/from-single-instance-to-split-brain-a-database-scaling-journey-8b6a27a65023?source=rss----905ea2b3d4d1---4>)

Author: Alok Mishra

Published: 2026-03-31T18:40:52Z

Content type: tutorial

Language: en

Sources: [Walmart Global Tech](<https://devfeed.tech/sources/walmart-global-tech.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [Self-hosted](<https://devfeed.tech/topics/self-hosted.md>), [backups](<https://devfeed.tech/topics/backups.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [backups](<https://devfeed.tech/tags/backups.md>), [bare-metal](<https://devfeed.tech/tags/bare-metal.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-services](<https://devfeed.tech/tags/cloud-services.md>), [cloud-sql](<https://devfeed.tech/tags/cloud-sql.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [database](<https://devfeed.tech/tags/database.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [failover](<https://devfeed.tech/tags/failover.md>), [google](<https://devfeed.tech/tags/google.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [google-cloud-sql](<https://devfeed.tech/tags/google-cloud-sql.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [production](<https://devfeed.tech/tags/production.md>), [read-replica](<https://devfeed.tech/tags/read-replica.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [software-architecture](<https://devfeed.tech/tags/software-architecture.md>), [software-development](<https://devfeed.tech/tags/software-development.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>)

### AI overview

This article explains how database scaling commonly uses a single-leader architecture with asynchronous read replicas. It discusses replication lag, stale reads, split-brain risks, read/write traffic separation, and the operational responsibilities of self-managed versus fully managed database services.

### Source excerpt

I used to think adding a 'Read Replica' was a magic button for scaling applications. I was wrong. While splitting read and write traffic is a standard system design pattern, implementing it introduces a world of pain - from stale reads to the dreaded Split-Brain problem. Here is how database replication actually works, and how to survive the transition. When people talk about "scaling databases" or "adding read replicas", they are almost always thinking about one specific architecture: Single-leader (Primary-Replica) architecture with asynchronous replication This is the architecture used by: MySQL + replicas PostgreSQL + streaming replication Google Cloud SQL PlanetScale, Neon, Supabase, etc. There is exactly one node that accepts writes -> called the Primary (or Leader/Master). All other nodes are Read Replicas -> they apply changes from the primary as fast as they can, but always with some delay (replication lag). This is the default and dominant model in 99% of applications today. Alternative architectures exist (multi-primary, leaderless, CRDTs, etc.), but they are rare and come with their own very different trade-offs. The second axis that actually matters in practice is: Who manages the replicas and failover for you?1. Self-hosted / Self-managed You run MySQL or PostgreSQL yourself (on EC2, Kubernetes, bare metal, etc.). You are 100% responsible for: Setting up replication Promoting a new primary when the old one dies Routing traffic correctly Handling replication lag Monitoring, backups, point-in-time recovery, etc. 2. Fully-managed cloud services RDS, Aurora, PlanetScale, Neon, Supabase, CockroachDB, Spanner, YugabyteDB, etc. The provider gives you a single connection string (or two: one for writes, one for reads) and magically keeps it pointing to healthy nodes, handles failover in seconds, and often hides (or eliminates) replication lag headaches. This second axis is the one that determines how much pain you will actually feel in production. Now, suppose yo

## Wix Migrated Its MySQL EC2 Fleet from Intel CPUs to Graviton

DevFeed: [Wix Migrated Its MySQL EC2 Fleet from Intel CPUs to Graviton](<https://devfeed.tech/articles/1-000-servers-160-clusters-30-days-zero-downtime-migrating-wix-s-mysql-fleet-to-graviton-22627.md>)

Original publisher: [Read original article](<https://www.wix.engineering/post/1-000-servers-160-clusters-30-days-zero-downtime-migrating-wix-s-mysql-fleet-to-graviton>)

Author: Wix Engineering

Published: 2026-03-23T11:27:39Z

Content type: article

Language: en

Sources: [Wix Engineering](<https://devfeed.tech/sources/wix-engineering.md>)

Topics: [Graviton](<https://devfeed.tech/topics/graviton.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Ansible](<https://devfeed.tech/topics/ansible.md>), [awx](<https://devfeed.tech/topics/awx.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [intel](<https://devfeed.tech/topics/intel.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [DevOps](<https://devfeed.tech/topics/devops.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [Server](<https://devfeed.tech/topics/server.md>), [Linux](<https://devfeed.tech/topics/linux.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [ansible](<https://devfeed.tech/tags/ansible.md>), [awx](<https://devfeed.tech/tags/awx.md>), [devops](<https://devfeed.tech/tags/devops.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [graviton](<https://devfeed.tech/tags/graviton.md>), [infra](<https://devfeed.tech/tags/infra.md>), [intel](<https://devfeed.tech/tags/intel.md>), [linux](<https://devfeed.tech/tags/linux.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [rdbms](<https://devfeed.tech/tags/rdbms.md>), [replication](<https://devfeed.tech/tags/replication.md>)

### AI overview

Wix's DB Infra team migrated more than 1,000 MySQL EC2 servers across over 160 clusters from Intel-based CPUs to Graviton in one month, ahead of the original three-month timeline. The project also enabled a move to Amazon Linux 2023 and away from CentOS 7.0, using AWX workflows and Wix's Dev Portal to automate the process while accounting for replication and cluster reliability concerns.

### Source excerpt

Last January, the DB Infra team at Wix embarked on a strategic and very important mission: Migrating all of our MySQL EC2 servers from Intel based CPUs to the new shiny Graviton ones. Not only will this allow us to work on better CPUs, saving us money and improving our databases, we'll be able to move to Amazon Linux 2023 OS and move from the EOL CentOS 7.0. That sounds simple enough, until you realize that we're dealing with over 1K servers spread over 160+ MySQL clusters. The initial...

## HubSpot Incident Report for October 20, 2025

DevFeed: [HubSpot Incident Report for October 20, 2025](<https://devfeed.tech/articles/hubspot-incident-report-for-october-20-2025-29106.md>)

Original publisher: [Read original article](<https://product.hubspot.com/blog/incident-report-for-october-20-2025>)

Author: Kartik Vishwanath

Published: 2025-11-11T17:09:36Z

Content type: news

Language: en

Sources: [HubSpot](<https://devfeed.tech/sources/hubspot.md>)

Topics: [incident](<https://devfeed.tech/topics/incident.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Amazon Simple Queue Service (SQS)](<https://devfeed.tech/topics/amazon-simple-queue-service-sqs.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [AWS IAM](<https://devfeed.tech/topics/aws-iam.md>), [DynamoDB](<https://devfeed.tech/topics/dynamodb.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [amazon-sqs](<https://devfeed.tech/tags/amazon-sqs.md>), [analysis](<https://devfeed.tech/tags/analysis.md>), [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [dynamodb](<https://devfeed.tech/tags/dynamodb.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [failover](<https://devfeed.tech/tags/failover.md>), [hubspot](<https://devfeed.tech/tags/hubspot.md>), [iam](<https://devfeed.tech/tags/iam.md>), [improvements](<https://devfeed.tech/tags/improvements.md>), [incident](<https://devfeed.tech/tags/incident.md>), [message-queue](<https://devfeed.tech/tags/message-queue.md>), [outage](<https://devfeed.tech/tags/outage.md>), [recovery](<https://devfeed.tech/tags/recovery.md>), [report](<https://devfeed.tech/tags/report.md>), [resilience](<https://devfeed.tech/tags/resilience.md>)

### AI overview

HubSpot reports that a severe AWS outage in the us-east-1 region on October 20, 2025 disrupted multiple product features and affected third-party vendors. The incident caused failures involving DynamoDB, IAM, SQS, and EC2, and degraded HubSpot's TQ2 background task processing. HubSpot says it completed an analysis and is implementing resilience improvements.

### Source excerpt

On October 20, 2025, HubSpot experienced a significant service disruption affecting multiple product features due to a severe AWS outage in the us-east-1 region. While our infrastructure remained intact, the widespread nature of the cloud provider failure impacted both our services and critical third-party vendors we rely on. We've completed a thorough analysis of this incident and are implementing comprehensive improvements to strengthen our resilience against future cloud provider disruptions.

## ECS on EC2: Covering Gaps in IMDS Hardening

DevFeed: [ECS on EC2: Covering Gaps in IMDS Hardening](<https://devfeed.tech/articles/ecs-on-ec2-covering-gaps-in-imds-hardening-29187.md>)

Original publisher: [Read original article](<https://www.latacora.com/blog/2025/10/02/ecs-on-ec2-covering-gaps-in-imds-hardening/>)

Published: 2025-10-02T18:18:59Z

Content type: tutorial

Language: en

Sources: [Latacora](<https://devfeed.tech/sources/latacora.md>)

Topics: [Amazon Elastic Container Service](<https://devfeed.tech/topics/amazon-elastic-container-service.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Security](<https://devfeed.tech/topics/security.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Containers](<https://devfeed.tech/topics/containers.md>), [Credential theft](<https://devfeed.tech/topics/credential-theft.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [containers](<https://devfeed.tech/tags/containers.md>), [credential-theft](<https://devfeed.tech/tags/credential-theft.md>), [credentials](<https://devfeed.tech/tags/credentials.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [ecs](<https://devfeed.tech/tags/ecs.md>), [hardening](<https://devfeed.tech/tags/hardening.md>), [security](<https://devfeed.tech/tags/security.md>), [sensitive-data](<https://devfeed.tech/tags/sensitive-data.md>)

### AI overview

This article examines security gaps in Amazon ECS workloads running on EC2, focusing on task isolation and restricting access to the EC2 Instance Metadata Service. It discusses how weak isolation can expose credentials and sensitive data and outlines the need for more comprehensive hardening guidance.

### Source excerpt

Introduction # AWS ECS is a widely-adopted service across industries. To illustrate the scale and ubiquity of this service, over 2.4 billion Amazon Elastic Container Service tasks are launched every week (source) and over 65% of all new AWS containers customers use Amazon ECS (source). There are two primary launch types for ECS: Fargate and EC2. The choice between them depends on factors like cost, performance, operational overhead, and the variability of your workload.

## Cost Optimisation in ECS: Integrating Spot Instances at Scale

DevFeed: [Cost Optimisation in ECS: Integrating Spot Instances at Scale](<https://devfeed.tech/articles/cost-optimisation-in-ecs-integrating-spot-instances-at-scale-19717.md>)

Original publisher: [Read original article](<https://deliveroo.engineering/2025/09/12/cost-optimisation-in-ecs.html>)

Author: Aakash Singhal

Published: 2025-09-12T00:00:00Z

Content type: article

Language: en

Sources: [Deliveroo](<https://devfeed.tech/sources/deliveroo.md>)

Topics: [Amazon Elastic Container Service](<https://devfeed.tech/topics/amazon-elastic-container-service.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Automation](<https://devfeed.tech/topics/automation.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [container](<https://devfeed.tech/topics/container.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [availability](<https://devfeed.tech/tags/availability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [container](<https://devfeed.tech/tags/container.md>), [cost](<https://devfeed.tech/tags/cost.md>), [cost-optimisation](<https://devfeed.tech/tags/cost-optimisation.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [ecs](<https://devfeed.tech/tags/ecs.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [load-balancer](<https://devfeed.tech/tags/load-balancer.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [scale](<https://devfeed.tech/tags/scale.md>), [stateless](<https://devfeed.tech/tags/stateless.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

Deliveroo describes how it integrated EC2 Spot Instances into Amazon ECS to reduce compute costs while maintaining service stability. The approach routes only eligible workloads to Spot capacity and uses criteria such as fast shutdown, task redundancy, statelessness, and load balancer deregistration timing.

### Source excerpt

At Deliveroo, we're always refining how we scale - especially when it comes to managing compute costs in the cloud. After optimising our Amazon ECS workloads with Reserved Instances and Savings Plans, we saw an opportunity to push further using EC2 Spot Instances, which offer up to 90% savings compared to On-Demand prices. But Spot comes with challenges: Their availability can fluctuate, and they can be terminated with just a two-minute warning. To unlock these savings without compromising service stability, we had to engineer a robust solution across infrastructure, workload qualification, and automation. The Challenge: Balancing Cost and Reliability Our ECS infrastructure initially relied entirely on On-Demand EC2 instances, provisioned through Auto Scaling Groups (ASGs) connected to ECS Capacity Providers. While reliable, this approach didn't take advantage of AWS's surplus compute capacity. We aimed to layer Spot Instances into our clusters, but selectively. Our goal was clear: route only eligible workloads to Spot capacity while ensuring no service degradation during unexpected terminations. Spot Instances: Power and Pitfalls Spot Instances provide dramatic cost reductions but introduce several operational caveats: Ephemeral by nature: AWS can terminate them at any time with a two-minute warning. Capacity variability: Availability depends on AWS's excess capacity in each AZ and can shift unpredictably. Scaling limitations: Auto Scaling may fail if the desired instance types are not currently available. To avoid introducing fragility into our stack, we established technical eligibility criteria that workloads must meet before being scheduled on Spot. Defining Spot Eligibility We formalised the following criteria to assess whether a workload could safely tolerate Spot interruptions: Fast Shutdown Support Constraint: stopTimeout must be < 120 seconds in the container definition. Reason: Ensures ECS has time to gracefully shut down the task before AWS's 2-minute te

## How to run Firecracker without KVM on cloud VMs

DevFeed: [How to run Firecracker without KVM on cloud VMs](<https://devfeed.tech/articles/how-to-run-firecracker-without-kvm-on-cloud-vms-26643.md>)

Original publisher: [Read original article](<https://blog.alexellis.io/how-to-run-firecracker-without-kvm-on-regular-cloud-vms/>)

Author: Alex Ellis

Published: 2025-02-12T09:05:21Z

Content type: tutorial

Language: en

Sources: [Alex Ellis' Blog](<https://devfeed.tech/sources/alex-ellis-blog.md>)

Topics: [Firecracker](<https://devfeed.tech/topics/firecracker.md>), [virtualization](<https://devfeed.tech/topics/virtualization.md>), [virtual machines](<https://devfeed.tech/topics/virtual-machines.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [bare-metal](<https://devfeed.tech/tags/bare-metal.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [firecracker](<https://devfeed.tech/tags/firecracker.md>), [github-actions](<https://devfeed.tech/tags/github-actions.md>), [kvm](<https://devfeed.tech/tags/kvm.md>), [qemu](<https://devfeed.tech/tags/qemu.md>), [virtual-machines](<https://devfeed.tech/tags/virtual-machines.md>), [virtualization](<https://devfeed.tech/tags/virtualization.md>)

### AI overview

This tutorial introduces a way to run microVMs on cloud virtual machines without KVM, using the PVM virtualization framework. It explains the limitations of nested virtualization and compares the cost of AWS bare-metal EC2 with alternatives.

### Source excerpt

MicroVMs need bare-metal or nested virtualisation with /dev/kvm. But what if that's not available? The PVM virtualisation framework may be the answer.

## Failover in Amazon RDS Multi-AZ Architectures

DevFeed: [Failover in Amazon RDS Multi-AZ Architectures](<https://devfeed.tech/articles/failover-in-amazon-rds-multi-az-architectures-18013.md>)

Original publisher: [Read original article](<https://blog.guilleojeda.com/failover-in-amazon-rds-multi-az-architectures>)

Author: Guillermo Ojeda

Published: 2024-12-18T23:22:08Z

Content type: tutorial

Language: en

Sources: [Guille Ojeda](<https://devfeed.tech/sources/guille-ojeda.md>)

Topics: [Amazon RDS](<https://devfeed.tech/topics/amazon-rds.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Cloud Architecture](<https://devfeed.tech/topics/cloud-architecture.md>), [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Database](<https://devfeed.tech/topics/database.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [VPC](<https://devfeed.tech/topics/vpc.md>)

Tags: [amazon-rds](<https://devfeed.tech/tags/amazon-rds.md>), [amazon-web-services](<https://devfeed.tech/tags/amazon-web-services.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [databases](<https://devfeed.tech/tags/databases.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [failover](<https://devfeed.tech/tags/failover.md>), [replication](<https://devfeed.tech/tags/replication.md>), [vpc](<https://devfeed.tech/tags/vpc.md>)

### AI overview

This tutorial explains how Amazon RDS Multi-AZ deployments handle database failures and failover. It examines RDS architecture, including EC2 compute, EBS storage, VPC connectivity, the control plane, storage-level replication, and Availability Zone isolation.

### Source excerpt

Database failures are inevitable. Even with the most reliable hardware and software, something will eventually break. AWS RDS Multi-AZ deployments promise to handle these failures gracefully, automatically failing over to a standby database when prob...

## Simplify and Secure AWS Access to Accelerate Outcomes: 3 Best Practices

DevFeed: [Simplify and Secure AWS Access to Accelerate Outcomes: 3 Best Practices](<https://devfeed.tech/articles/simplify-and-secure-aws-access-to-accelerate-outcomes-3-best-practices-29941.md>)

Original publisher: [Read original article](<https://goteleport.com/blog/three-best-practices-secure-and-simplify-aws-access/>)

Author: info@goteleport.com (Jack Pitts)

Published: 2024-12-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Teleport](<https://devfeed.tech/sources/teleport.md>)

Topics: [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Security](<https://devfeed.tech/topics/security.md>), [Development](<https://devfeed.tech/topics/development.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [cloud-infrastructure](<https://devfeed.tech/tags/cloud-infrastructure.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [complexity](<https://devfeed.tech/tags/complexity.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [databases](<https://devfeed.tech/tags/databases.md>), [development](<https://devfeed.tech/tags/development.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [security](<https://devfeed.tech/tags/security.md>), [visibility](<https://devfeed.tech/tags/visibility.md>)

### AI overview

This article presents three best practices for simplifying and securing AWS access as cloud infrastructure grows. It focuses on infrastructure sprawl, granular access controls, and improved access visibility, while describing effects on engineering productivity, management overhead, misconfiguration risk, and compliance.

### Source excerpt

Discover three best practices to simplify and secure AWS access: addressing infrastructure sprawl, adding granular controls, and improving access visibility.

## Best Practices for Testing Zone Redundancy

DevFeed: [Best Practices for Testing Zone Redundancy](<https://devfeed.tech/articles/best-practices-for-testing-zone-redundancy-11561.md>)

Original publisher: [Read original article](<https://www.gremlin.com/blog/best-practices-for-testing-zone-redundancy>)

Author: Sam Rossoff

Published: 2024-10-16T00:00:00Z

Content type: article

Language: en

Sources: [Gremlin Blog](<https://devfeed.tech/sources/gremlin-blog.md>)

Topics: [Testing](<https://devfeed.tech/topics/testing.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [data-centers](<https://devfeed.tech/tags/data-centers.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

The article explains why deploying software across multiple zones or regions does not by itself ensure resilience. It presents testing as a way to verify zonal redundancy and discusses failures involving packet loss, configuration, dependencies, capacity, quotas, fleet balance, and split-brain tolerance.

### Source excerpt

Gremlin Principal Software Engineer Sam Rossoff shares key best practices and strategies for effectively testing zone redundancy.

## A Batch of nixbuild.net Updates

DevFeed: [A Batch of nixbuild.net Updates](<https://devfeed.tech/articles/a-batch-of-nixbuild-net-updates-34142.md>)

Original publisher: [Read original article](<https://blog.nixbuild.net/posts/2024-10-16-a-batch-of-nixbuild-net-updates.html>)

Author: support@nixbuild.net

Published: 2024-10-16T00:00:00Z

Content type: article

Language: en

Sources: [nixbuild.net blog](<https://devfeed.tech/sources/nixbuild-net-blog.md>)

Topics: [Self-hosted](<https://devfeed.tech/topics/self-hosted.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Provisioning](<https://devfeed.tech/topics/provisioning.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [builds](<https://devfeed.tech/topics/builds.md>), [configuration](<https://devfeed.tech/topics/configuration.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [builds](<https://devfeed.tech/tags/builds.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [provisioning](<https://devfeed.tech/tags/provisioning.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [self-hosting](<https://devfeed.tech/tags/self-hosting.md>)

### AI overview

This blog post summarizes nixbuild.net updates, focusing on enterprise self-hosting. It describes an EC2 AMI-based deployment using NixOS, EC2 provisioning, and configuration files or cloud-init, with support for configurable builder instances and build provisioning.

### Source excerpt

We have not written about nixbuild.net for over a year now, but that doesn't mean nothing has happened. Quite the opposite - we've been really busy! This blog post will try to summarize some of what we have been working on behind the scene the last couple of years. Read on to hear about nixbuild.net going enterprise, a brand new web UI and lots of new functionality! Self-Hosting nixbuild.net A big thing that has kept us busy the last year is packaging nixbuild.net for self-hosting inside enterprise setups. We've had a number of companies reaching out to us inquiring about such setups. The reason for wanting to self-host nixbuild.net are several: internal legal requirements, network bottlenecks, cost control and the ability to use hardware that is tricky to offer in the public service. One of our enterprise customers self-hosts nixbuild.net in production running 10K+ builds daily. A very welcome side-effect of adapting nixbuild.net for self-hosting is that many scenarios that are seldom or never triggered by users of our public service are uncovered by enterprise usage. This has made nixbuild.net even more reliable and capable under high load. We are in a stage where we are happy to work closely with companies, figuring out how to best deploy and integrate nixbuild.net into their environments. Our goal is to eventually be able to offer self-hosted nixbuild.net as an off-the-shelf product. Actually, if you are on AWS, the deployment of nixbuild.net is already very simple. We will publish a blog post with more details on this in the near future, but the short rundown is this: We provide an EC2 AMI that runs NixOS and all necessary nixbuild.net services. The user deploys this AMI using her preferred method of EC2 provisioning. The user provides nixbuild.net configuration through cloud-init or by simply putting configuration files in place on the EC2 instance. The nixbuild.net configuration includes details on what types of EC2 instances to use for running the builds, an

## Creating NetBSD EC2 AMIs

DevFeed: [Creating NetBSD EC2 AMIs](<https://devfeed.tech/articles/creating-netbsd-ec2-amis-30169.md>)

Original publisher: [Read original article](<https://www.netmeister.org/blog/creating-netbsd-ec2-amis.html>)

Published: 2024-08-17T19:40:12Z

Content type: tutorial

Language: en

Sources: [Signs of Triviality](<https://devfeed.tech/sources/signs-of-triviality.md>)

Topics: [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Tool](<https://devfeed.tech/topics/tool.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

A short guide to creating a NetBSD Amazon Machine Image for Amazon EC2 using the bsdec2-image-upload tool.

### Source excerpt

A short description of how to create a NetBSD AMI for Amazon AWS EC2 using the 'bsdec2-image-upload' tool.

## How Indeed Replaced Its CI Platform with Gitlab CI

DevFeed: [How Indeed Replaced Its CI Platform with Gitlab CI](<https://devfeed.tech/articles/how-indeed-replaced-its-ci-platform-with-gitlab-ci-29990.md>)

Original publisher: [Read original article](<https://engineering.indeedblog.com/blog/2024/08/indeed-gitlab-ci-migration/>)

Author: Carl Myers

Published: 2024-08-06T15:03:51Z

Content type: article

Language: en

Sources: [Indeed](<https://devfeed.tech/sources/indeed.md>)

Topics: [GitLab](<https://devfeed.tech/topics/gitlab.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [Jenkins](<https://devfeed.tech/topics/jenkins.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [ci](<https://devfeed.tech/tags/ci.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [gitlab](<https://devfeed.tech/tags/gitlab.md>), [gitlab-ci](<https://devfeed.tech/tags/gitlab-ci.md>), [jenkins](<https://devfeed.tech/tags/jenkins.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [unsorted](<https://devfeed.tech/tags/unsorted.md>)

### AI overview

Indeed's engineering platform evolved from Hudson to Jenkins and later encountered architectural and scaling limitations as the company grew and adopted AWS EC2 and Kubernetes. The article discusses replacing that CI platform with GitLab CI.

### Source excerpt

Here at Indeed, our mission is to help people get jobs. Indeed is the #1 job site in the world with over 580M+ Job Seeker Profiles. For Indeed's Engineering Platform teams, we have a slightly different motto: "We help people to help people get jobs". As part of a data-driven engineering culture that has spent [...]

## Bazaarvoice Notification System for Transactional Email Delivery

DevFeed: [Bazaarvoice Notification System for Transactional Email Delivery](<https://devfeed.tech/articles/cloud-native-marvel-driving-6-million-daily-notifications-38724.md>)

Original publisher: [Read original article](<https://blog.developer.bazaarvoice.com/2024/04/24/cloud-native-marvel-driving-6-million-daily-notifications/>)

Author: Someswar Bhowmick

Published: 2024-04-24T10:24:56Z

Content type: article

Language: en

Sources: [Bazaarvoice](<https://devfeed.tech/sources/bazaarvoice.md>)

Topics: [notifications](<https://devfeed.tech/topics/notifications.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [AWS CloudFormation](<https://devfeed.tech/topics/aws-cloudformation.md>), [Jenkins](<https://devfeed.tech/topics/jenkins.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloudformation](<https://devfeed.tech/tags/cloudformation.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [email](<https://devfeed.tech/tags/email.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [jenkins](<https://devfeed.tech/tags/jenkins.md>), [notifications](<https://devfeed.tech/tags/notifications.md>), [s3](<https://devfeed.tech/tags/s3.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [software-architecture](<https://devfeed.tech/tags/software-architecture.md>)

### AI overview

Bazaarvoice describes its notification system for sending transactional email on behalf of clients. The article outlines data ingestion and a decision engine for scheduled delivery, then discusses operational challenges in its earlier AWS-based architecture, including EC2 scaling, updates, logging, and prolonged file-processing batch jobs.

### Source excerpt

Bazaarvoice notification system stands as a testament to cutting-edge technology, designed to seamlessly dispatch transactional email messages (post-interaction email or PIE) on behalf of our clients. The heartbeat of our system lies in the constant influx of new content, driven by active content solicitations. Equipped with an array of tools, including email message styling, default [...]

## Monitoring and Troubleshooting on AWS: CloudWatch, X-Ray, and Beyond

DevFeed: [Monitoring and Troubleshooting on AWS: CloudWatch, X-Ray, and Beyond](<https://devfeed.tech/articles/monitoring-and-troubleshooting-on-aws-cloudwatch-x-ray-and-beyond-18002.md>)

Original publisher: [Read original article](<https://blog.guilleojeda.com/aws-monitoring-troubleshooting-cloudwatch-xray>)

Author: Guillermo Ojeda

Published: 2024-03-23T16:33:24Z

Content type: tutorial

Language: en

Sources: [Guille Ojeda](<https://devfeed.tech/sources/guille-ojeda.md>)

Topics: [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Amazon CloudWatch Logs](<https://devfeed.tech/topics/amazon-cloudwatch-logs.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Cloud Architecture](<https://devfeed.tech/topics/cloud-architecture.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-architecture](<https://devfeed.tech/tags/cloud-architecture.md>), [devops](<https://devfeed.tech/tags/devops.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [troubleshooting](<https://devfeed.tech/tags/troubleshooting.md>)

### AI overview

This tutorial introduces AWS monitoring and troubleshooting, focusing on CloudWatch and X-Ray. It explains how CloudWatch collects operational data such as metrics, logs, and events, and describes using metrics and logs to understand resource performance and investigate application issues.

### Source excerpt

As an AWS user, I'm sure you know that monitoring and troubleshooting are essential for keeping your applications running smoothly. After all, you can't fix what you can't see. But with the sheer number of services and tools available on AWS, it can ...

## SSD Performance Has Advanced Faster Than Cloud Vendor NVMe Instances

DevFeed: [SSD Performance Has Advanced Faster Than Cloud Vendor NVMe Instances](<https://devfeed.tech/articles/ssds-have-become-ridiculously-fast-except-in-the-cloud-25084.md>)

Original publisher: [Read original article](<https://databasearchitects.blogspot.com/2024/02/ssds-have-become-ridiculously-fast.html>)

Author: Viktor Leis (noreply@blogger.com)

Published: 2024-02-19T08:00:00Z

Content type: opinion

Language: en

Sources: [Database Architects](<https://devfeed.tech/sources/database-architects.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [NVMe](<https://devfeed.tech/topics/nvme.md>), [pcie](<https://devfeed.tech/topics/pcie.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [capacity](<https://devfeed.tech/tags/capacity.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [io](<https://devfeed.tech/tags/io.md>), [nvme](<https://devfeed.tech/tags/nvme.md>), [pcie](<https://devfeed.tech/tags/pcie.md>), [performance](<https://devfeed.tech/tags/performance.md>), [server](<https://devfeed.tech/tags/server.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

The article compares rapid advances in commodity PCIe and NVMe SSD performance with slower progress in cloud-based NVMe instances. It argues that AWS EC2 instance performance per SSD has remained around 2 GB/s since the i3 launch, leaving a substantial gap between leading data center SSDs and those offered by major cloud vendors.

### Source excerpt

In recent years, flash-based SSDs have largely replaced disks for most storage use cases. Internally, each SSD consists of many independent flash chips, each of which can be accessed in parallel. Assuming the SSD controller keeps up, the throughput of an SSD therefore primarily depends on the interface speed to the host. In the past six years, we have seen a rapid transition from SATA to PCIe 3.0 to PCIe 4.0 to PCIe 5.0. As a result, there was an explosion in SSD throughput: At the same time, we saw not just better performance, but also more capacity per dollar: The two plots illustrate the power of a commodity market. The combination of open standards (NVMe and PCIe), huge demand, and competing vendors led to great benefits for customers. Today, top PCIe 5.0 data center SSDs such as the Kioxia CM7-R or Samsung PM1743 achieve up to 13 GB/s read throughput and 2.7M+ random read IOPS. Modern servers have around 100 PCIe lanes, making it possible to have a dozen of SSDs (each usually using 4 lanes) in a single server at full bandwidth. For example, in our lab we have a single-socket Zen 4 server with 8 Kioxia CM7-R SSDs, which achieves 100GB/s (!) I/O bandwidth: AWS EC2 was an early NVMe pioneer, launching the i3 instance with 8 physically-attached NVMe SSDs in early 2017. At that time, NVMe SSDs were still expensive, and having 8 in a single server was quite remarkable. The per-SSD read (2 GB/s) and write (1 GB/s) performance was considered state of the art as well. Another step forward occurred in 2019 with the launch of i3en instances, which doubled storage capacity per dollar. Since then, several NVMe instance types, including i4i and im4gn, have been launched. Surprisingly, however, the performance has not increased; seven years after the i3 launch, we are still stuck with 2 GB/s per SSD. Indeed, the venerable i3 and i3en instances basically remain the best EC2 has to offer in terms of IO-bandwidth/$ and SSD-capacity/$, respectively. Personally, I find this very s

## Teleport's 2023 product retrospective highlights Device Trust, automatic agent updates, agentless access, and Teleport Team

DevFeed: [Teleport's 2023 product retrospective highlights Device Trust, automatic agent updates, agentless access, and Teleport Team](<https://devfeed.tech/articles/is-santa-an-insider-threat-29622.md>)

Original publisher: [Read original article](<https://goteleport.com/blog/december-newsletter/>)

Author: ben@goteleport.com (Ben Arent)

Published: 2023-12-14T00:00:00Z

Content type: article

Language: en

Sources: [Teleport](<https://devfeed.tech/sources/teleport.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [releases](<https://devfeed.tech/topics/releases.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>), [Device Trust](<https://devfeed.tech/topics/device-trust.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [Amazon EKS](<https://devfeed.tech/topics/amazon-eks.md>), [trust](<https://devfeed.tech/topics/trust.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [device-trust](<https://devfeed.tech/tags/device-trust.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [releases](<https://devfeed.tech/tags/releases.md>), [retrospective](<https://devfeed.tech/tags/retrospective.md>), [saas](<https://devfeed.tech/tags/saas.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

Teleport's 2023 retrospective reviews releases and product highlights, including Device Trust, Windows host support, automatic agent updates, TLS routing, identity and security features, agentless access, and the Teleport Team SaaS offering.

### Source excerpt

A quick retrospective and focus on a few highlights for 2023.

## How Accurate Clocks Improve Observability in Distributed Systems

DevFeed: [How Accurate Clocks Improve Observability in Distributed Systems](<https://devfeed.tech/articles/it-s-about-time-12547.md>)

Original publisher: [Read original article](<http://brooker.co.za/blog/2023/11/27/about-time.html>)

Author: Marc Brooker

Published: 2023-11-27T00:00:00Z

Content type: article

Language: en

Sources: [Marc Brooker's Blog](<https://devfeed.tech/sources/marc-brooker-s-blog.md>), [Marc Brooker's Blog](<https://devfeed.tech/sources/marc-brooker-s-blog-2.md>)

Topics: [distributed-systems](<https://devfeed.tech/topics/distributed-systems.md>), [Amazon EC2](<https://devfeed.tech/topics/amazon-ec2.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [causality](<https://devfeed.tech/tags/causality.md>), [distributed-systems](<https://devfeed.tech/tags/distributed-systems.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [observability](<https://devfeed.tech/tags/observability.md>)

### AI overview

This article examines how increasingly accurate time synchronization, including microsecond-accurate clocks in Amazon EC2, can affect distributed-system design. It focuses on observability and explains that accurate timestamps make establishing causality and ordering events easier.

### Source excerpt

It's About Time! What's the time? Time to get a watch. My friend Al Vermeulen used to say time is for the amusement of humans1. Al's sentiment is still the common one among distributed systems builders: real wall-clock physical time is great for human-consumption (like log timestamps and UI presentation), but shouldn't be relied on by computer for things like actually affect the operation of the system. This remains a solid starting point, the right default position, but the picture has always been more subtle. Recently, the availability of ever-better time synchronization has made it even more subtle. This post will attempt to unravel some of that subtlety. Today is a good day to talk about time, because last week AWS announced (more details here) microsecond-accurate time synchronization in EC2, improving on what was already very good. All this means is that if you have an EC2 instance2 you can expect its clock to by accurate to within microseconds of the physical time. It turns out that having microsecond-level time accuracy makes some distributed systems stuff much easier than it was in the past. In hopes of understanding the controversy over using real time in systems, let's descend level-by-level into how we might entangle physical time more deeply in our system designs. Level 0: Observability, and the Amusement of Humans ... reality, the name we give to the common experience3 When we try understand how a system works, or why it's not working, the first task is to establish causality. Thing A caused thing B. Here in our weird little universe, we need thing A to have happened before thing B for A to have caused B. Time is useful for this. Prosecutor: Why, Mr Load Balancer, did you stop sending traffic to Mrs Server? Mr LB: Simply, sir, because she stopped processing my traffic! Mrs Server, from the gallery: Liar! Liar! I only stopped processing because you stopped sending! If we can't trust the order of our logs (or other events), finding causality is difficult.

[Next page](<https://devfeed.tech/tags/ec2.md?cursor=WyIyMDIzLTExLTI3VDAwOjAwOjAwKzAwOjAwIiwgImFkYjc3NjI0LTEyZWEtNGY5NS05ZGI2LWU4YjA2MzA0ZjJkOSJd>)