# infrastructure monitoring

Published articles for infrastructure monitoring.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations

DevFeed: [Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations](<https://devfeed.tech/articles/monitoring-production-agent-lifecycle-with-aws-devops-agent-and-agentcore-evaluations-4737.md>)

Original publisher: [Read original article](<https://aws.amazon.com/blogs/machine-learning/monitoring-production-agent-lifecycle-with-aws-devops-agent-and-agentcore-evaluations/>)

Author: Meghana Ashok

Published: 2026-09-11T18:26:38Z

Content type: article

Language: en

Sources: [Artificial Intelligence](<https://devfeed.tech/sources/artificial-intelligence.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [AWS IAM](<https://devfeed.tech/topics/aws-iam.md>), [incident](<https://devfeed.tech/topics/incident.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>)

Tags: [advanced-300](<https://devfeed.tech/tags/advanced-300.md>), [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [amazon-bedrock](<https://devfeed.tech/tags/amazon-bedrock.md>), [amazon-bedrock-agentcore](<https://devfeed.tech/tags/amazon-bedrock-agentcore.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-identity-and-access-management-iam](<https://devfeed.tech/tags/aws-identity-and-access-management-iam.md>), [best-practices](<https://devfeed.tech/tags/best-practices.md>), [devops](<https://devfeed.tech/tags/devops.md>), [incident](<https://devfeed.tech/tags/incident.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [production](<https://devfeed.tech/tags/production.md>), [tracing](<https://devfeed.tech/tags/tracing.md>)

### AI overview

The article describes monitoring production multi-agent systems with Amazon Bedrock AgentCore Evaluations for continuous quality assessment and AWS DevOps Agent for autonomous infrastructure incident investigation.

### Source excerpt

Multi-agent systems fail in ways traditional monitoring misses. This post presents a dual-layer approach to monitoring production agents: Amazon Bedrock AgentCore Evaluations for continuous quality scoring and AWS DevOps Agent for autonomous infrastructure investigation, shown on a four-agent airline reservation system.

## Avoid Azure secret rotation with secretless authentication

DevFeed: [Avoid Azure secret rotation with secretless authentication](<https://devfeed.tech/articles/avoid-azure-secret-rotation-with-secretless-authentication-2231.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/azure-secretless-authentication/>)

Author: Cody Murray-Bruce; Ben Johnson-Staub

Published: 2026-08-12T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Azure](<https://devfeed.tech/topics/azure.md>), [Authentication](<https://devfeed.tech/topics/authentication.md>), [Entra ID](<https://devfeed.tech/topics/entra-id.md>), [OpenID connect (OIDC)](<https://devfeed.tech/topics/oidc.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Terraform](<https://devfeed.tech/topics/terraform.md>), [migration](<https://devfeed.tech/topics/migration.md>)

Tags: [authentication](<https://devfeed.tech/tags/authentication.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cli](<https://devfeed.tech/tags/cli.md>), [entra-id](<https://devfeed.tech/tags/entra-id.md>), [identity-and-access-management](<https://devfeed.tech/tags/identity-and-access-management.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [integration](<https://devfeed.tech/tags/integration.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [migration](<https://devfeed.tech/tags/migration.md>), [oidc](<https://devfeed.tech/tags/oidc.md>), [secrets](<https://devfeed.tech/tags/secrets.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [terraform](<https://devfeed.tech/tags/terraform.md>)

### AI overview

This article explains how Datadog's secretless authentication for Azure replaces long-lived client secrets with federated identity using Microsoft Entra ID and OpenID Connect. It describes the short-lived token flow and configuration or migration options through Quickstart, Terraform, the Azure CLI, or the Azure portal.

### Source excerpt

Learn how secretless authentication for Datadog's Azure integration helps prevent telemetry interruptions that result from expired client secrets.

## Provision Datadog on Stripe Projects

DevFeed: [Provision Datadog on Stripe Projects](<https://devfeed.tech/articles/provision-datadog-on-stripe-projects-2262.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/datadog-stripe-projects/>)

Author: Christy Smith; Vidur Khanna; Elizabeth George

Published: 2026-07-23T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [stripe](<https://devfeed.tech/topics/stripe.md>), [Provisioning](<https://devfeed.tech/topics/provisioning.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>), [datadog agent](<https://devfeed.tech/topics/datadog-agent.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [cli](<https://devfeed.tech/tags/cli.md>), [datadog-agent](<https://devfeed.tech/tags/datadog-agent.md>), [developer](<https://devfeed.tech/tags/developer.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [saas](<https://devfeed.tech/tags/saas.md>), [stripe](<https://devfeed.tech/tags/stripe.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

Stripe Projects lets developers provision Datadog from the Stripe CLI in two commands, creating a Datadog organization, trial, and API key while using an existing Stripe payment method.

### Source excerpt

Learn how to provision Datadog on Stripe Projects, start a 14-day free trial, and manage billing through your existing Stripe account.

## Use OpenTelemetry-native observability with Datadog from ingestion to investigation

DevFeed: [Use OpenTelemetry-native observability with Datadog from ingestion to investigation](<https://devfeed.tech/articles/use-opentelemetry-native-observability-with-datadog-from-ingestion-to-investigation-2300.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/native-otel-with-datadog/>)

Author: Shanel Huang

Published: 2026-07-20T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [OpenTelemetry](<https://devfeed.tech/topics/opentelemetry.md>), [observability](<https://devfeed.tech/topics/observability.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Application Performance Management (APM)](<https://devfeed.tech/topics/apm.md>), [infrastructure monitoring](<https://devfeed.tech/topics/infrastructure-monitoring.md>), [Instrumentation](<https://devfeed.tech/topics/instrumentation.md>), [Traces](<https://devfeed.tech/topics/traces.md>), [log management](<https://devfeed.tech/topics/log-management.md>)

Tags: [apm](<https://devfeed.tech/tags/apm.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [instrumentation](<https://devfeed.tech/tags/instrumentation.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [opentelemetry](<https://devfeed.tech/tags/opentelemetry.md>), [sdks](<https://devfeed.tech/tags/sdks.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [traces](<https://devfeed.tech/tags/traces.md>), [tracing](<https://devfeed.tech/tags/tracing.md>)

### AI overview

This article explains how Datadog supports OpenTelemetry-native observability from instrumentation through ingestion and investigation. It covers vendor-neutral OTLP ingestion through the OTel Collector's standard HTTP exporter and Datadog's direct OTLP intake for metrics, logs, and traces, including use cases such as managed services and serverless environments.

### Source excerpt

Learn how you can use vendor-neutral telemetry while preserving Datadog's infrastructure and APM experiences.

## Monitor Apigee X API traffic and security with Datadog

DevFeed: [Monitor Apigee X API traffic and security with Datadog](<https://devfeed.tech/articles/monitor-apigee-x-api-traffic-and-security-with-datadog-2292.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/monitor-apigee-x-api-traffic-and-security-with-datadog/>)

Author: Sydney McLaughlin; Cody Murray-Bruce; Eddie Cai

Published: 2026-07-15T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [telemetry](<https://devfeed.tech/topics/telemetry.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [integration](<https://devfeed.tech/tags/integration.md>), [latency](<https://devfeed.tech/tags/latency.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [security](<https://devfeed.tech/tags/security.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

This tutorial explains how to use Datadog's Google Cloud Apigee X integration to monitor API traffic, latency, anomalies, security scores, and related logs and traces.

### Source excerpt

Learn how to monitor Apigee X with Datadog so you can track API traffic, latency, anomalies, and security posture and catch issues before clients do.

## Datadog achieves GovRAMP High authorization

DevFeed: [Datadog achieves GovRAMP High authorization](<https://devfeed.tech/articles/datadog-achieves-govramp-high-authorization-2254.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/datadog-achieves-govramp-high-authorization/>)

Author: Rukshan Gunawardana; Sophie Wang; Jason Hansen

Published: 2026-06-29T00:00:00Z

Content type: release

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Security](<https://devfeed.tech/topics/security.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Resilience](<https://devfeed.tech/topics/resilience.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [agent-observability](<https://devfeed.tech/tags/agent-observability.md>), [ai-observability](<https://devfeed.tech/tags/ai-observability.md>), [aws](<https://devfeed.tech/tags/aws.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [dashboards](<https://devfeed.tech/tags/dashboards.md>), [database-monitoring](<https://devfeed.tech/tags/database-monitoring.md>), [fedramp](<https://devfeed.tech/tags/fedramp.md>), [government](<https://devfeed.tech/tags/government.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [log-management](<https://devfeed.tech/tags/log-management.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [networks](<https://devfeed.tech/tags/networks.md>), [observability](<https://devfeed.tech/tags/observability.md>), [resilience](<https://devfeed.tech/tags/resilience.md>), [security](<https://devfeed.tech/tags/security.md>), [systems](<https://devfeed.tech/tags/systems.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

Datadog announced GovRAMP High authorization for its government offering, extending its FedRAMP High certification and supporting state, local, tribal, territorial, and education organizations. The article explains how unified observability and security can consolidate monitoring data across applications, logs, networks, databases, infrastructure, and AI initiatives.

### Source excerpt

Datadog achieves GovRAMP High authorization, bringing unified observability and security to state and local government agencies for critical systems.

## Monitor Scaleway with Datadog

DevFeed: [Monitor Scaleway with Datadog](<https://devfeed.tech/articles/monitor-scaleway-with-datadog-2298.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/monitor-scaleway-with-datadog/>)

Author: Ellie Cohen; Eddie Cai

Published: 2026-06-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [log management](<https://devfeed.tech/topics/log-management.md>), [datadog agent](<https://devfeed.tech/topics/datadog-agent.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [audit trail](<https://devfeed.tech/topics/audit-trail.md>), [incident](<https://devfeed.tech/topics/incident.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [API](<https://devfeed.tech/topics/api.md>), [configuration](<https://devfeed.tech/topics/configuration.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [audit-trail](<https://devfeed.tech/tags/audit-trail.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [compliance](<https://devfeed.tech/tags/compliance.md>), [database](<https://devfeed.tech/tags/database.md>), [datadog-agent](<https://devfeed.tech/tags/datadog-agent.md>), [incident](<https://devfeed.tech/tags/incident.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [log-management](<https://devfeed.tech/tags/log-management.md>), [logs](<https://devfeed.tech/tags/logs.md>), [scaleway](<https://devfeed.tech/tags/scaleway.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [traces](<https://devfeed.tech/tags/traces.md>), [tracing](<https://devfeed.tech/tags/tracing.md>)

### AI overview

This post explains how the Datadog Scaleway integration collects telemetry from Scaleway Cockpit and Audit Trail, deploys the Datadog Agent on Scaleway compute instances, and monitors Kubernetes workloads. It also shows how to search and correlate Scaleway logs with infrastructure metrics, application logs, and traces in Datadog.

### Source excerpt

Use the Datadog Scaleway integration to collect logs from Cockpit and Audit Trail and monitor Scaleway alongside your other cloud providers.

## Monitor OVHcloud with Datadog

DevFeed: [Monitor OVHcloud with Datadog](<https://devfeed.tech/articles/monitor-ovhcloud-with-datadog-2296.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/monitor-ovhcloud-with-datadog/>)

Author: Ellie Cohen; Eddie Cai

Published: 2026-06-09T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [log management](<https://devfeed.tech/topics/log-management.md>), [observability](<https://devfeed.tech/topics/observability.md>), [Application Performance Management (APM)](<https://devfeed.tech/topics/apm.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [data](<https://devfeed.tech/topics/data.md>), [AWS IAM](<https://devfeed.tech/topics/aws-iam.md>)

Tags: [apm](<https://devfeed.tech/tags/apm.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [iam](<https://devfeed.tech/tags/iam.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [log-management](<https://devfeed.tech/tags/log-management.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

This post explains how to centralize OVHcloud logs in Datadog Log Management and collect metrics and APM traces from OVHcloud infrastructure. It focuses on consistent observability across multi-cloud environments, including centralized search, alerting, analysis, and out-of-the-box monitoring.

### Source excerpt

Centralize OVHcloud logs, host metrics, and APM traces in Datadog for consistent visibility across your entire multi-cloud environment.

## Investigate Kubernetes resources with Datadog MCP tools

DevFeed: [Investigate Kubernetes resources with Datadog MCP tools](<https://devfeed.tech/articles/investigate-kubernetes-resources-with-datadog-mcp-tools-2288.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/kubernetes-mcp-tools/>)

Author: Allie Rittman; Christine Chun Alias Yang

Published: 2026-06-09T00:00:00Z

Content type: tutorial

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Agent](<https://devfeed.tech/topics/ai-agent.md>), [Protocol (disambiguation)](<https://devfeed.tech/topics/protocol.md>), [context window](<https://devfeed.tech/topics/context-window.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [bits-ai](<https://devfeed.tech/tags/bits-ai.md>), [containers](<https://devfeed.tech/tags/containers.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

Datadog MCP Server's Kubernetes toolset gives MCP-compatible AI agents read-only, cross-cluster access to Kubernetes resource context in Datadog. The article explains how its tools support resource searches, manifest inspection, issue investigation, metadata enrichment, scoped structured responses, pagination, and context-window management.

### Source excerpt

Use Datadog MCP Server Kubernetes tools to help AI agents search resources, inspect manifests, and investigate Kubernetes issues.

## Trace network paths from devices to SaaS applications

DevFeed: [Trace network paths from devices to SaaS applications](<https://devfeed.tech/articles/trace-network-paths-from-devices-to-saas-applications-2265.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/devices-to-saas-network-paths/>)

Author: Erica Ho; Natasha Silva

Published: 2026-06-09T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [networking](<https://devfeed.tech/topics/networking.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [devices](<https://devfeed.tech/tags/devices.md>), [digital-experience-monitoring](<https://devfeed.tech/tags/digital-experience-monitoring.md>), [end-user-device-monitoring](<https://devfeed.tech/tags/end-user-device-monitoring.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [latency](<https://devfeed.tech/tags/latency.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [network-device-monitoring](<https://devfeed.tech/tags/network-device-monitoring.md>), [networking](<https://devfeed.tech/tags/networking.md>), [performance](<https://devfeed.tech/tags/performance.md>), [root-cause-analysis](<https://devfeed.tech/tags/root-cause-analysis.md>), [routing](<https://devfeed.tech/tags/routing.md>), [saas](<https://devfeed.tech/tags/saas.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

The article explains how Datadog End User Device Monitoring and Network Path trace network paths from user devices to SaaS applications, exposing per-hop latency and packet loss to locate slowdowns.

### Source excerpt

Diagnose per-hop latency across network hops from user devices to SaaS applications with Datadog End User Device Monitoring and Network Path.

## Monitor Nebius AI Cloud with Datadog

DevFeed: [Monitor Nebius AI Cloud with Datadog](<https://devfeed.tech/articles/monitor-nebius-ai-cloud-with-datadog-2295.md>)

Original publisher: [Read original article](<https://www.datadoghq.com/blog/monitor-nebius-ai-cloud-with-datadog/>)

Author: Ellie Cohen; Eddie Cai

Published: 2026-06-09T00:00:00Z

Content type: article

Language: en

Sources: [Datadog | The Monitor blog](<https://devfeed.tech/sources/datadog-the-monitor-blog.md>)

Topics: [Cloud](<https://devfeed.tech/topics/cloud.md>), [observability](<https://devfeed.tech/topics/observability.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [log management](<https://devfeed.tech/topics/log-management.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Application Performance Management (APM)](<https://devfeed.tech/topics/apm.md>), [agent observability](<https://devfeed.tech/topics/agent-observability.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>)

Tags: [agent-observability](<https://devfeed.tech/tags/agent-observability.md>), [apm](<https://devfeed.tech/tags/apm.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [gpu-monitoring](<https://devfeed.tech/tags/gpu-monitoring.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [llm](<https://devfeed.tech/tags/llm.md>), [log-management](<https://devfeed.tech/tags/log-management.md>), [logs](<https://devfeed.tech/tags/logs.md>), [observability](<https://devfeed.tech/tags/observability.md>), [performance](<https://devfeed.tech/tags/performance.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>)

### AI overview

This article explains how to monitor Nebius AI Cloud workloads with Datadog by centralizing logs, collecting infrastructure metrics and APM traces, tracing LLM applications, and correlating signals across GPU, training, inference, Kubernetes, and other cloud environments.

### Source excerpt

Monitor Nebius AI Cloud workloads with Datadog. Centralize logs, track GPU performance, trace LLM apps, and get unified multi-cloud observability.

## Elastic Observability: A solution for today's "always-on" world

DevFeed: [Elastic Observability: A solution for today's "always-on" world](<https://devfeed.tech/articles/elastic-observability-a-solution-for-today-s-always-on-world-4831.md>)

Original publisher: [Read original article](<https://www.elastic.co/blog/observability-powerful-flexible-efficient>)

Author: Bahubali Shetti

Published: 2023-03-27T00:00:00Z

Content type: article

Language: en

Sources: [Elastic Blog - Elasticsearch, Kibana, and ELK Stack](<https://devfeed.tech/sources/elastic-blog-elasticsearch-kibana-and-elk-stack.md>)

Topics: [telemetry](<https://devfeed.tech/topics/telemetry.md>), [serverless monitoring](<https://devfeed.tech/topics/serverless-monitoring.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [apm](<https://devfeed.tech/tags/apm.md>), [aws](<https://devfeed.tech/tags/aws.md>), [azure](<https://devfeed.tech/tags/azure.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [logs](<https://devfeed.tech/tags/logs.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [machine-learning-cloud-native-infrastructure-monitoring-aiops](<https://devfeed.tech/tags/machine-learning-cloud-native-infrastructure-monitoring-aiops.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [observability](<https://devfeed.tech/tags/observability.md>), [opentelemetry](<https://devfeed.tech/tags/opentelemetry.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

Elastic presents its observability platform as a unified way for SREs to analyze application, infrastructure, user, and business telemetry. It covers metrics, logs, traces, APM, cloud environments, Kubernetes, serverless systems, OpenTelemetry, and machine-learning-assisted insights.

### Source excerpt

Achieve operational excellence and efficiency with the Elastic Observability platform. Our unique approach to unifying observability data on an open and flexible platform with sophisticated AIOps and machine learning will prepare you for the future.

## Sending emails to our half million and growing user community

DevFeed: [Sending emails to our half million and growing user community](<https://devfeed.tech/articles/sending-emails-to-our-half-million-and-growing-user-community-20005.md>)

Original publisher: [Read original article](<http://engineering.hackerearth.com/2016/02/11/sending-emails-to-our-half-million-and-growing-user-community/>)

Published: 2016-02-11T00:00:00Z

Content type: article

Language: en

Sources: [HackerEarth](<https://devfeed.tech/sources/hackerearth.md>)

Topics: [API](<https://devfeed.tech/topics/api.md>), [MongoDB](<https://devfeed.tech/topics/mongodb.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [Django](<https://devfeed.tech/topics/django.md>), [HTML](<https://devfeed.tech/topics/html.md>), [html elements](<https://devfeed.tech/topics/html-elements.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [asynchronous-architecture](<https://devfeed.tech/tags/asynchronous-architecture.md>), [database](<https://devfeed.tech/tags/database.md>), [infrastructure-monitoring](<https://devfeed.tech/tags/infrastructure-monitoring.md>), [mongodb](<https://devfeed.tech/tags/mongodb.md>), [mysql](<https://devfeed.tech/tags/mysql.md>), [queue](<https://devfeed.tech/tags/queue.md>), [rabbitmq](<https://devfeed.tech/tags/rabbitmq.md>), [sendgrid](<https://devfeed.tech/tags/sendgrid.md>)

### AI overview

HackerEarth describes an asynchronous email delivery architecture for sending large volumes of user notifications. Emails are serialized and stored, metadata is placed in RabbitMQ queues, and workers reconstruct and deliver messages through SendGrid, with priority queues used to reduce waiting time.

### Source excerpt

At hackerearth we send emails to keep our users updated on upcoming challenges and their activities, for example, when a user successfully solves a problem, receives test-invitation, updates on user comments. Basically whenever it is appropriate. Architecture It takes lot of computational power to send emails in such large quantities synchronously. So we have implemented an asynchronous architecture to send emails. Here is brief overview of the architecture: Step 1: Construct an email and save the serialized email object in database. Step 2: Queue the metadata for later consumption. Step 3: Consume the metadata, recreate the email object and deliver. The diagram below shows high level architecture of emailing system. The solid line represents the data flow between different components. The dotted line represents the communications. Hackerearth email infrastructure consists of MySQL database, MongoDB database, RabbitMQ queues. Journey Of Email Step 1 - Construct email: There are two different type of emails. Text - Plain text emails Html - Emails with rich interface using html elements. These emails are made using django templates API used by hackerearth developers for sending email - send_email(ctx, template, subject, from_email, html=False, async=True, **kwargs) The above API creates Sendgrid Mail object, serializes and saves it in the db with some additional data. A piece of code similar to the bit shown below is used to create sendgrid Mail object import sendgrid sg = sendgrid.SendGridClient('YOUR_SENDGRID_API_KEY') message = sendgrid.Mail() message.add_to('John Doe <john@email.com>') message.set_subject('Example') message.set_html('Body') message.set_text('Body') message.set_from('Doe John <doe@email.com>') status, msg = sg.send(message) Model below is used for storing the serialized mail object and additional data. class Message(): # The actual data - a pickled sendgrid.Mail object message_data = models.TextField() when_added = models.DateTimeField(default=date