# model-monitoring

Published articles for model-monitoring.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## From Data to Insight: Helpshift's Journey with ML Observability

DevFeed: [From Data to Insight: Helpshift's Journey with ML Observability](<https://devfeed.tech/articles/from-data-to-insight-helpshift-s-journey-with-ml-observability-30515.md>)

Original publisher: [Read original article](<https://medium.com/helpshift-engineering/from-data-to-insight-helpshifts-journey-with-ml-observability-9680e27d1d01?source=rss----3229f31ca4f4---4>)

Author: Sujit Singh

Published: 2025-11-26T14:00:15Z

Content type: article

Language: en

Sources: [Helpshift](<https://devfeed.tech/sources/helpshift.md>)

Topics: [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [observability](<https://devfeed.tech/topics/observability.md>), [telemetry](<https://devfeed.tech/topics/telemetry.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [generalization in machine learning](<https://devfeed.tech/topics/generalization-in-machine-learning.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [inference](<https://devfeed.tech/tags/inference.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [model-monitoring](<https://devfeed.tech/tags/model-monitoring.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>)

### AI overview

Helpshift describes its journey toward building a custom machine learning observability solution. The article explains ML observability, outlines system, inference, and model monitoring, discusses limitations of existing tools, and introduces an approach based on "Wide Events."

### Source excerpt

Introduction In an age where artificial intelligence (AI) and machine learning (ML) are integral to almost every aspect of our lives, ensuring the effectiveness, fairness, and reliability of ML models is paramount. Observability plays a crucial role in maintaining the performance of these models, allowing us to detect and resolve issues promptly. At Helpshift, we recognized the need for robust ML observability to keep our models running smoothly and efficiently. This blog post explores our journey in building a custom ML observability solution tailored to our specific needs. We'll delve into the concept of ML observability, discuss the limitations of existing tools, and share how we implemented our own solution based on the idea of "Wide Events." Understanding ML Observability ML observability is the ability to monitor and understand the performance, behavior, and outputs of machine learning models in real-time. It enables us to proactively identify potential issues and anomalies, facilitating timely interventions and mitigating risks. ML observability encompasses several key components: System Monitoring: Tracking the performance of the infrastructure where ML services are deployed, including metrics like CPU and memory usage, network traffic, disk space, and service performance. Inference Monitoring: Evaluating and auditing the real-time performance of deployed ML models in production by tracking incoming requests and the accuracy of model predictions. Model Monitoring: Observing the long-term accuracy of ML models by monitoring key metrics such as accuracy, precision, recall, and F1-score, and detecting any drift over time. Why Observability Matters If you've ever played Age of Empires, you know how crucial it is to explore the map to manage resources proactively and strategize effectively. Similarly, ML observability is about exploring properties and patterns not determined in advance. It allows us to be proactive in debugging and improving our systems, ensuring

## Implementing Model Monitoring on Databricks

DevFeed: [Implementing Model Monitoring on Databricks](<https://devfeed.tech/articles/implementing-model-monitoring-on-databricks-28604.md>)

Original publisher: [Read original article](<https://www.marvelousmlops.io/p/lecture-10-implementing-model-monitoring>)

Author: Başak Tuğçe Eskili

Published: 2025-08-06T16:55:23Z

Content type: tutorial

Language: en

Sources: [MarvelousMLOps](<https://devfeed.tech/sources/marvelousmlops.md>)

Topics: [databricks](<https://devfeed.tech/topics/databricks.md>), [MLOps](<https://devfeed.tech/topics/mlops.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [model-serving](<https://devfeed.tech/topics/model-serving.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Logging](<https://devfeed.tech/topics/logging.md>), [Data Quality](<https://devfeed.tech/topics/data-quality.md>)

Tags: [databricks](<https://devfeed.tech/tags/databricks.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [model-monitoring](<https://devfeed.tech/tags/model-monitoring.md>), [model-serving](<https://devfeed.tech/tags/model-serving.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [request](<https://devfeed.tech/tags/request.md>), [table](<https://devfeed.tech/tags/table.md>)

### AI overview

Lecture 10 in an MLOps with Databricks course demonstrates model monitoring using inference tables and Lakehouse Monitoring. It covers collecting inference logs, creating a structured monitoring table, scheduling refreshes, and building a dashboard to visualize metrics and detect drift.

### Source excerpt

Lecture 10 of MLOps with Databricks course

## Glassdoor Decreases Latency Overhead and Improves Data Monitoring with WhyLabs

DevFeed: [Glassdoor Decreases Latency Overhead and Improves Data Monitoring with WhyLabs](<https://devfeed.tech/articles/glassdoor-decreases-latency-overhead-and-improves-data-monitoring-with-whylabs-22609.md>)

Original publisher: [Read original article](<https://medium.com/glassdoor-engineering/glassdoor-decreases-latency-overhead-and-improves-data-monitoring-with-whylabs-ad399576624d?source=rss----288d984af747---4>)

Author: Lanqi Fei

Published: 2023-09-06T21:57:52Z

Content type: article

Language: en

Sources: [Glassdoor Engineering](<https://devfeed.tech/sources/glassdoor-engineering.md>)

Topics: [Latency](<https://devfeed.tech/topics/latency.md>), [Logging](<https://devfeed.tech/topics/logging.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Library](<https://devfeed.tech/topics/library.md>), [async](<https://devfeed.tech/topics/async.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [async](<https://devfeed.tech/tags/async.md>), [complexity](<https://devfeed.tech/tags/complexity.md>), [graph](<https://devfeed.tech/tags/graph.md>), [latency](<https://devfeed.tech/tags/latency.md>), [latency-optimization](<https://devfeed.tech/tags/latency-optimization.md>), [library](<https://devfeed.tech/tags/library.md>), [logging](<https://devfeed.tech/tags/logging.md>), [model-monitoring](<https://devfeed.tech/tags/model-monitoring.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

This article examines how Glassdoor and WhyLabs addressed latency when integrating data monitoring into a real-time service. It describes changes to whylogs, an open-source data logging library, and compares architectural options including asynchronous calls, DAG-based restructuring, and keeping work in a linear execution path.

### Source excerpt

Authors: Lanqi Fei, Jamie, Natalia This blog was written by Lanqi Fei, Senior ML Scientist at Glassdoor, Jamie Broomall, Senior Software Engineer at WhyLabs, and Natalia Skaczkowska-Drabczyk, Customer Success Data Scientist at WhyLabs. The challenge of integration latency Consider the scenario where we want to integrate a new tool into an existing service that potentially operates in real-time and involves some user interface. We need to make sure that the latency of the service in production is acceptable after the integration, while still keeping the overall maintenance costs low. In this scenario, there are trade-offs to be made and the right choice will depend on the individual characteristics of the service and the newly integrated function. Simplifying this function is a common path to gaining a significant advantage in this optimization game. This blog, written in collaboration between Glassdoor and WhyLabs, describes a real-world instance of an integration latency challenge and gives a detailed walk-through of the changes applied within whylogs (an open-source data logging library maintained by WhyLabs) to mitigate it. What are the best options for reducing latency? There are a couple of options for reducing latency when integrating a new function into an existing service. Restructuring your service or architecture to allow an early response to the caller before doing the additional work (this may be as simple as using an async call pattern with a log statement or as complex as a DAG framework). You can think of your service as a graph -- its nodes should be the latency-critical tasks and ideally those should be executed, instrumented and tested independently. Using a DAG can be a good way of scaling out a service to a large number of new features and integrations while maintaining latency requirements and managing the complexity of the critical path to generating a high quality user response. The downside of this approach is the additional complexity as well