# Algo Hour - Large Scale Data & ML Monitoring with whylogs | Alessya Visnjic

DevFeed: [Algo Hour - Large Scale Data & ML Monitoring with whylogs | Alessya Visnjic](<https://devfeed.tech/articles/algo-hour-large-scale-data-ml-monitoring-with-whylogs-alessya-visnjic-29338.md>)

Original publisher: [Read original article](<https://multithreaded.stitchfix.com/blog/2022/09/29/alessya-algo-hour-announcement/>)

Published: 2022-09-29T09:00:00Z

Content type: article

Language: en

Sources: [Stitch Fix](<https://devfeed.tech/sources/stitch-fix.md>)

Topics: [Data Quality](<https://devfeed.tech/topics/data-quality.md>), [data observability](<https://devfeed.tech/topics/data-observability.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [data-ml](<https://devfeed.tech/tags/data-ml.md>), [data-pipelines](<https://devfeed.tech/tags/data-pipelines.md>), [data-quality](<https://devfeed.tech/tags/data-quality.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [observability](<https://devfeed.tech/tags/observability.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

## AI overview

This talk explains how the open-source whylogs library supports end-to-end data quality and monitoring across machine learning pipelines. It covers whylogs' lightweight statistical data collection, language- and platform-agnostic approach, architecture, and application to existing data and ML pipelines.

## Source excerpt

Title: Large Scale Data & ML Monitoring with whylogs Talk Abstract: In the era of microservices, decentralized ML architectures and complex data pipelines, data quality has become a bigger challenge than ever. When data is involved in complex business processes and decisions, bad data can, and will, affect the bottom line. As a result, ensuring data quality across the entire ML pipeline is both costly, and cumbersome while data monitoring is often fragmented and performed ad hoc. An open source library called whylogs is built to address these challenges. It is a lightweight data profiling library that enables end-to-end data monitoring across the entire software stack. The library implements a language and platform agnostic approach to data quality and data monitoring. It's been deployed at massive-scale data environments, on structured and unstructured data modalities, and across a range of points in the ML lifecycle. In this talk, we will provide an overview of the whylogs architecture, including its lightweight statistical data collection approach and we will show how users can apply this library to existing data and ML pipelines. Date and Time: The talk will be held on Tuesday, October 11th at 1:00PM PDT. Recording Info: This talk was recorded live and is viewable below: Speaker Info: Alessya Visnjic is the CEO of WhyLabs, the AI Observability company building tools that power robust and responsible AI deployment. Prior to WhyLabs, Alessya was a CTO-in-residence at the Allen Institute for AI, where she evaluated commercial potential for the latest AI research. Earlier, Alessya spent 9 years at Amazon leading ML initiatives, including forecasting and data science platforms. Alessya is also the founder of Rsqrd AI, a global community of 1,000+ AI practitioners who are committed to making enterprise AI technology responsible.