# How to Reliably Scale Your Data Platform for High Volumes

DevFeed: [How to Reliably Scale Your Data Platform for High Volumes](<https://devfeed.tech/articles/how-to-reliably-scale-your-data-platform-for-high-volumes-1546.md>)

Original publisher: [Read original article](<https://shopify.engineering/reliably-scale-data-platform>)

Author: Arbab Ahmed

Published: 2020-12-08T17:30:29Z

Content type: article

Language: en

Sources: [Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering.md>), [Shopify Engineering - Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering-shopify-engineering.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Platform Engineering](<https://devfeed.tech/topics/platform-engineering.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [MySQL](<https://devfeed.tech/topics/mysql.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [apache](<https://devfeed.tech/tags/apache.md>), [apache-parquet](<https://devfeed.tech/tags/apache-parquet.md>), [apache-spark](<https://devfeed.tech/tags/apache-spark.md>), [data](<https://devfeed.tech/tags/data.md>), [data-platform-engineering](<https://devfeed.tech/tags/data-platform-engineering.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [insights](<https://devfeed.tech/tags/insights.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [platform](<https://devfeed.tech/tags/platform.md>), [scale](<https://devfeed.tech/tags/scale.md>), [spark](<https://devfeed.tech/tags/spark.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [systems](<https://devfeed.tech/tags/systems.md>)

## AI overview

Shopify's Data Platform Engineering team describes how it prepared the data platform to handle the high-volume Black Friday and Cyber Monday event. The platform experienced an average throughput increase of 150 percent and processes data through ingestion, batch or stream processing, and delivery to merchants, partners, and internal teams. The article covers the use of Apache Parquet, Apache Spark, dbt, MySQL, Kafka, and tiered services to prioritize reliability and infrastructure investment.

## Source excerpt

In this post, we'll outline the approach we took to reliably scale our data platform in preparation for Black Friday and Cyber Monday.