# warehouse

Published articles for warehouse.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## MYWAI ports its VILMA visual imitation learning toolkit to Arduino UNO Q and VENTUNO Q

DevFeed: [MYWAI ports its VILMA visual imitation learning toolkit to Arduino UNO Q and VENTUNO Q](<https://devfeed.tech/articles/mywaitm-vilmatm-is-designed-to-bring-human-like-learning-to-robots-via-one-shot-demonstration-26776.md>)

Original publisher: [Read original article](<https://blog.arduino.cc/2026/09/15/mywai-vilma-is-designed-to-bring-human-like-learning-to-robots-via-one-shot-demonstration/>)

Author: Arduino Team

Published: 2026-09-15T14:26:17Z

Content type: article

Language: en

Sources: [Arduino Blog](<https://devfeed.tech/sources/arduino-blog.md>)

Topics: [Robotics](<https://devfeed.tech/topics/robotics.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Computer vision](<https://devfeed.tech/topics/computer-vision.md>), [Arduino](<https://devfeed.tech/topics/arduino.md>), [Qualcomm](<https://devfeed.tech/topics/qualcomm.md>), [UNO Q](<https://devfeed.tech/topics/uno-q.md>), [VENTUNO Q](<https://devfeed.tech/topics/ventuno-q.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [arduino](<https://devfeed.tech/tags/arduino.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [edge-ai](<https://devfeed.tech/tags/edge-ai.md>), [industrial](<https://devfeed.tech/tags/industrial.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [qualcomm](<https://devfeed.tech/tags/qualcomm.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [robots](<https://devfeed.tech/tags/robots.md>), [uno-q](<https://devfeed.tech/tags/uno-q.md>), [ventuno-q](<https://devfeed.tech/tags/ventuno-q.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article describes VILMA, an AI-powered toolkit from MYWAI that enables robots and humanoids to learn manipulation tasks from one-shot human demonstrations. It reports that the toolkit is being ported to Arduino UNO Q and VENTUNO Q boards powered by Qualcomm Dragonwing processors.

### Source excerpt

Every day, hundreds of thousands of kits are prepared in warehouses before components ever reach an automotive production line. While robots have become commonplace in modern manufacturing, many upstream logistics activities still rely heavily on human operators performing repetitive pick-and-place and kitting tasks. What if robots could learn these operations the same way humans do: [...] The post MYWAI™ VILMA™ is designed to bring human-like learning to robots via one-shot demonstration appeared first on Arduino Blog.

## The hard part of an MCP gateway is auth

DevFeed: [The hard part of an MCP gateway is auth](<https://devfeed.tech/articles/the-hard-part-of-an-mcp-gateway-is-auth-16029.md>)

Original publisher: [Read original article](<https://workos.com/blog/mcp-gateway-hard-part-is-auth>)

Author: WorkOS

Published: 2026-09-11T15:22:28Z

Content type: opinion

Language: en

Sources: [WorkOS Blog](<https://devfeed.tech/sources/workos-blog.md>)

Topics: [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [Authorization](<https://devfeed.tech/topics/authorization.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [audit](<https://devfeed.tech/tags/audit.md>), [auth](<https://devfeed.tech/tags/auth.md>), [github](<https://devfeed.tech/tags/github.md>), [integration](<https://devfeed.tech/tags/integration.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-gateway](<https://devfeed.tech/tags/mcp-gateway.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [salesforce](<https://devfeed.tech/tags/salesforce.md>), [scopes](<https://devfeed.tech/tags/scopes.md>), [slack](<https://devfeed.tech/tags/slack.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article examines Sierra's internal MCP gateway and argues that its most difficult challenges are authentication, identity, per-tool scopes, consent, and audit rather than MCP protocol design. It also discusses integration work and the lack of standardized gateway behavior.

### Source excerpt

Sierra's MCP gateway iceberg is a field report on agent auth: identity, per-tool scopes, consent, and audit are the submerged mass, and they ship off the shelf.

## Your Data Warehouse Isn't Integrated Just Because the Tables Are in One Place

DevFeed: [Your Data Warehouse Isn't Integrated Just Because the Tables Are in One Place](<https://devfeed.tech/articles/your-data-warehouse-isn-t-integrated-just-because-the-tables-are-in-one-place-37156.md>)

Original publisher: [Read original article](<https://seattledataguy.substack.com/p/your-data-warehouse-isnt-integrated>)

Author: SeattleDataGuy

Published: 2026-07-18T23:26:52Z

Content type: opinion

Language: en

Sources: [SeattleDataGuy's Newsletter](<https://devfeed.tech/sources/seattledataguy-s-newsletter.md>)

Topics: [centralization](<https://devfeed.tech/topics/centralization.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [centralization](<https://devfeed.tech/tags/centralization.md>), [data](<https://devfeed.tech/tags/data.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article argues that placing tables in one data warehouse centralizes data but does not by itself integrate it.

### Source excerpt

Centralization is not integration!

## How We Refresh Razorpay's Data Warehouse 10x Faster with Graphs and Indexes

DevFeed: [How We Refresh Razorpay's Data Warehouse 10x Faster with Graphs and Indexes](<https://devfeed.tech/articles/how-we-refresh-razorpay-s-data-warehouse-10x-faster-with-graphs-and-indexes-24040.md>)

Original publisher: [Read original article](<https://engineering.razorpay.com/how-we-refresh-razorpays-data-warehouse-10x-faster-with-graphs-and-indexes-538abc244703?source=rss----6407ad2e59af---4>)

Author: Amit Prabhu

Published: 2026-07-14T14:06:16Z

Content type: article

Language: en

Sources: [Razorpay Engineering - Medium](<https://devfeed.tech/sources/razorpay-engineering-medium.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [microservices architecture](<https://devfeed.tech/topics/microservices-architecture.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [Architecture & Design](<https://devfeed.tech/topics/architecture-design.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [batch](<https://devfeed.tech/tags/batch.md>), [data](<https://devfeed.tech/tags/data.md>), [graphs](<https://devfeed.tech/tags/graphs.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [razorpay](<https://devfeed.tech/tags/razorpay.md>), [spark](<https://devfeed.tech/tags/spark.md>), [trino](<https://devfeed.tech/tags/trino.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

Razorpay describes its data warehouse refresh pipeline, which builds wide denormalized Facts by joining data from multiple microservices. The article covers the original Airflow- and Spark-based full-refresh process, the underlying lake formats and query layer, and the scaling challenges that led the team to reconsider refresh strategy, data layout, and high-cardinality dimensions.

### Source excerpt

Contributors: Utkarsh Koppikar Rohan Background Razorpay provides the payment infrastructure for millions of merchants globally. Behind every payment, settlement, and refund is a microservices architecture where each service owns its own database. While this keeps services independent and scalable, it creates a challenge for stakeholders who need to see across those boundaries. The Data Platform team manages the infrastructure that bridges this gap. Transactional data flows into the lake via CDC pipelines, ingested onto S3 in Delta Lake, Apache Iceberg, or plain Parquet formats. On top of the lake, we build domain-specific warehouse tables -- wide, pre-joined tables that co-locate all the data a consumer needs, queryable via Trino. These power two use cases: Analytics (internal dashboards on Tableau and Superset) and Reporting (merchants and regulated entities who download structured data exports; Razorpay generates nearly a million such reports per month). The warehouse tables that power both use cases are called Facts. A Fact is a flat denormalised table on S3, produced by joining 10 to 30 microservice tables and materialising the result once. A settlement Fact, for example, merges payments, refunds, adjustments, and card details into a single wide row so that a dashboard or report reads from a single table instead of joining across services in real time. It is closer to a domain-specific materialised view than a classical data warehouse fact table. We maintain over 50 such Facts, and approximately 40% of all merchant reports are served directly from them. As data volumes and the number of entities per fact grew, the batch generation pipeline began to show its limits, prompting us to rethink the refresh strategy, the data layout, and how to handle high-cardinality dimensions. The rest of this post covers that journey. The Full Refresh Pipeline: Our Baseline and the Pain The original full-refresh pipeline was straightforward. Schedule: Airflow schedules Spark jobs o

## How the 5 major cloud data warehouses compare on cost-performance

DevFeed: [How the 5 major cloud data warehouses compare on cost-performance](<https://devfeed.tech/articles/how-the-5-major-cloud-data-warehouses-compare-on-cost-performance-5209.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/cloud-data-warehouses-cost-performance-comparison>)

Author: Tom Schreiber; Lionel Palacin

Published: 2025-12-02T00:00:00Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [databricks](<https://devfeed.tech/topics/databricks.md>), [dataset](<https://devfeed.tech/topics/dataset.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-data](<https://devfeed.tech/tags/cloud-data.md>), [compare](<https://devfeed.tech/tags/compare.md>), [comparison](<https://devfeed.tech/tags/comparison.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data](<https://devfeed.tech/tags/data.md>), [databricks](<https://devfeed.tech/tags/databricks.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [performance](<https://devfeed.tech/tags/performance.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [storage](<https://devfeed.tech/tags/storage.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

This article compares the cost-performance of Snowflake, Databricks, ClickHouse Cloud, BigQuery, and Redshift across analytical workloads containing 1 billion, 10 billion, and 100 billion rows. Using each system's real compute billing model, the benchmark concludes that ClickHouse Cloud provides substantially better value than the other systems at scale.

### Source excerpt

We benchmarked the five major cloud data warehouses at 1B-100B rows using their real billing models to measure performance per dollar. Results show how cost-performance shifts as data grows.

## Building an Enterprise Data Warehouse on Heroku: From Complex ETL to Seamless Salesforce Integration

DevFeed: [Building an Enterprise Data Warehouse on Heroku: From Complex ETL to Seamless Salesforce Integration](<https://devfeed.tech/articles/building-an-enterprise-data-warehouse-on-heroku-from-complex-etl-to-seamless-salesforce-integration-26383.md>)

Original publisher: [Read original article](<https://www.heroku.com/blog/building-an-enterprise-data-warehouse-on-heroku/>)

Author: Sudarshan Hiray

Published: 2025-11-05T20:05:38Z

Content type: article

Language: en

Sources: [Heroku](<https://devfeed.tech/sources/heroku.md>)

Topics: [Heroku](<https://devfeed.tech/topics/heroku.md>), [data-architecture](<https://devfeed.tech/topics/data-architecture.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [applications](<https://devfeed.tech/tags/applications.md>), [article](<https://devfeed.tech/tags/article.md>), [building](<https://devfeed.tech/tags/building.md>), [data](<https://devfeed.tech/tags/data.md>), [ecosystems](<https://devfeed.tech/tags/ecosystems.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [etl](<https://devfeed.tech/tags/etl.md>), [heroku](<https://devfeed.tech/tags/heroku.md>), [heroku-postgres](<https://devfeed.tech/tags/heroku-postgres.md>), [integration](<https://devfeed.tech/tags/integration.md>), [platforms](<https://devfeed.tech/tags/platforms.md>), [salesforce](<https://devfeed.tech/tags/salesforce.md>), [services](<https://devfeed.tech/tags/services.md>), [systems](<https://devfeed.tech/tags/systems.md>), [tools](<https://devfeed.tech/tags/tools.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

Heroku describes an enterprise data warehouse architecture that unifies Salesforce and data from multiple application databases for real-time analytics. The approach uses Heroku Connect, Heroku Postgres follower databases, and a unified analytics platform to address API limits, ETL complexity, production database load, and fragmented monitoring.

### Source excerpt

Modern businesses don't just run on Salesforce--they run on entire ecosystems of applications. At Heroku, we operate dozens of services alongside our Salesforce instance such as billing systems, user management platforms, analytics engines, and support tools. Traditional approaches to unifying this data create more problems than they solve. In this article, we'll see how we [...] The post Building an Enterprise Data Warehouse on Heroku: From Complex ETL to Seamless Salesforce Integration appeared first on Heroku.

## A Date-Stamped Data Warehouse Setup

DevFeed: [A Date-Stamped Data Warehouse Setup](<https://devfeed.tech/articles/the-data-warehouse-setup-no-one-taught-you-27258.md>)

Original publisher: [Read original article](<https://blog.dataexpert.io/p/the-data-warehouse-setup-no-one-taught>)

Author: Sahar Massachi

Published: 2025-10-24T21:03:30Z

Content type: article

Language: en

Sources: [DataExpert.io Newsletter](<https://devfeed.tech/sources/dataexpert-io-newsletter.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [data](<https://devfeed.tech/topics/data.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>)

Tags: [ab-testing](<https://devfeed.tech/tags/ab-testing.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [storage](<https://devfeed.tech/tags/storage.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article presents date-stamping data as a simple organizing principle for building resilient data warehouses and pipelines. It explains how this approach can manage changes over time, support experimentation and metrics, and work with systems including Hive metastore, Iceberg, Delta, and Hudi.

### Source excerpt

Storage is cheap, your time is not!

## Postgres Meets Analytics: CDC From Neon to ClickHouse Via PeerDB

DevFeed: [Postgres Meets Analytics: CDC From Neon to ClickHouse Via PeerDB](<https://devfeed.tech/articles/postgres-meets-analytics-cdc-from-neon-to-clickhouse-via-peerdb-5735.md>)

Original publisher: [Read original article](<https://neon.com/blog/postgres-meets-analytics-cdc-from-neon-to-clickhouse-via-peerdb>)

Author: Sai Srirampur

Published: 2024-10-02T16:39:48Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Data analysis](<https://devfeed.tech/topics/data-analysis.md>), [Replication](<https://devfeed.tech/topics/replication.md>), [data](<https://devfeed.tech/topics/data.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [web applications](<https://devfeed.tech/topics/web-applications.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [code](<https://devfeed.tech/tags/code.md>), [community](<https://devfeed.tech/tags/community.md>), [data](<https://devfeed.tech/tags/data.md>), [databases](<https://devfeed.tech/tags/databases.md>), [latency](<https://devfeed.tech/tags/latency.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [olap](<https://devfeed.tech/tags/olap.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [postgres](<https://devfeed.tech/tags/postgres.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [replication](<https://devfeed.tech/tags/replication.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>), [web-applications](<https://devfeed.tech/tags/web-applications.md>)

### AI overview

This article explains how Neon serverless Postgres and ClickHouse can be combined for transactional workloads and real-time analytics. It presents PeerDB's Change Data Capture and replication capabilities as the bridge for synchronizing Neon data to ClickHouse for customer-facing analytics and data warehousing.

### Source excerpt

If you're building a data-driven application that handles large amounts of data, you may need to balance two different types of databases: a purpose-built operational DB for your transactions, and an analytical database for large-scale data analysis. A purpose-built analytical da...

## Building a Data Pipeline to Track Strava's Bad Events

DevFeed: [Building a Data Pipeline to Track Strava's Bad Events](<https://devfeed.tech/articles/an-eventful-summer-at-strava-26571.md>)

Original publisher: [Read original article](<https://medium.com/strava-engineering/an-eventful-summer-at-strava-5692882e5f4f?source=rss----89d4108ce2a3---4>)

Author: Bisman Sodhi

Published: 2024-01-08T20:19:46Z

Content type: opinion

Language: en

Sources: [Strava Engineering](<https://devfeed.tech/sources/strava-engineering.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [airflow](<https://devfeed.tech/topics/airflow.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Python](<https://devfeed.tech/topics/python.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [JSON](<https://devfeed.tech/topics/json.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Computer science](<https://devfeed.tech/topics/computer-science.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [apache-airflow](<https://devfeed.tech/tags/apache-airflow.md>), [aws](<https://devfeed.tech/tags/aws.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [integrity](<https://devfeed.tech/tags/integrity.md>), [json](<https://devfeed.tech/tags/json.md>), [python](<https://devfeed.tech/tags/python.md>), [s3](<https://devfeed.tech/tags/s3.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [strava](<https://devfeed.tech/tags/strava.md>), [tableau](<https://devfeed.tech/tags/tableau.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

A software engineering intern describes building a daily Apache Airflow pipeline that extracts schema-invalid user behavior events from S3, decompresses them into JSON, and loads them into Snowflake. Staging tables protect production data from partial loads, while materialized SQL views and a Tableau dashboard improve querying and monitoring.

### Source excerpt

Hi my name is Bisman and I studied Computer Science at University of California, Santa Barbara. During summer of 2022, I had the most amazing experience working as a Software Engineer Intern on Strava's Data Platform Team. In the first fews weeks, I learned the tools my team uses and then spent the rest of the time working on my project. TRACKING BAD EVENTS For my major summer project, I created a data pipeline that pulls user behavior data out of external storage and persists it in our data warehouse. Strava uses a service called Snowplow to collect this user behavior data, like loading a club page or uploading a profile photo. Sometimes, this data fails to match the schema that we've set, and a piece of data that fails this schema validation is called a bad event. Previously, these bad events were temporarily stored in an Elastic Search. Persisting this data in Snowflake, our data warehouse, makes it accessible to a wider audience. It also makes it easier to incorporate the bad events data with other services used at Strava. To start my project, I created a directed acyclic graph in Apache Airflow, a scheduling framework, using python that extracts bad events data from the S3, AWS's storage service, buckets on a daily cadence. This data was stored as gzip files on S3 which I decompressed and stored the data as JSON blobs. As I was working with billions of rows of data, it was important to maintain data integrity and take measures in case data failed to load from S3. Therefore, I loaded data into a staging table in Snowflake. The staging table ensured that if loading from S3 failed, the production table would remain untouched. This data was then loaded into the production table free of any partial data. After all the data was loaded into the production table, I created six view tables because there were six different types of bad events stored in the production table. I collaborated with our stakeholders -- data analysts -- throughout this process to craft tables bas

## Build a real-time dashboard over BigQuery

DevFeed: [Build a real-time dashboard over BigQuery](<https://devfeed.tech/articles/build-a-real-time-dashboard-over-bigquery-18395.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/bigquery-real-time-dashboard>)

Author: Cameron Archer

Published: 2023-10-20T00:00:00Z

Content type: tutorial

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [bigquery](<https://devfeed.tech/tags/bigquery.md>), [build](<https://devfeed.tech/tags/build.md>), [data](<https://devfeed.tech/tags/data.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [i-built-this](<https://devfeed.tech/tags/i-built-this.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

A tutorial about building a real-time dashboard over BigQuery, with sub-second updates to warehouse data.

### Source excerpt

BigQuery real-time dashboard sounds impossible. It's not. Here's how to add sub-second updates to your warehouse data.

## How the Guardian's Ophan Analytics Tool Moved from Elasticsearch Rollups to BigQuery

DevFeed: [How the Guardian's Ophan Analytics Tool Moved from Elasticsearch Rollups to BigQuery](<https://devfeed.tech/articles/roll-over-rollups-the-big-future-of-ophan-s-historical-data-19962.md>)

Original publisher: [Read original article](<https://www.theguardian.com/info/2023/jun/07/roll-over-rollups-the-big-future-of-ophans-historical-data>)

Author: Sam Hession

Published: 2023-06-07T12:44:33Z

Content type: article

Language: en

Sources: [Guardian](<https://devfeed.tech/sources/guardian.md>)

Topics: [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [elasticsearch](<https://devfeed.tech/topics/elasticsearch.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [data](<https://devfeed.tech/tags/data.md>), [elasticsearch](<https://devfeed.tech/tags/elasticsearch.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The Guardian describes how Ophan, its real-time analytics tool, expanded from short-term pageview monitoring to historical data analysis. Because Elasticsearch Rollups remained in technical preview and lacked service-level guarantees, the team considered moving the long-term data pipeline to BigQuery.

### Source excerpt

How the Guardian's real time analytics tool pivoted from ElasticSearch Rollups to BigQuery and what we learnt along the way Ophan is the Guardian's in-house developed real time analytics tool which allows us to see how our content is performing in real-time, providing our Editorial teams with the insights they need to curate and promote our journalism. Its intuitive ways of monitoring reader engagement with content help the teams to evaluate whether it's performing optimally. Ophan's distinguishing feature has always been to provide journalists with visualised real time pageview information, with features based on the previous two weeks. More recently we have expanded these capabilities to provide insights for wider timescales with years of data available at their fingertips. Continue reading...

## Adding Speed to Snowflake Warehouses for Real-Time Solutions

DevFeed: [Adding Speed to Snowflake Warehouses for Real-Time Solutions](<https://devfeed.tech/articles/building-real-time-solutions-with-snowflake-at-a-fraction-of-cost-18633.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/real-time-solutions-with-snowflake>)

Author: Alejandro Martín

Published: 2023-04-19T00:00:00Z

Content type: article

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [real-time](<https://devfeed.tech/topics/real-time.md>)

Tags: [product-updates](<https://devfeed.tech/tags/product-updates.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [solutions](<https://devfeed.tech/tags/solutions.md>), [speed](<https://devfeed.tech/tags/speed.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article discusses workarounds for adding speed to Snowflake warehouses for real-time solutions without replacing the warehouse.

### Source excerpt

Real-time solutions with Snowflake require workarounds. Here's how to add speed to your warehouse without replacing it.

## Using Trino with Apache Airflow for (almost) all your data problems

DevFeed: [Using Trino with Apache Airflow for (almost) all your data problems](<https://devfeed.tech/articles/using-trino-with-apache-airflow-for-almost-all-your-data-problems-8706.md>)

Original publisher: [Read original article](<https://trino.io/blog/2022/12/21/trino-summit-2022-astronomer-recap.html>)

Author: Philippe Gagnon, Brian Olsen

Published: 2022-12-21T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [airflow](<https://devfeed.tech/topics/airflow.md>), [data](<https://devfeed.tech/topics/data.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>), [Availability](<https://devfeed.tech/topics/availability.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [Software as a service](<https://devfeed.tech/topics/saas.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [apache](<https://devfeed.tech/tags/apache.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [availability](<https://devfeed.tech/tags/availability.md>), [batch](<https://devfeed.tech/tags/batch.md>), [data](<https://devfeed.tech/tags/data.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [databases](<https://devfeed.tech/tags/databases.md>), [integration](<https://devfeed.tech/tags/integration.md>), [saas](<https://devfeed.tech/tags/saas.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [summit](<https://devfeed.tech/tags/summit.md>), [trading](<https://devfeed.tech/tags/trading.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

The article recaps a Trino Summit 2022 talk about using Apache Airflow to orchestrate Trino batch queries. It explains Trino's fault-tolerant execution mode and discusses moving queries closer to federated data sources to improve data availability, performance, scalability, and collaboration.

### Source excerpt

As we close in on the final talks from Trino Summit 2022, this next talk dives into how to set up Trino for batch processing. Trino has historically been well-known for facilitating fast adhoc analytics queries as opposed to long-running, resource intensive batch/ETL queries. This is due to the fact that Trino kills queries that run out of resources in order to prioritize faster query execution. Earlier this year, Trino added features to better support batch queries with a new fault-tolerant execution mode. This mode backs up intermediate data during execution time, allowing Trino to restart individual query tasks on failure rather than a query stage or the query itself. Batch queries don't typically involve human intervention and run asynchronously. These tasks may depend on each other and have a complex workflow. This talk describes how to orchestrate this complexity using Airflow's new Trino integration to run Trino batch queries to solve (almost) all your data problems.

## Understanding the Data Warehouse

DevFeed: [Understanding the Data Warehouse](<https://devfeed.tech/articles/understanding-the-data-warehouse-18770.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/why-data-warehouses>)

Author: Alasdair Brown

Published: 2022-11-29T00:00:00Z

Content type: article

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>)

Tags: [blog](<https://devfeed.tech/tags/blog.md>), [data](<https://devfeed.tech/tags/data.md>), [series](<https://devfeed.tech/tags/series.md>), [the-data-base](<https://devfeed.tech/tags/the-data-base.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

This article introduces the origins of data warehouses and explains what they enable. It is presented as the first article in a two-part series.

### Source excerpt

What gave rise to the Data Warehouse? What do they enable? I answer these questions in this first of a two-blog series.

## Let me automate that for you II, Electric Bugaloo

DevFeed: [Let me automate that for you II, Electric Bugaloo](<https://devfeed.tech/articles/let-me-automate-that-for-you-ii-electric-bugaloo-19838.md>)

Original publisher: [Read original article](<https://tech.gc.com/let-me-automate-that-for-you-ii-electric-bugaloo/>)

Author: GameChanger

Published: 2021-05-07T18:29:38Z

Content type: article

Language: en

Sources: [GameChanger](<https://devfeed.tech/sources/gamechanger.md>)

Topics: [SQL](<https://devfeed.tech/topics/sql.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [Pull Request](<https://devfeed.tech/topics/pull-request.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Structured-data](<https://devfeed.tech/topics/structured-data.md>), [Script](<https://devfeed.tech/topics/script.md>), [Slack](<https://devfeed.tech/topics/slack.md>), [Documentation](<https://devfeed.tech/topics/documentation.md>)

Tags: [automation](<https://devfeed.tech/tags/automation.md>), [blog](<https://devfeed.tech/tags/blog.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [building](<https://devfeed.tech/tags/building.md>), [data](<https://devfeed.tech/tags/data.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [github](<https://devfeed.tech/tags/github.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [reactive](<https://devfeed.tech/tags/reactive.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [sql](<https://devfeed.tech/tags/sql.md>), [systems](<https://devfeed.tech/tags/systems.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article describes evolving an embedded SQL generator and related scripts into a standalone SQL producer for warehouse schema and table management. The system generates SQL migrations, creates or updates tables, optimizes table performance, documents proposed changes, opens pull requests, and notifies engineers in Slack for review. The article also discusses limitations of the reactive original implementation and outlines the design of the improved system.

### Source excerpt

Improving our original, embedded SQL generator and some related scripts by converting them to a better, long term, stand alone SQL producer that's faster, more reliable, and more obvious. About seventeen years ago, in 2019, I published my blog post "Let me automate that for you" about a design for automating creating warehouse tables based on schemas for new event data. The idea was when our ETL system couldn't load waiting data into a warehouse table (as there was no table to be found), it would look up the schema for that data, convert the schema to a SQL statement, then issue a PR to the repo where SQL migrations for such needs are kept. Eventually creating tables made a friend, updating tables when there was a mismatch between the schema of the data we were loading and the schema of the table in the warehouse, and a third buddy joined the part, optimizing a table to improve its performance. The system had some absolutely great qualities: it automated acting on errors it saw, it generated great documentation in the PR and the SQL statement (with comments for discussions and places to review more closely), and it posted to Slack to let engineers know that there was something for them to do a final review on. However... it wasn't perfect. Reading is going toward something that is about to be, and no one yet knows what it will be. [1] Let me take you through the evolution of our embedded SQL generator to stand-alone SQL producer. Limitations of previous implementation Opportunities to build it better Building blocks of a stand-alone SQL producer Joining the human needs with the computer's logic Detailed breakdown of the services available Troubleshooting live Final thoughts Appendix A: Redshift optimization queries Appendix B: select Github logic Limitations of previous implementation While the SQL generator eased so much work for so many different people in the company, it had some... strange caveats, shall we say. Some were more noticable than others but all were, in

## From pipeline to beyond

DevFeed: [From pipeline to beyond](<https://devfeed.tech/articles/from-pipeline-to-beyond-19833.md>)

Original publisher: [Read original article](<https://tech.gc.com/from-pipeline-to-beyond/>)

Author: GameChanger

Published: 2021-05-05T09:00:18Z

Content type: article

Language: en

Sources: [GameChanger](<https://devfeed.tech/sources/gamechanger.md>)

Topics: [Kafka](<https://devfeed.tech/topics/kafka.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [data](<https://devfeed.tech/topics/data.md>), [streaming-data-processing](<https://devfeed.tech/topics/streaming-data-processing.md>), [Back end](<https://devfeed.tech/topics/backend.md>)

Tags: [analysis](<https://devfeed.tech/tags/analysis.md>), [backend](<https://devfeed.tech/tags/backend.md>), [consul](<https://devfeed.tech/tags/consul.md>), [data](<https://devfeed.tech/tags/data.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [s3](<https://devfeed.tech/tags/s3.md>), [terraform](<https://devfeed.tech/tags/terraform.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article examines options for moving data out of Kafka into a data warehouse and archive. It discusses off-the-shelf tools such as Kafka Connect, Secor, and Gobblin, the limitations encountered, and the development of a custom solution. The requirements include preserving Avro data and schemas, writing to S3, and partitioning data by processing time or event time.

### Source excerpt

An overview of off-the-shelf solutions for moving data out of Kafka, problems we had making those systems work, and how we wrote our own solution and stood it up for those in a similar situation. You cannot know everything a system will be used for when you start: it is only at the end of its life you can have such certainty. [1] Many moons ago I wrote about our design for upgrading our data pipeline, which lightly touched on how we'd move data out of our pipeline (Kafka) to downstream systems, namely our data warehouse and data archive. At that time we hadn't really been able to dive into focusing on getting the data out of Kafka, because getting data in to Kafka is often much more custom and complex, and we thought we'd be able to use an off the shelf solution like Kafka Connect to move data out, don't even worry about it. We were, uh -- we were wrong. Let me take you on our journey, in case you're on this journey too. The problem space Solution 1: Kafka Connect Solution 2: Secor or Gobblin Solution 3: we'll do this ourselves Tangent: naming things How you can do this yourselves, code edition How you can do this yourselves, infrastructure and metrics edition Takeaways The problem space Programmers are not to be measured by their ingenuity and their logic but by the completeness of their case analysis. [4] At a high level, the problem we needed a solution for was as follows: Data enters the data pipeline from numerous backend systems. This crossover point is producers into the pipeline, which we'd already implemented. Data from the data pipeline needs to move into the data warehouse and the data archive. This crossover point would be a consumer on the pipeline. Ideally we'd like the same consumer for both needs that we can simply configure differently. We want to preserve our data's Avro format along side its schemas. This would allow every system that interacts with the data to use the same language. We want to write our data to S3. data warehouse: This will be our

## Crash Course to Redshift

DevFeed: [Crash Course to Redshift](<https://devfeed.tech/articles/crash-course-to-redshift-19828.md>)

Original publisher: [Read original article](<https://tech.gc.com/crash-course-to-redshift/>)

Author: GameChanger

Published: 2020-03-30T14:36:13Z

Content type: tutorial

Language: en

Sources: [GameChanger](<https://devfeed.tech/sources/gamechanger.md>)

Topics: [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [Databases](<https://devfeed.tech/topics/databases.md>), [data](<https://devfeed.tech/topics/data.md>), [IO](<https://devfeed.tech/topics/io.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Development](<https://devfeed.tech/topics/development.md>), [debugging](<https://devfeed.tech/topics/debugging.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [development](<https://devfeed.tech/tags/development.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [documentation](<https://devfeed.tech/tags/documentation.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [io](<https://devfeed.tech/tags/io.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [scale](<https://devfeed.tech/tags/scale.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

A crash course to Amazon Redshift that introduces its database architecture, large-scale table design considerations, data distribution, data loading, debugging, and query-performance optimization.

### Source excerpt

Redshift. It can store insane amounts of data. It can also store insane amounts of surprises, considerations, new ideas to learn, skewed tables to fix, distributions to get in line, what's a WLM, what am I doing‽ This post is meant to give you a crash course into working with Redshift, to get you off and running until you have the time and resources to come back and internalize what it all means. This is by no means a comprehensive review of Redshift, as then it'd no longer be a crash course, nor does this dive into data warehousing specifics, which I can cover in another post if people want. At a high level what I'll be covering is: Introduction to Redshift Table design Table analysis Data loading Debugging The vast majority of this post actually comes from our internal documentation, so you can trust that we do use this to help educate those less familiar with Redshift, and get them ramped up and feeling comfortable. Introduction to Redshift On Redshift The Redshift database will behave like other databases you've encountered, but under the hood it has some extra considerations to take into account. The main difference between Redshift and most other databases you'll have encountered is due to scale, with the cluster being important to keep in mind in table design along with standard table design considerations. And since the scale is so much larger, the impact of IO can go up considerably, especially if the cluster needs to move or share data to perform a query. The reasons for this and how to best avoid these inefficiencies are detailed below. More on Redshift database development here. On distributing data Within a Redshift cluster, there is a leader node and many compute nodes. The leader node helps orchestrate the work the compute nodes do. For example, if a query is operating only on data from May of 2017, and all of that data is stored on a single compute node, the leader only needs that node to perform the work. If instead a query is operating on data from

## Let me automate that for you

DevFeed: [Let me automate that for you](<https://devfeed.tech/articles/let-me-automate-that-for-you-19839.md>)

Original publisher: [Read original article](<https://tech.gc.com/let-me-automate-that-for-you/>)

Author: GameChanger

Published: 2019-09-20T18:29:38Z

Content type: article

Language: en

Sources: [GameChanger](<https://devfeed.tech/sources/gamechanger.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [systems](<https://devfeed.tech/topics/systems.md>), [Refactoring](<https://devfeed.tech/topics/refactoring.md>), [Back end](<https://devfeed.tech/topics/backend.md>), [Database](<https://devfeed.tech/topics/database.md>), [SQL](<https://devfeed.tech/topics/sql.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [automation](<https://devfeed.tech/tags/automation.md>), [backend](<https://devfeed.tech/tags/backend.md>), [data](<https://devfeed.tech/tags/data.md>), [database](<https://devfeed.tech/tags/database.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [refactor](<https://devfeed.tech/tags/refactor.md>), [refactoring](<https://devfeed.tech/tags/refactoring.md>), [sql](<https://devfeed.tech/tags/sql.md>), [tests](<https://devfeed.tech/tags/tests.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

The article describes GameChanger's data pipeline and warehouse, where engineers configure producers, pipeline connections, and warehouse tables. A refactoring project simplified producers, removed boilerplate, and added tests, but creating or updating warehouse tables remained a difficult manual task because the warehouse uses different SQL syntax and lacks library-generated code.

### Source excerpt

As GameChanger's data engineer, I oversee the data pipeline and data warehouse. Sounds simple, right? And at a high level, it is! Fig 1.1: high level architecture diagram. Some complexity removed due to it being kinda boring for this post. Producers produce into the pipeline, and our main consumer is the ETL job which moves data to our warehouse, enabling anybody to come get answers to their questions and see what's happening across all our systems. Boom: easy. Well, not quite. Who owns what Since data can come from any number of backend systems and teams, engineers are responsible for writing the setup that shepherds their data through the system: a producer that lives near their data, the pipe it travels through in the pipeline, and the warehouse table. This often means new data that I'm unfamiliar with arrives in our warehouse without me even knowing it's been set up, which is actually kind of neat: the system should be so easy to work with that you don't need the data engineer. After a recent refactoring project, producers were made as simple as possible with removed boilerplate and plenty of tests to automatically catch the most common bugs engineers encounter. Typically, engineers have no problems with making their producers. Fig 2.1: engineers before and after producer refactor project. Studies have shown that engineers prefer to be happy. The pipe their data travels through is set up by filling in a form and pressing a button. Again, engineers typically have no problems with this. It's the warehouse table that becomes a pain point. Follow the readme There are two times non data engineers need to interact with warehouse tables: they've created a new producer which needs a table for their data to land in. they've updated an existing producer which needs its table updated as well. The second point is trickier and easier to get wrong, but the first point proved just as difficult for many engineers and far more common, especially if the engineers in question had

## Engineering a Historic Moment: Shopify Gets Ready for Cannabis in Canada

DevFeed: [Engineering a Historic Moment: Shopify Gets Ready for Cannabis in Canada](<https://devfeed.tech/articles/engineering-a-historic-moment-shopify-gets-ready-for-cannabis-in-canada-1381.md>)

Original publisher: [Read original article](<https://shopify.engineering/engineering-a-historic-moment-shopify-gets-ready-for-cannabis-in-canada>)

Author: Jason Hiltz-Laforge

Published: 2019-02-07T16:30:00Z

Content type: article

Language: en

Sources: [Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering.md>), [Shopify Engineering - Shopify Engineering](<https://devfeed.tech/sources/shopify-engineering-shopify-engineering.md>)

Topics: [Shopify](<https://devfeed.tech/topics/shopify.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>), [Terraform](<https://devfeed.tech/topics/terraform.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [migration](<https://devfeed.tech/topics/migration.md>), [data](<https://devfeed.tech/topics/data.md>), [Networks](<https://devfeed.tech/topics/networks.md>), [DDoS](<https://devfeed.tech/topics/ddos.md>)

Tags: [attacks](<https://devfeed.tech/tags/attacks.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-platform](<https://devfeed.tech/tags/cloud-platform.md>), [clusters](<https://devfeed.tech/tags/clusters.md>), [compute](<https://devfeed.tech/tags/compute.md>), [core](<https://devfeed.tech/tags/core.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [google-cloud-platform](<https://devfeed.tech/tags/google-cloud-platform.md>), [industry](<https://devfeed.tech/tags/industry.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [migration](<https://devfeed.tech/tags/migration.md>), [networks](<https://devfeed.tech/tags/networks.md>), [platform](<https://devfeed.tech/tags/platform.md>), [resiliency](<https://devfeed.tech/tags/resiliency.md>), [retail](<https://devfeed.tech/tags/retail.md>), [shopify](<https://devfeed.tech/tags/shopify.md>), [terraform](<https://devfeed.tech/tags/terraform.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

Shopify describes the engineering work required to support cannabis retailers in Canada after legalization in 2018. The team built a Montreal-based regional deployment on Google Cloud Platform to satisfy Canadian data-residency requirements, using Google Kubernetes Engine, Google Compute Engine, Terraform, regional clusters, workload segregation, and a regional data warehouse. It also modeled traffic scenarios and provisioned capacity for launches, media-driven surges, sales peaks, and possible denial-of-service attacks.

### Source excerpt

On October 17th, 2018, Canada ended a 95-year history of cannabis prohibition. For Shopify, the legalization of cannabis marked a new industry entering the Canadian retail market and we worked with governments and licensed sellers across the country to provide a safe, reliable and scalable platform for their business. For our engineering team, it meant significant changes to our platform to meet the strict requirements for this newly regulated industry.

## Things to consider before building a serverless data warehouse

DevFeed: [Things to consider before building a serverless data warehouse](<https://devfeed.tech/articles/things-to-consider-before-building-a-serverless-data-warehouse-14449.md>)

Original publisher: [Read original article](<https://www.serverless.com/blog/things-consider-building-serverless-data-warehouse>)

Author: Ashan Fernando

Published: 2018-08-29T00:00:00Z

Content type: article

Language: en

Sources: [Serverless Blog](<https://devfeed.tech/sources/serverless-blog.md>)

Topics: [Serverless](<https://devfeed.tech/topics/serverless.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [cloud-computing](<https://devfeed.tech/tags/cloud-computing.md>), [data](<https://devfeed.tech/tags/data.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [faas](<https://devfeed.tech/tags/faas.md>), [function-as-a-service](<https://devfeed.tech/tags/function-as-a-service.md>), [infrastructure-as-code](<https://devfeed.tech/tags/infrastructure-as-code.md>), [post](<https://devfeed.tech/tags/post.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [serverless-architecture](<https://devfeed.tech/tags/serverless-architecture.md>), [serverless-framework](<https://devfeed.tech/tags/serverless-framework.md>), [tips](<https://devfeed.tech/tags/tips.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

A post about considerations and practical tips for building a serverless data warehouse.

### Source excerpt

Is it time for the rise of the serverless data warehouse? Read this post to find out, and for some serverless data warehousing pro-tips and considerations.

## A Journey Towards a Custom Data Warehouse Solution Part 2: We Need Storage

DevFeed: [A Journey Towards a Custom Data Warehouse Solution Part 2: We Need Storage](<https://devfeed.tech/articles/a-journey-towards-a-custom-data-warehouse-solution-part-2-we-need-storage-35108.md>)

Original publisher: [Read original article](<https://upday.github.io/blog/dwh-part2-we-need-storage/>)

Author: Robert Bordo (robert@upday.com)

Published: 2017-08-22T04:39:55Z

Content type: article

Language: en

Sources: [Upday](<https://devfeed.tech/sources/upday.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [data lake](<https://devfeed.tech/topics/data-lake.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [business-intelligence](<https://devfeed.tech/tags/business-intelligence.md>), [data](<https://devfeed.tech/tags/data.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [s3](<https://devfeed.tech/tags/s3.md>), [scalability](<https://devfeed.tech/tags/scalability.md>), [storage](<https://devfeed.tech/tags/storage.md>), [warehouse](<https://devfeed.tech/tags/warehouse.md>)

### AI overview

This article examines storage choices for a custom data warehouse. It describes application log data, the limitations of time-series databases for additional master and historical data, business intelligence access needs, and the use of Amazon S3 as a data lake while considering other storage options including AWS Redshift.

### Source excerpt

In the beginning we created a cluster. And the cluster was without form, and void; and nulls were upon the face of the storage. As we learned in part 1 of our series, a data warehouse consists of several components. The key component is the storage. All the others group around it. But how can one draw a decision on which storage solution to adopt? What is out there anyway? Preface In a perfect world there would be only one kind of storage that fits all the needs of current DWH development and analysis. But since we are not living in that kind of place, we have several options. And the number of options increase the deeper one dives into the topic. There seem to be solutions for every use case you can think of. That might be a good starting point. What is our most common use case? What are we going to store? And how would we like to access our data in the end? Our major source is a massive amount of log data coming from our app. Everything the user does (e.g swiping through articles, selecting categories, leaving the app) is tracked, enriched with metadata (e.g. the user's location, app version, article identifier) and stored by a third-party service in big, semi-structured log files. Having only this source, a time series database like Graphite or InfluxDB could do the job. But also having slow changing master data, like user profiles, article metadata and maybe even to keep a history of data, this solution would not satisfy our current and future needs. Another thing that comes to my mind is how the data will be accessed by our final consumer (namely: Business Intelligence). Usually they use tools like Jasper Reports or Tableau for generating reports. For analyses we have to pre-aggregate the data to make queries more performant and translate raw information into a digestible format. What else is on the market? Storage good at bad at Example S3/Flat Files scalability, easy to use, data lake querying S3 Time Series DB handling time series data non time series data G