# Hadoop

Published articles for Hadoop.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## The Creator of Pandas on AI, Apache Arrow, and the Future of Software Engineering

DevFeed: [The Creator of Pandas on AI, Apache Arrow, and the Future of Software Engineering](<https://devfeed.tech/articles/the-creator-of-pandas-on-ai-apache-arrow-and-the-future-of-software-engineering-38717.md>)

Original publisher: [Read original article](<https://dataengineeringcentral.substack.com/p/the-creator-of-pandas-on-ai-apache>)

Author: Daniel Beach

Published: 2026-07-08T12:16:09Z

Content type: article

Language: en

Sources: [Data Engineering Central](<https://devfeed.tech/sources/data-engineering-central.md>)

Topics: [pandas](<https://devfeed.tech/topics/pandas.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [DuckDB](<https://devfeed.tech/topics/duckdb.md>), [data-platforms](<https://devfeed.tech/topics/data-platforms.md>), [software-development](<https://devfeed.tech/topics/software-development.md>), [future of software](<https://devfeed.tech/topics/future-of-software.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [parquet](<https://devfeed.tech/topics/parquet.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [apache-arrow](<https://devfeed.tech/tags/apache-arrow.md>), [arrow](<https://devfeed.tech/tags/arrow.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [data-platforms](<https://devfeed.tech/tags/data-platforms.md>), [duckdb](<https://devfeed.tech/tags/duckdb.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pandas](<https://devfeed.tech/tags/pandas.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [software-development](<https://devfeed.tech/tags/software-development.md>)

### AI overview

An interview with Wes McKinney covers the origins of pandas and Apache Arrow, the evolution of modern data engineering from Hadoop to lakehouse architectures, and the roles of tools such as Parquet, DuckDB, DataFusion, and Spark. McKinney also discusses how AI affects software development, arguing that it can improve experienced engineers' productivity but does not replace software engineering, architecture, or judgment.

### Source excerpt

interview with Wes McKinney

## Beyond the warehouse: How METRO Markets built a do-it-all data platform on ClickHouse Cloud

DevFeed: [Beyond the warehouse: How METRO Markets built a do-it-all data platform on ClickHouse Cloud](<https://devfeed.tech/articles/beyond-the-warehouse-how-metro-markets-built-a-do-it-all-data-platform-on-clickhouse-cloud-5416.md>)

Original publisher: [Read original article](<https://clickhouse.com/blog/metro-markets-data-warehouse>)

Author: ClickHouse

Published: 2026-06-18T20:31:40Z

Content type: article

Language: en

Sources: [ClickHouse Blog](<https://devfeed.tech/sources/clickhouse-blog.md>)

Topics: [clickhouse](<https://devfeed.tech/topics/clickhouse.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [data](<https://devfeed.tech/topics/data.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [observability](<https://devfeed.tech/topics/observability.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [log management](<https://devfeed.tech/topics/log-management.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [bigquery](<https://devfeed.tech/tags/bigquery.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [logging](<https://devfeed.tech/tags/logging.md>), [observability](<https://devfeed.tech/tags/observability.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

METRO Markets replaced its self-hosted Hadoop-based data stack with ClickHouse Cloud, consolidating company data into a single warehouse that supports data warehousing, real-time seller analytics, credit risk modeling, observability, operational logging, and new AI use cases.

### Source excerpt

How METRO Markets replaced a failing Hadoop-based stack with ClickHouse Cloud to build a single platform now powering data warehousing, real-time seller analytics, credit risk modeling, observability, and AI across the whole company.

## Why HTAP Databases Are Giving Way to Disaggregated Architectures

DevFeed: [Why HTAP Databases Are Giving Way to Disaggregated Architectures](<https://devfeed.tech/articles/htap-is-dead-5421.md>)

Original publisher: [Read original article](<https://neon.com/blog/htap-is-dead>)

Author: Zhou Sun

Published: 2025-05-04T10:00:00Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [olap](<https://devfeed.tech/topics/olap.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [consistency](<https://devfeed.tech/topics/consistency.md>), [NoSQL](<https://devfeed.tech/topics/nosql.md>), [MongoDB](<https://devfeed.tech/topics/mongodb.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [hdfs](<https://devfeed.tech/topics/hdfs.md>), [CockroachDB](<https://devfeed.tech/topics/cockroachdb.md>), [Amazon Redshift](<https://devfeed.tech/topics/amazon-redshift.md>), [vitess](<https://devfeed.tech/topics/vitess.md>)

Tags: [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-data](<https://devfeed.tech/tags/cloud-data.md>), [cockroachdb](<https://devfeed.tech/tags/cockroachdb.md>), [consistency](<https://devfeed.tech/tags/consistency.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [databases](<https://devfeed.tech/tags/databases.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hdfs](<https://devfeed.tech/tags/hdfs.md>), [latency](<https://devfeed.tech/tags/latency.md>), [mongodb](<https://devfeed.tech/tags/mongodb.md>), [olap](<https://devfeed.tech/tags/olap.md>), [redshift](<https://devfeed.tech/tags/redshift.md>), [sql](<https://devfeed.tech/tags/sql.md>), [vitess](<https://devfeed.tech/tags/vitess.md>)

### AI overview

The article traces the separation of transactional and analytical database workloads, explaining how differing storage and scaling requirements led to specialized OLTP and OLAP systems. It argues that HTAP as a single database architecture is declining while its underlying ideas persist in today's disaggregated data stack.

### Source excerpt

This blog is inspired by Jordan Tigani's "Big Data is Dead." Jordan and I actually spent some time building an HTAP database at SingleStore. From the one database that did everything in the '80s, to the great divide, to HTAP, to today's disaggregated stack--here's why HTAP as a database is dead, but its spirit lives on.

## Out with the old file system

DevFeed: [Out with the old file system](<https://devfeed.tech/articles/out-with-the-old-file-system-8769.md>)

Original publisher: [Read original article](<https://trino.io/blog/2025/02/10/old-file-system.html>)

Author: Manfred Moser, David Phillips, Mateusz Gajewski

Published: 2025-02-10T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Filesystems](<https://devfeed.tech/topics/filesystems.md>), [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>)

Tags: [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [apache-iceberg](<https://devfeed.tech/tags/apache-iceberg.md>), [apache-parquet](<https://devfeed.tech/tags/apache-parquet.md>), [azure](<https://devfeed.tech/tags/azure.md>), [compression](<https://devfeed.tech/tags/compression.md>), [data-lake](<https://devfeed.tech/tags/data-lake.md>), [deprecated](<https://devfeed.tech/tags/deprecated.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [maintenance](<https://devfeed.tech/tags/maintenance.md>), [release](<https://devfeed.tech/tags/release.md>)

### AI overview

Trino 470 deprecated its legacy Hadoop-based file system support, which will be removed in a future release. The article explains Trino's move toward custom file system implementations for cloud storage and describes the migration path for users, including catalog-level file system configuration and warnings for deprecated properties.

### Source excerpt

What a long journey it has been! From the start Trino supported querying Hive data and used libraries from the Hive and Hadoop ecosystem. With the release of Trino 470 we mark another milestone to more features and better performance for data lake and lakehouse querying with Trino. We deprecated the legacy file system support, and will permanently remove them in an upcoming release.

## The long journey to Apache Ranger

DevFeed: [The long journey to Apache Ranger](<https://devfeed.tech/articles/the-long-journey-to-apache-ranger-8766.md>)

Original publisher: [Read original article](<https://trino.io/blog/2024/12/02/ranger.html>)

Author: Manfred Moser

Published: 2024-12-02T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [releases](<https://devfeed.tech/topics/releases.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [apache](<https://devfeed.tech/tags/apache.md>), [data](<https://devfeed.tech/tags/data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [integration](<https://devfeed.tech/tags/integration.md>), [pull-request](<https://devfeed.tech/tags/pull-request.md>), [release](<https://devfeed.tech/tags/release.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Trino 466 adds an Apache Ranger plugin for data access control. The article recounts the history of the integration, including prior pull requests, container-image support for testing, and renewed work in 2024.

### Source excerpt

Apache Ranger has arrived! With the new Trino 466 you all get another jam-packed release of Trino awesomeness. One of the goodies is a new plugin for access control for your data with Apache Ranger, and it has gone through a long story to get here. Apache Ranger has a long history and wide adoption as an access control system for data lakes using Hadoop and Hive. Since Trino brings fast analytics to this space, and also supports modern data lakehouses and other data sources, Apache Ranger is a natural fit for access control on a Trino-powered data platform.

## How Uber Reduced Their Log Size By 99%

DevFeed: [How Uber Reduced Their Log Size By 99%](<https://devfeed.tech/articles/how-uber-reduced-their-log-size-by-99-17981.md>)

Original publisher: [Read original article](<https://newsletter.betterstack.com/p/how-uber-reduced-their-log-size-by>)

Author: Richard Oliver Bray

Published: 2024-10-09T13:02:55Z

Content type: article

Language: en

Sources: [Hacking Scale by Better Stack](<https://devfeed.tech/sources/hacking-scale-by-better-stack.md>)

Topics: [Logging](<https://devfeed.tech/topics/logging.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Filesystems](<https://devfeed.tech/topics/filesystems.md>), [big-data](<https://devfeed.tech/topics/big-data.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [apache-spark](<https://devfeed.tech/tags/apache-spark.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [cli](<https://devfeed.tech/tags/cli.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data](<https://devfeed.tech/tags/data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hdfs](<https://devfeed.tech/tags/hdfs.md>), [logging](<https://devfeed.tech/tags/logging.md>)

### AI overview

The article explains how Uber addressed the storage cost of generating roughly 5 PB of INFO-level logs each month. It describes Uber's use of HDFS and related data-processing tools, while reporting that the company reduced log storage size by 99%.

### Source excerpt

Uber broke apart an open source tool to massively compress their logs

## The crushing success of relational databases

DevFeed: [The crushing success of relational databases](<https://devfeed.tech/articles/the-crushing-success-of-relational-databases-5768.md>)

Original publisher: [Read original article](<https://neon.com/blog/relational-databases-success>)

Author: Andy Hattemer

Published: 2024-08-13T20:27:05Z

Content type: opinion

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Graphs](<https://devfeed.tech/topics/graphs.md>)

Tags: [databases](<https://devfeed.tech/tags/databases.md>), [graphs](<https://devfeed.tech/tags/graphs.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [postgres](<https://devfeed.tech/tags/postgres.md>), [relational-databases](<https://devfeed.tech/tags/relational-databases.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

The article argues that relational databases remain dominant because they repeatedly absorb capabilities associated with competing data models and query systems. It presents document, vector, and graph functionality as features that can coexist within a relational database, and revisits arguments that the relational model will continue to outlast attempted replacements.

### Source excerpt

Every once in a while, a revolutionary product comes along and changes everything. And today, we're talking about three of these phenomenal products. The first one is a document store. The second is a vector database. And the third is a graph database. So, three things: document...

## Expanded Memory and Compute with Heroku's New Larger Dynos

DevFeed: [Expanded Memory and Compute with Heroku's New Larger Dynos](<https://devfeed.tech/articles/expanded-memory-and-compute-with-heroku-s-new-larger-dynos-26437.md>)

Original publisher: [Read original article](<https://www.heroku.com/blog/heroku-larger-dyno-types/>)

Author: Ethan Limchayseng

Published: 2024-03-28T02:25:00Z

Content type: release

Language: en

Sources: [Heroku](<https://devfeed.tech/sources/heroku.md>)

Topics: [Heroku](<https://devfeed.tech/topics/heroku.md>), [cpu](<https://devfeed.tech/topics/cpu.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [big-data](<https://devfeed.tech/topics/big-data.md>), [Data analysis](<https://devfeed.tech/topics/data-analysis.md>)

Tags: [big-data](<https://devfeed.tech/tags/big-data.md>), [cache](<https://devfeed.tech/tags/cache.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cloud-infrastructure](<https://devfeed.tech/tags/cloud-infrastructure.md>), [compute](<https://devfeed.tech/tags/compute.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [data](<https://devfeed.tech/tags/data.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [dynos](<https://devfeed.tech/tags/dynos.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [heroku](<https://devfeed.tech/tags/heroku.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [memory](<https://devfeed.tech/tags/memory.md>), [news](<https://devfeed.tech/tags/news.md>), [performance-optimization](<https://devfeed.tech/tags/performance-optimization.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [private-spaces](<https://devfeed.tech/tags/private-spaces.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [spark](<https://devfeed.tech/tags/spark.md>), [video](<https://devfeed.tech/tags/video.md>)

### AI overview

Heroku introduces nine larger dyno types across its Performance, Private, and Shield tiers, adding higher memory and CPU limits for compute-intensive workloads. The new sizes support use cases including real-time analytics, caching, machine learning, video encoding, and simulations.

### Source excerpt

Introduction Heroku is excited to introduce nine new dyno types to our fleets and product offerings. In 2014, we introduced Performance-tier dynos, giving our customers fully dedicated resources to run their most compute-intensive workloads. Now in 2024, today's standards are rapidly increasing as complex applications and growing data volumes consume more memory and carry heavier [...] The post Expanded Memory and Compute with Heroku's New Larger Dynos appeared first on Heroku.

## Implementing Data Validation with Great Expectations in Hybrid Environments

DevFeed: [Implementing Data Validation with Great Expectations in Hybrid Environments](<https://devfeed.tech/articles/implementing-data-validation-with-great-expectations-in-hybrid-environments-28036.md>)

Original publisher: [Read original article](<https://tech.trivago.com/post/2023-04-25-implementing-data-validation-with-great-expectations-in-hybrid-environments/>)

Author: Kamila Widyanto Full time DevOps; Site Reliability Engineer; Part Time Rubberduck Linkedin Profile

Published: 2023-04-25T00:00:00Z

Content type: tutorial

Language: en

Sources: [Trivago](<https://devfeed.tech/sources/trivago.md>)

Topics: [data-processing](<https://devfeed.tech/topics/data-processing.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [hdfs](<https://devfeed.tech/topics/hdfs.md>), [integrity](<https://devfeed.tech/topics/integrity.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Library](<https://devfeed.tech/topics/library.md>), [Python](<https://devfeed.tech/topics/python.md>), [JSON](<https://devfeed.tech/topics/json.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [YAML](<https://devfeed.tech/topics/yaml.md>), [version-control](<https://devfeed.tech/topics/version-control.md>)

Tags: [configuration](<https://devfeed.tech/tags/configuration.md>), [data](<https://devfeed.tech/tags/data.md>), [data-pipeline](<https://devfeed.tech/tags/data-pipeline.md>), [data-processing](<https://devfeed.tech/tags/data-processing.md>), [data-validation](<https://devfeed.tech/tags/data-validation.md>), [devops](<https://devfeed.tech/tags/devops.md>), [distributed](<https://devfeed.tech/tags/distributed.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hdfs](<https://devfeed.tech/tags/hdfs.md>), [integrity](<https://devfeed.tech/tags/integrity.md>), [json](<https://devfeed.tech/tags/json.md>), [library](<https://devfeed.tech/tags/library.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [python](<https://devfeed.tech/tags/python.md>), [version-control](<https://devfeed.tech/tags/version-control.md>), [workflow](<https://devfeed.tech/tags/workflow.md>), [yaml](<https://devfeed.tech/tags/yaml.md>)

### AI overview

This article describes implementing Great Expectations for data validation in a hybrid Hadoop environment. It explains the framework's core concepts and how the authors ran it as a PySpark job in an automated data pipeline, including configuring the Data Context for HDFS constraints.

### Source excerpt

Data validation is an essential step in any data processing pipeline, as it ensures the integrity and accuracy of the data to be used across all subsequent processing steps.

## Journey to Iceberg with Trino

DevFeed: [Journey to Iceberg with Trino](<https://devfeed.tech/articles/journey-to-iceberg-with-trino-8705.md>)

Original publisher: [Read original article](<https://trino.io/blog/2022/12/19/trino-summit-2022-sk-telecom-recap.html>)

Author: JaeChang Song, Jennifer Oh, Brian Olsen

Published: 2022-12-19T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [big-data](<https://devfeed.tech/topics/big-data.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>)

Tags: [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hdfs](<https://devfeed.tech/tags/hdfs.md>), [logs](<https://devfeed.tech/tags/logs.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [post](<https://devfeed.tech/tags/post.md>), [scale](<https://devfeed.tech/tags/scale.md>), [speed](<https://devfeed.tech/tags/speed.md>), [summit](<https://devfeed.tech/tags/summit.md>), [switching](<https://devfeed.tech/tags/switching.md>)

### AI overview

SK Telecom describes its journey from Hive-based Trino deployments to Iceberg after encountering scaling and performance problems. The company used Trino across Hadoop and HDFS-based data platforms, collected query plans, JMX statistics, system metrics, and logs, and built a dashboard to investigate blocked queries and cluster behavior.

### Source excerpt

This post comes from the second half of Trino Summit 2022 session. Our friends JaeChang and Jennifer from SK Telecom traveled across the globe from South Korea to join us in person! SK Telecom recently had some issues scaling Trino on the Hive model, among other issues that come with Hive. While some initial tweaking helped speed things up, it ultimately never solved the problem. After switching to Iceberg, SK Telecom ran initial performance tests with some very impressive results. In this talk, Jennifer and JaeChang describe their journey to Iceberg with Trino.

## Building A Modern Data Stack for QazAI

DevFeed: [Building A Modern Data Stack for QazAI](<https://devfeed.tech/articles/building-a-modern-data-stack-for-qazai-8676.md>)

Original publisher: [Read original article](<https://trino.io/blog/2022/06/08/building-a-modern-data-stack-for-qaz-ai.html>)

Author: Baurzhan Kuspayev

Published: 2022-06-08T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [data](<https://devfeed.tech/topics/data.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [building](<https://devfeed.tech/tags/building.md>), [clickhouse](<https://devfeed.tech/tags/clickhouse.md>), [data](<https://devfeed.tech/tags/data.md>), [databases](<https://devfeed.tech/tags/databases.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [s3](<https://devfeed.tech/tags/s3.md>), [speed](<https://devfeed.tech/tags/speed.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

QazAI describes moving from a data stack built around S3, Hive, and Clickhouse toward Trino for faster, lower-cost analytics and ETL exploration. The article highlights Trino's SQL support, federated queries, setup simplicity, and a lack of fault tolerance for ETL pipelines.

### Source excerpt

At QazAI, we build data lakes as a service for companies. In the original architecture, we get raw data in S3, transform the S3 data with Hive, and then delivered the data to business units via our datamart built on Clickhouse (for optimal delivery speeds). Over time, we were dragged down by the slower speeds and high costs of running Hive, and started shopping for a faster and cheaper open source engine to do our ETL data transformations.

## A Serverless Data Engineering Stack at Teamwork Using Go and AWS

DevFeed: [A Serverless Data Engineering Stack at Teamwork Using Go and AWS](<https://devfeed.tech/articles/the-go-serverless-data-engineering-revolution-at-teamwork-golang-aws-35103.md>)

Original publisher: [Read original article](<https://engineroom.teamwork.com/the-go-serverless-data-engineering-revolution-at-teamwork-golang-aws-f2fd3cb1f563?source=rss----cea4eecd5960---4>)

Author: Joe Minichino

Published: 2021-10-19T13:58:52Z

Content type: article

Language: en

Sources: [Teamwork](<https://devfeed.tech/sources/teamwork.md>)

Topics: [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>), [Go Language](<https://devfeed.tech/topics/go-language.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [AWS Lambda](<https://devfeed.tech/topics/aws-lambda.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [GitHub Actions](<https://devfeed.tech/topics/github-actions.md>), [Docker](<https://devfeed.tech/topics/docker.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [airflow](<https://devfeed.tech/topics/airflow.md>)

Tags: [airflow](<https://devfeed.tech/tags/airflow.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-athena](<https://devfeed.tech/tags/aws-athena.md>), [aws-lambda](<https://devfeed.tech/tags/aws-lambda.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [docker](<https://devfeed.tech/tags/docker.md>), [github-actions](<https://devfeed.tech/tags/github-actions.md>), [go](<https://devfeed.tech/tags/go.md>), [golang](<https://devfeed.tech/tags/golang.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [s3](<https://devfeed.tech/tags/s3.md>), [serverless](<https://devfeed.tech/tags/serverless.md>), [serverless-architecture](<https://devfeed.tech/tags/serverless-architecture.md>)

### AI overview

This article describes Teamwork's serverless-oriented data engineering stack and the reasons for choosing it. It discusses AWS services including S3, Athena, Glue, QuickSight, Kinesis, and Lambda, with Go used for Lambda processing and GitHub Actions used for deployment.

### Source excerpt

A real data lake. Traditional Data Engineering relies on products such as Airflow, Hadoop, Spark and Spark-based architectures, or similar technologies. These are still viable solutions for a number of reason, not least the fact that Data Engineers are few and far between, and the vast majority of them will be familiar in the above technologies or similar products/frameworks. Go Serverless I wrote an article about our tech stack, which includes S3, Athena, Glue, Quicksight, Kinesis etc. What is not immediately apparent is the serverless nature of our stack, which was a deliberate choice taken in the context of a Data Analytics department which was started as an experiment and had to pick its battles very wisely. Sysops, cluster / server management, CI/CD were not top of our list. Creating dashboards was. Also before Teamwork I had a nearly 2-year run working with a company that was entirely serverless in their set up AND mentality, and it was a career-changing experience. We worked prevalently with AWS Lambda when Lambda was a new toy, and we loved it. Finally, despite the "buzz" over Functional Programming being seemingly over, I am an arduous fanatic of it and of what it represents philosophically, a way to represent each problem in terms of an input, some transformation, an output and when possible, no side effects. NOTE: some of the tools I mention are not strictly serverless as much as they are fully managed (eg. Kinesis or Github Actions), but they still involve little or no sysops / devops. Serverless Processing: AWS Lambda Once upon a time you would have to settle for Node.js or Python to write Lambda code, but nowadays not only you can use a whole lot of runtimes, you can simply use your own docker image et voila', you're ready to go. In our case we are ready to Go, since at Teamwork Go(lang) is our programming language of choice. It's fast, easy to use, very safe and reliable, and given that it can complete the same operation faster than most other program

## Плагин Big Data Tools теперь поддерживает IntelliJ IDEA Ultimate, PyCharm Professional, DataGrip 2021.3 EAP и DataSpell

DevFeed: [Плагин Big Data Tools теперь поддерживает IntelliJ IDEA Ultimate, PyCharm Professional, DataGrip 2021.3 EAP и DataSpell](<https://devfeed.tech/articles/big-data-tools-intellij-idea-ultimate-pycharm-professional-datagrip-2021-3-eap-dataspell-23936.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/JetBrains/articles/580344/>)

Author: olegchir (JetBrains)

Published: 2021-09-28T06:17:21Z

Content type: release

Language: ru

Sources: [JetBrains RU](<https://devfeed.tech/sources/jetbrains-ru.md>)

Topics: [Apache Spark](<https://devfeed.tech/topics/spark.md>), [ide](<https://devfeed.tech/topics/ide.md>), [IntelliJ IDEA](<https://devfeed.tech/topics/intellij-idea.md>), [PyCharm](<https://devfeed.tech/topics/pycharm.md>), [Data Science](<https://devfeed.tech/topics/data-science.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Kafka](<https://devfeed.tech/topics/kafka.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [amazon](<https://devfeed.tech/tags/amazon.md>), [apache-kafka](<https://devfeed.tech/tags/apache-kafka.md>), [apache-spark](<https://devfeed.tech/tags/apache-spark.md>), [apache-zeppelin](<https://devfeed.tech/tags/apache-zeppelin.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-s3](<https://devfeed.tech/tags/aws-s3.md>), [configuration](<https://devfeed.tech/tags/configuration.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [ide](<https://devfeed.tech/tags/ide.md>), [intellij-idea](<https://devfeed.tech/tags/intellij-idea.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [pycharm](<https://devfeed.tech/tags/pycharm.md>), [python](<https://devfeed.tech/tags/python.md>), [s3](<https://devfeed.tech/tags/s3.md>), [spark](<https://devfeed.tech/tags/spark.md>), [zeppelin](<https://devfeed.tech/tags/zeppelin.md>)

### AI overview

JetBrains released a new Big Data Tools plugin build compatible with IntelliJ IDEA Ultimate and PyCharm Professional 2021.3, with planned support for DataGrip 2021.3 EAP and support for running in DataSpell. The update adds features including Spark Submit run configurations, Kafka monitoring, AWS S3 named profiles, Zeppelin notebook search, and Python interpreter selection.

### Source excerpt

Недавно мы выпустили новую сборку плагина Big Data Tools, совместимую со свежими (2021.3) версиями IntelliJ IDEA Ultimate и PyCharm Professional. Когда в октябре выйдет DataGrip 2021.3, эта сборка тоже будет с ним работать. Более того, теперь мы умеем запускаться в DataSpell -- новой IDE для Data Science. Если вы используете старые версии Big Data Tools, сейчас самое время обновиться и попробовать новую версию плагина вместе со свежей версией IDE! В этом году мы много чего улучшили и добавили совершенно новые фичи (например, запуск Spark Submit в виде Run Configuration). Вот небольшой список изменений за этот год. Этот список -- лишь небольшая капля в море того, что изменилось с прошлого года. Читать далее

## Обзор плагина Big Data Tools

DevFeed: [Обзор плагина Big Data Tools](<https://devfeed.tech/articles/big-data-tools-23923.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/JetBrains/articles/570088/>)

Author: olegchir (JetBrains)

Published: 2021-07-28T10:41:36Z

Content type: article

Language: ru

Sources: [JetBrains RU](<https://devfeed.tech/sources/jetbrains-ru.md>)

Topics: [ide](<https://devfeed.tech/topics/ide.md>), [big-data](<https://devfeed.tech/topics/big-data.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [data](<https://devfeed.tech/topics/data.md>), [Scala](<https://devfeed.tech/topics/scala.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [big-data](<https://devfeed.tech/tags/big-data.md>), [big-data-tools](<https://devfeed.tech/tags/big-data-tools.md>), [data](<https://devfeed.tech/tags/data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [ide](<https://devfeed.tech/tags/ide.md>), [jetbrains](<https://devfeed.tech/tags/jetbrains.md>), [kafka](<https://devfeed.tech/tags/kafka.md>), [python](<https://devfeed.tech/tags/python.md>), [scala](<https://devfeed.tech/tags/scala.md>), [spark](<https://devfeed.tech/tags/spark.md>), [tools](<https://devfeed.tech/tags/tools.md>), [zeppelin](<https://devfeed.tech/tags/zeppelin.md>)

### AI overview

This article reviews JetBrains' Big Data Tools plugin for working with cloud file systems, Hadoop, Spark, and Zeppelin directly from an IDE. It explains the plugin's role in data-engineering workflows, including ETL, and notes support for Scala and Python.

### Source excerpt

Храните файлы в облачных файловых системах или, может быть, используете Hadoop, Spark и Zeppelin? А пробовали ли вы работать с ними напрямую из IDE? Привет, меня зовут Олег, я из команды плагина Big Data Tools. В этой статье мы поговорим, зачем этот плагин нужен, как применяется и где его достать. За последний год плагин прошёл большой путь и из экспериментального продукта превратился в боевое решение, на которое стоит взглянуть специалистам по Big Data. В JetBrains мы создаем IDE и другие инструменты, которые делают жизнь разработчиков лучше. Big Data Tools -- это очень узкоспециализированный, редкоземельный плагин, который предназначен для конкретного вида разработчиков -- для дата-инженеров. Если вам интересно подробней узнать о мире Big Data и работе дата-инженеров, рекомендую развернутую серию статей Паши Финкельштейна. Здесь мы рассмотрим одну из самых популярных схем. Читать далее

## Trino on ice I: A gentle introduction To Iceberg

DevFeed: [Trino on ice I: A gentle introduction To Iceberg](<https://devfeed.tech/articles/trino-on-ice-i-a-gentle-introduction-to-iceberg-8663.md>)

Original publisher: [Read original article](<https://trino.io/blog/2021/05/03/a-gentle-introduction-to-iceberg.html>)

Author: Brian Olsen

Published: 2021-05-03T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Apache Iceberg](<https://devfeed.tech/topics/apache-iceberg.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [big-data](<https://devfeed.tech/tags/big-data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

An introduction to Apache Iceberg's table format with Trino, framed by the limitations and informal conventions of the Hive connector model.

### Source excerpt

Welcome to the Trino on ice series, covering the details around how the Iceberg table format works with the Trino query engine. The examples build on each previous post, so it's recommended to read the posts sequentially and reference them as needed later. Here are links to the posts in this series: Trino on ice I: A gentle introduction to Iceberg Trino on ice II: In-place table evolution and cloud compatibility with Iceberg Trino on ice III: Iceberg concurrency model, snapshots, and the Iceberg spec Trino on ice IV: Deep dive into Iceberg internals Back in the Gentle introduction to the Hive connector blog post, I discussed a commonly misunderstood architecture and uses of the Trino Hive connector. In short, while some may think the name indicates Trino makes a call to a running Hive instance, the Hive connector does not use the Hive runtime to answer queries. Instead, the connector is named Hive connector because it relies on Hive conventions and implementation details from the Hadoop ecosystem - the invisible Hive specification.

## We're rebranding PrestoSQL as Trino

DevFeed: [We're rebranding PrestoSQL as Trino](<https://devfeed.tech/articles/we-re-rebranding-prestosql-as-trino-8657.md>)

Original publisher: [Read original article](<https://trino.io/blog/2020/12/27/announcing-trino.html>)

Author: Martin Traverso, Dain Sundstrom, David Phillips

Published: 2020-12-27T00:00:00Z

Content type: release

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Software](<https://devfeed.tech/topics/software.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Data Infrastructure](<https://devfeed.tech/topics/data-infrastructure.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [GitHub](<https://devfeed.tech/topics/github.md>), [Slack](<https://devfeed.tech/topics/slack.md>)

Tags: [community](<https://devfeed.tech/tags/community.md>), [contributors](<https://devfeed.tech/tags/contributors.md>), [data-infrastructure](<https://devfeed.tech/tags/data-infrastructure.md>), [github](<https://devfeed.tech/tags/github.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [rebranding](<https://devfeed.tech/tags/rebranding.md>)

### AI overview

The article announces that PrestoSQL is being rebranded as Trino. It explains that the software and community remain intact while describing the project's origins in low-latency analytics over Hadoop data and its commitment to an open, independent, collaborative community.

### Source excerpt

We're rebranding PrestoSQL as Trino. The software and the community you have come to love and depend on aren't going anywhere, we are simply renaming. Trino is the new name for PrestoSQL, the project supported by the founders and creators of Presto® along with the major contributors - just under a shiny new name. And now you can find us here: GitHub: https://github.com/trinodb/trino. Please give it a star! Twitter: @trinodb Slack: https://trino.io/slack.html If you want to learn why we're doing this, read on...

## A gentle introduction to the Hive connector

DevFeed: [A gentle introduction to the Hive connector](<https://devfeed.tech/articles/a-gentle-introduction-to-the-hive-connector-8654.md>)

Original publisher: [Read original article](<https://trino.io/blog/2020/10/20/intro-to-hive-connector.html>)

Author: Brian Olsen

Published: 2020-10-20T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [data](<https://devfeed.tech/topics/data.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>), [systems](<https://devfeed.tech/topics/systems.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [CSV](<https://devfeed.tech/topics/csv.md>), [JSON](<https://devfeed.tech/topics/json.md>)

Tags: [amazon-s3](<https://devfeed.tech/tags/amazon-s3.md>), [blog](<https://devfeed.tech/tags/blog.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-storage](<https://devfeed.tech/tags/cloud-storage.md>), [code](<https://devfeed.tech/tags/code.md>), [data](<https://devfeed.tech/tags/data.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [json](<https://devfeed.tech/tags/json.md>), [object-storage](<https://devfeed.tech/tags/object-storage.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [s3](<https://devfeed.tech/tags/s3.md>), [spark](<https://devfeed.tech/tags/spark.md>), [systems](<https://devfeed.tech/tags/systems.md>)

### AI overview

This article gently explains Trino's Hive connector, clarifying that it reads data organized according to Hive conventions without using the Hive runtime. It introduces the main components of Hive architecture and explains how object storage and metadata connect SQL tables to files.

### Source excerpt

TL;DR: The Hive connector is what you use in Trino for reading data from object storage that is organized according to the rules laid out by Hive, without using the Hive runtime code. One of the most confusing aspects when starting Trino is the Hive connector. Typically, you seek out the use of Trino when you experience an intensely slow query turnaround from your existing Hadoop, Spark, or Hive infrastructure. In fact, the genesis of Trino, formerly known as Presto, came about due to these slow Hive query conditions at Facebook back in 2012. So when you learn that Trino has a Hive connector, it can be rather confusing since you moved to Trino to circumvent the slowness of your current Hive cluster. Another common source of confusion is when you want to query your data from your cloud object storage, such as AWS S3, MinIO, and Google Cloud Storage. This too uses the Hive connector. If that confuses you, don't worry, you are not alone. This blog aims to explain this commonly confusing nomenclature.

## Hive 3 support in Presto

DevFeed: [Hive 3 support in Presto](<https://devfeed.tech/articles/hive-3-support-in-presto-8631.md>)

Original publisher: [Read original article](<https://trino.io/blog/2019/12/28/hive-3.html>)

Author: Piotr Findeisen, Starburst Data

Published: 2019-12-28T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [data](<https://devfeed.tech/topics/data.md>), [ci](<https://devfeed.tech/topics/ci.md>), [Availability](<https://devfeed.tech/topics/availability.md>)

Tags: [compatibility](<https://devfeed.tech/tags/compatibility.md>), [continuous-integration](<https://devfeed.tech/tags/continuous-integration.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [hdfs](<https://devfeed.tech/tags/hdfs.md>), [integration](<https://devfeed.tech/tags/integration.md>)

### AI overview

This article summarizes Presto's compatibility with Hive 3, including support for Hadoop Erasure Coding, transactional tables, timestamp values stored in ORC, and Hive bucketing v2. It also describes compatibility improvements delivered across Presto releases and ongoing work toward fuller Hive 3 integration.

### Source excerpt

The Hive community is centered around a few different Hive distributions, one of them being Hortonworks Data Platform (HDP). Even after the Cloudera-Hortonworks merger there is vivid interest in HDP 3, featuring Hive 3. Presto is ready for the game. In this post, we summarize which Hive 3 features Presto already supports, covering all the work that went into Presto to achieve that. We also outline next steps lying ahead.

## A Case Study on Anaconda: How it Delivers its Platform

DevFeed: [A Case Study on Anaconda: How it Delivers its Platform](<https://devfeed.tech/articles/a-case-study-on-anaconda-how-it-delivers-its-platform-29597.md>)

Original publisher: [Read original article](<https://goteleport.com/blog/case-study-anaconda/>)

Author: info@goteleport.com (Jon Silvers)

Published: 2019-08-29T00:00:00Z

Content type: article

Language: en

Sources: [Teleport](<https://devfeed.tech/sources/teleport.md>)

Topics: [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [Data Science](<https://devfeed.tech/topics/data-science.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Apache Spark](<https://devfeed.tech/topics/spark.md>)

Tags: [case-study](<https://devfeed.tech/tags/case-study.md>), [cloud-native](<https://devfeed.tech/tags/cloud-native.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [spark](<https://devfeed.tech/tags/spark.md>)

### AI overview

This case study explains why Anaconda selected Teleport's Gravity to package and distribute its Kubernetes-based Anaconda Enterprise application for enterprise customers. It describes Anaconda's need to support cloud, hybrid, and on-premises environments while meeting security and compliance requirements.

### Source excerpt

Learn why Anaconda selected Teleport to package their Kubernetes-based application to Enterprise customers.

## Even Faster ORC

DevFeed: [Even Faster ORC](<https://devfeed.tech/articles/even-faster-orc-8608.md>)

Original publisher: [Read original article](<https://trino.io/blog/2019/04/23/even-faster-orc.html>)

Author: Dain Sundstrom, Martin Traverso

Published: 2019-04-23T00:00:00Z

Content type: article

Language: en

Sources: [Trino Blog](<https://devfeed.tech/sources/trino-blog.md>)

Topics: [Optimization](<https://devfeed.tech/topics/optimization.md>), [Code](<https://devfeed.tech/topics/code.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Programming](<https://devfeed.tech/topics/programming.md>)

Tags: [batch](<https://devfeed.tech/tags/batch.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [performance](<https://devfeed.tech/tags/performance.md>), [sql](<https://devfeed.tech/tags/sql.md>)

### AI overview

This article explains how Trino improved its custom ORC reader. The changes optimize type-specific decoding, bulk reads, null handling, and dynamic dispatch, reducing TPC-DS query time by about 5% and CPU usage by about 9%.

### Source excerpt

Trino is known for being the fastest SQL on Hadoop engine, and our custom ORC reader implementation is a big reason for this speed - now it is even faster!

## How to Setup a Scheduled Scala Spark Job

DevFeed: [How to Setup a Scheduled Scala Spark Job](<https://devfeed.tech/articles/how-to-setup-a-scheduled-scala-spark-job-26528.md>)

Original publisher: [Read original article](<http://engineering.curalate.com/2019/03/27/scheduled-scala-spark-job.html>)

Published: 2019-03-27T00:00:00Z

Content type: tutorial

Language: en

Sources: [Curalate](<https://devfeed.tech/sources/curalate.md>)

Topics: [Apache Spark](<https://devfeed.tech/topics/spark.md>), [Scala](<https://devfeed.tech/topics/scala.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>)

Tags: [aws](<https://devfeed.tech/tags/aws.md>), [cli](<https://devfeed.tech/tags/cli.md>), [daily](<https://devfeed.tech/tags/daily.md>), [data](<https://devfeed.tech/tags/data.md>), [ec2](<https://devfeed.tech/tags/ec2.md>), [emr](<https://devfeed.tech/tags/emr.md>), [github](<https://devfeed.tech/tags/github.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [job](<https://devfeed.tech/tags/job.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [s3](<https://devfeed.tech/tags/s3.md>), [scala](<https://devfeed.tech/tags/scala.md>), [scheduled](<https://devfeed.tech/tags/scheduled.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [spark](<https://devfeed.tech/tags/spark.md>), [weekly](<https://devfeed.tech/tags/weekly.md>)

### AI overview

A tutorial for packaging and deploying a Scala Spark job as a fat JAR, uploading it to Amazon S3, and configuring it to run on a schedule through AWS Data Pipeline and EMR.

### Source excerpt

Have you written a Scala Spark job that processes a massive amount of data on an intimidating amount of RAM and you want to run it daily/weekly/monthly on a schedule on AWS? I had to do this recently, and couldn't find a good tutorial on the full process to get the spark job running. Included in this article and accompanying repository is everything you need to get your Scala Spark job running on AWS Data Pipeline and EMR. Code Repo This tutorial is not going to walk you through the process of actually writing your specific Scala Spark job to do whatever number crunching you need. There are already plenty of resources available (1, 2, 3) to get you started on that. The code template for setting up a Spark Scala job is available in this GitHub repo. Assuming that you have already written your Spark Job and are only using the AWS Java SDK to connect to your AWS data stores, drop your code in the Main function of SparkJob.scala and run the deploy.sh script to upload the fat jar to your S3 bucket. If you do take other dependencies, then it may take some extra work on your part. To run a Scala Spark job on AWS you need to compile a fat jar that contains the byte code for your job and all of the libraries it needs to run. This project already has the sbt-assembly plugin setup and a assemblyMergeStrategy set up to package the Spark, Hadoop, and AWS SDK together in the fat jar. If you need to add in other libraries that do not play well with each other, or are using a noncompatible version of Spark for this current repo, there are a few good resources available to help you through the needed build.sbt modifications. Outside of the previously mentioned needed changes you need to set a few parameters in the deploy.sh script. Mainly the deploymentPath to your specific S3 bucket, adding a profile to the AWS CLI command to upload to your specific S3 bucket if it's private, and changing the resulting fat jar name if you please. The deploy script uses the AWS CLI to upload the fat

## Announcing OpenTSDB 2.4.0: Rollup and Pre-Aggregation Storage, Histograms, Sketches, and More

DevFeed: [Announcing OpenTSDB 2.4.0: Rollup and Pre-Aggregation Storage, Histograms, Sketches, and More](<https://devfeed.tech/articles/announcing-opentsdb-2-4-0-rollup-and-pre-aggregation-storage-histograms-sketches-and-more-20496.md>)

Original publisher: [Read original article](<https://yahooeng.tumblr.com/post/181461332311>)

Author: amberwilsonla-blog

Published: 2018-12-27T17:01:17Z

Content type: release

Language: en

Sources: [Yahoo](<https://devfeed.tech/sources/yahoo.md>)

Topics: [Databases](<https://devfeed.tech/topics/databases.md>), [Time Series](<https://devfeed.tech/topics/time-series.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [Data analysis](<https://devfeed.tech/topics/data-analysis.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Temporal data](<https://devfeed.tech/topics/temporal-data.md>)

Tags: [data-analysis](<https://devfeed.tech/tags/data-analysis.md>), [databases](<https://devfeed.tech/tags/databases.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [release](<https://devfeed.tech/tags/release.md>), [salesforce](<https://devfeed.tech/tags/salesforce.md>), [time-series](<https://devfeed.tech/tags/time-series.md>), [workflows](<https://devfeed.tech/tags/workflows.md>), [yahoo-engineering](<https://devfeed.tech/tags/yahoo-engineering.md>)

### AI overview

OpenTSDB 2.4.0 is released with rollup and pre-aggregation storage for time series data, along with histograms and sketches. The release focuses on retaining lower-resolution data for longer periods, reducing query workloads, and supporting more accurate percentile analysis across multiple series.

### Source excerpt

yahoodevelopers: By Chris Larsen, Architect OpenTSDB is one of the first dedicated open source time series databases built on top of Apache HBase and the Hadoop Distributed File System. Today, we are proud to share that version 2.4.0 is now available and has many new features developed in-house and with contributions from the open source community. This release would not have been possible without support from our monitoring team, the Hadoop and HBase developers, as well as contributors from other companies like Salesforce, Alibaba, JD.com, Arista and more. Thank you to everyone who contributed to this release! A few of the exciting new features include: Rollup and Pre-Aggregation Storage As time series data grows, storing the original measurements becomes expensive. Particularly in the case of monitoring workflows, users rarely care about last years' high fidelity data. It's more efficient to store lower resolution "rollups" for longer periods, discarding the original high-resolution data. OpenTSDB now supports storing and querying such data so that the raw data can expire from HBase or Bigtable, and the rollups can stick around longer. Querying for long time ranges will read from the lower resolution data, fetching fewer data points and speeding up queries. Likewise, when a user wants to query tens of thousands of time series grouped by, for example, data centers, the TSD will have to fetch and process a significant amount of data, making queries painfully slow. To improve query speed, pre-aggregated data can be stored and queried to fetch much less data at query time, while still retaining the raw data. We have an Apache Storm pipeline that computes these rollups and pre-aggregates, and we intend to open source that code in 2019. For more details, please visit http://opentsdb.net/docs/build/html/user_guide/rollups.html. Histograms and Sketches When monitoring or performing data analysis, users often like to explore percentiles of their measurements, such as the 9

## Omid 1.0.0 adds low-latency transaction processing for Apache Phoenix

DevFeed: [Omid 1.0.0 adds low-latency transaction processing for Apache Phoenix](<https://devfeed.tech/articles/a-new-chapter-for-omid-20494.md>)

Original publisher: [Read original article](<https://yahooeng.tumblr.com/post/180867271141>)

Author: amberwilsonla-blog

Published: 2018-12-06T19:32:04Z

Content type: release

Language: en

Sources: [Yahoo](<https://devfeed.tech/sources/yahoo.md>)

Topics: [phoenix](<https://devfeed.tech/topics/phoenix.md>), [Transactions](<https://devfeed.tech/topics/transactions.md>), [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [SQL](<https://devfeed.tech/topics/sql.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [big-data](<https://devfeed.tech/tags/big-data.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [phoenix](<https://devfeed.tech/tags/phoenix.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speed](<https://devfeed.tech/tags/speed.md>), [sql](<https://devfeed.tech/tags/sql.md>), [transactions](<https://devfeed.tech/tags/transactions.md>), [yahoo-engineering](<https://devfeed.tech/tags/yahoo-engineering.md>)

### AI overview

This release article introduces Omid 1.0.0, an open source transaction processing platform for Big Data, and announces its selection as the transaction management provider for Apache Phoenix. It describes the Omid Low Latency protocol, which reduces short-transaction latency by 5 times under light load and by 10 to 100 times under heavy load.

### Source excerpt

yahoodevelopers: By Ohad Shacham, Yonatan Gottesman, Edward Bortnikov Scalable Systems Research, Verizon/Oath Omid, an open source transaction processing platform for Big Data, was born as a research project at Yahoo (now part of Verizon), and became an Apache Incubator project in 2015. Omid complements Apache HBase, a distributed key-value store in Apache Hadoop suite, with a capability to clip multiple operations into logically indivisible (atomic) units named transactions. This programming model has been extremely popular since the dawn of SQL databases, and has more recently become indispensable in the NoSQL world. For example, it is the centerpiece for dynamic content indexing of search and media products at Verizon, powering a web-scale content management platform since 2015. Today, we are excited to share a new chapter in Omid's history. Thanks to its scalability, reliability, and speed, Omid has been selected as transaction management provider for Apache Phoenix, a real-time converged OLTP and analytics platform for Hadoop. Phoenix provides a standard SQL interface to HBase key-value storage, which is much simpler and in many cases more performant than the native HBase API. With Phoenix, big data and machine learning developers get the best of all worlds: increased productivity coupled with high scalability. Phoenix is designed to scale to 10,000 query processing nodes in one instance and is expected to process hundreds of thousands or even millions of transactions per second (tps). It is widely used in the industry, including by Alibaba, Bloomberg, PubMatic, Salesforce, Sogou and many others. We have just released a new and significantly improved version of Omid (1.0.0), the first major release since its original launch. We have extended the system with multiple functional and performance features to power a modern SQL database technology, ready for deployment on both private and public cloud platforms. A few of the significant innovations include: Protocol

## Hadoop Contributors Meetup at Oath

DevFeed: [Hadoop Contributors Meetup at Oath](<https://devfeed.tech/articles/hadoop-contributors-meetup-at-oath-20493.md>)

Original publisher: [Read original article](<https://yahooeng.tumblr.com/post/179901430546>)

Author: amberwilsonla-blog

Published: 2018-11-08T19:07:25Z

Content type: news

Language: en

Sources: [Yahoo](<https://devfeed.tech/sources/yahoo.md>)

Topics: [Hadoop](<https://devfeed.tech/topics/hadoop.md>), [Development](<https://devfeed.tech/topics/development.md>), [Containers](<https://devfeed.tech/topics/containers.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [big-data](<https://devfeed.tech/tags/big-data.md>), [containers](<https://devfeed.tech/tags/containers.md>), [development](<https://devfeed.tech/tags/development.md>), [docker](<https://devfeed.tech/tags/docker.md>), [hadoop](<https://devfeed.tech/tags/hadoop.md>), [low-latency](<https://devfeed.tech/tags/low-latency.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [yahoo-engineering](<https://devfeed.tech/tags/yahoo-engineering.md>)

### AI overview

Oath hosted a day-long Hadoop Contributors Meetup at its Sunnyvale campus, bringing together more than 80 Hadoop users, contributors, committers, and PMC members. Talks and breakout sessions covered challenges and solutions across the Hadoop ecosystem, including HDFS, Tez, containers, low-latency processing, deep learning, scalability, security, and cloud deployment.

### Source excerpt

yahoodevelopers: By Scott Bush, Director, Hadoop Software Engineering, Oath On Tuesday, September 25, we hosted a special day-long Hadoop Contributors Meetup at our Sunnyvale, California campus. Much of the early Hadoop development work started at Yahoo, now part of Oath, and has continued over the past decade. Our campus was the perfect setting for this meetup, as we continue to make Hadoop a priority. More than 80 Hadoop users, contributors, committers, and PMC members gathered to hear talks on key issues facing the Hadoop user community. Speakers from Ampool, Cloudera, Hortonworks, Microsoft, Oath, and Twitter detailed some of the challenges and solutions pertinent to their parts of the Hadoop ecosystem. The talks were followed by a number of parallel, birds of a feather breakout sessions to discuss HDFS, Tez, containers and low latency processing. The day ended with a reception and consensus that the event went well and should be repeated in the near future. Presentation recordings (YouTube playlist) and slides (links included in the video description) are available here: Hadoop {Submarine} Project: Running deep learning workloads on YARN, Wangda Tan, Hortonworks Apache YARN Federation and Tez at Microsoft, Anupam Upadhyay, Adrian Nicoara, Botong Huang "HDFS Scalability and Security", Daryn Sharp, Senior Engineer, Oath The Future of Hadoop in an AI World, Milind Bhandarkar, CEO, Ampool Moving the Oath Grid to Docker, Eric Badger, Software Developer Engineer, Oath Vespa: Open Source Big Data Serving Engine, Jon Bratseth, Distinguished Architect, Oath Containerized Services on Apache Hadoop YARN: Past, Present, and Future, Shane Kumpf, Hortonworks How Twitter Hadoop Chose Google Cloud, Joep Rottinghuis, Lohit VijayaRenu Thank you to all the presenters and the attendees both in person and remote! P.S. We're hiring! Learn more about career opportunities at Oath.

[Next page](<https://devfeed.tech/tags/hadoop.md?cursor=WyIyMDE4LTExLTA4VDE5OjA3OjI1KzAwOjAwIiwgIjVlMDI1OTZkLTdlNDgtNDI5ZS1iZTAzLTgyNTljMjFlMTgxMSJd>)