# xet

Xet is a custom storage system built for AI/ML development that provides chunk-level deduplication, smaller uploads, and faster downloads than Git LFS.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Introducing Storage Buckets on the Hugging Face Hub

DevFeed: [Introducing Storage Buckets on the Hugging Face Hub](<https://devfeed.tech/articles/introducing-storage-buckets-on-the-hugging-face-hub-7492.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/storage-buckets>)

Author: Lucain Pouget; Eliott Coyac; Adrien Carreira; Victor Mustar; Julien Chaumond; Quentin Lhoest; Pierric Cistac; Sylvestre Bcht; Hugo Larcher; Rajat Arya

Published: 2026-03-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [xet](<https://devfeed.tech/topics/xet.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [storage](<https://devfeed.tech/tags/storage.md>), [training](<https://devfeed.tech/tags/training.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

The article introduces Storage Buckets on the Hugging Face Hub, a mutable S3-like storage option for machine-learning artifacts. Built on Xet, Buckets use chunking and deduplication to reduce redundant transfers, storage use, and enterprise billing, while pre-warming places frequently accessed data near distributed-training compute.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Building for an Open Future - our new partnership with Google Cloud

DevFeed: [Building for an Open Future - our new partnership with Google Cloud](<https://devfeed.tech/articles/building-for-an-open-future-our-new-partnership-with-google-cloud-7218.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/google-cloud>)

Author: Jeff Boudier; Simon Pagezy

Published: 2025-11-13T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Cloud Run](<https://devfeed.tech/topics/cloud-run.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [xet](<https://devfeed.tech/topics/xet.md>), [data](<https://devfeed.tech/topics/data.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cache](<https://devfeed.tech/tags/cache.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [google](<https://devfeed.tech/tags/google.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [networking](<https://devfeed.tech/tags/networking.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [storage](<https://devfeed.tech/tags/storage.md>), [vertex](<https://devfeed.tech/tags/vertex.md>), [vertex-ai](<https://devfeed.tech/tags/vertex-ai.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

Hugging Face and Google Cloud announce a strategic partnership focused on making open AI models easier to use, customize, deploy, and govern. The article describes integrations across Vertex AI, GKE AI/ML, Cloud Run GPUs, and other Google Cloud infrastructure, plus a planned CDN Gateway using Hugging Face Xet and Google Cloud storage and networking to accelerate model and dataset downloads and improve supply-chain robustness.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Parquet Content-Defined Chunking

DevFeed: [Parquet Content-Defined Chunking](<https://devfeed.tech/articles/parquet-content-defined-chunking-7438.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/parquet-cdc>)

Author: Krisztian Szucs

Published: 2025-07-25T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [parquet](<https://devfeed.tech/topics/parquet.md>), [content defined chunking](<https://devfeed.tech/topics/content-defined-chunking.md>), [xet](<https://devfeed.tech/topics/xet.md>), [data-engineering](<https://devfeed.tech/topics/data-engineering.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [content-defined-chunking](<https://devfeed.tech/tags/content-defined-chunking.md>), [data](<https://devfeed.tech/tags/data.md>), [data-engineering](<https://devfeed.tech/tags/data-engineering.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [dedupe](<https://devfeed.tech/tags/dedupe.md>), [hub](<https://devfeed.tech/tags/hub.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [performance](<https://devfeed.tech/tags/performance.md>), [storage](<https://devfeed.tech/tags/storage.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

This article explains how Parquet Content-Defined Chunking (CDC) in PyArrow and Pandas improves deduplication of Parquet files on content-addressable storage systems such as Hugging Face Xet. It describes how CDC reduces data transfer and storage costs by identifying and reusing unchanged data chunks, and demonstrates the behavior across common table modifications.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Migrating the Hub from Git LFS to Xet

DevFeed: [Migrating the Hub from Git LFS to Xet](<https://devfeed.tech/articles/migrating-the-hub-from-git-lfs-to-xet-7351.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/migrating-the-hub-to-xet>)

Author: Jared Sulzdorf; Joseph Godlewski; Sam Horradarn

Published: 2025-07-15T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [xet](<https://devfeed.tech/topics/xet.md>), [migration](<https://devfeed.tech/topics/migration.md>), [content addressed store](<https://devfeed.tech/topics/content-addressed-store.md>), [content defined chunking](<https://devfeed.tech/topics/content-defined-chunking.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>), [Git](<https://devfeed.tech/topics/git.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [cas](<https://devfeed.tech/tags/cas.md>), [content-addressed-store](<https://devfeed.tech/tags/content-addressed-store.md>), [content-defined-chunking](<https://devfeed.tech/tags/content-defined-chunking.md>), [git](<https://devfeed.tech/tags/git.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [integration](<https://devfeed.tech/tags/integration.md>), [migration](<https://devfeed.tech/tags/migration.md>), [s3](<https://devfeed.tech/tags/s3.md>), [storage](<https://devfeed.tech/tags/storage.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

The article explains Hugging Face's migration of the Hub from Git LFS to Xet. It describes the Git LFS Bridge, background content migrations, content-defined chunking, the content addressed store, and S3-backed storage that enable gradual, large-scale migration without disrupting users.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Xet is on the Hub

DevFeed: [Xet is on the Hub](<https://devfeed.tech/articles/xet-is-on-the-hub-7570.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/xet-on-the-hub>)

Author: Assaf Vayner; Brian Ronan; Di Xiao; Joseph Godlewski; Sam Horradarn; Jared Sulzdorf

Published: 2025-03-18T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [xet](<https://devfeed.tech/topics/xet.md>), [content defined chunking](<https://devfeed.tech/topics/content-defined-chunking.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [migration](<https://devfeed.tech/topics/migration.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [SQLite](<https://devfeed.tech/topics/sqlite.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [content-addressed-store](<https://devfeed.tech/tags/content-addressed-store.md>), [content-defined-chunking](<https://devfeed.tech/tags/content-defined-chunking.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [migration](<https://devfeed.tech/tags/migration.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [performance](<https://devfeed.tech/tags/performance.md>), [sqlite](<https://devfeed.tech/tags/sqlite.md>), [storage](<https://devfeed.tech/tags/storage.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

Hugging Face describes migrating the first Model and Dataset repositories from LFS to Xet storage. The article explains how content-defined chunking enables byte-level deduplication, reducing the data transferred for small edits to large files and improving upload performance.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## From Chunks to Blocks: Accelerating Uploads and Downloads on the Hub

DevFeed: [From Chunks to Blocks: Accelerating Uploads and Downloads on the Hub](<https://devfeed.tech/articles/from-chunks-to-blocks-accelerating-uploads-and-downloads-on-the-hub-7205.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/from-chunks-to-blocks>)

Author: Jared Sulzdorf; yuchenglow; Zach Nation; saba noorassa

Published: 2025-02-12T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [xet](<https://devfeed.tech/topics/xet.md>), [content addressed store](<https://devfeed.tech/topics/content-addressed-store.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Amazon S3](<https://devfeed.tech/topics/amazon-s3.md>)

Tags: [cas](<https://devfeed.tech/tags/cas.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [content-addressed-store](<https://devfeed.tech/tags/content-addressed-store.md>), [content-defined-chunking](<https://devfeed.tech/tags/content-defined-chunking.md>), [dedupe](<https://devfeed.tech/tags/dedupe.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [integration](<https://devfeed.tech/tags/integration.md>), [network](<https://devfeed.tech/tags/network.md>), [performance](<https://devfeed.tech/tags/performance.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [rust](<https://devfeed.tech/tags/rust.md>), [s3](<https://devfeed.tech/tags/s3.md>), [storage](<https://devfeed.tech/tags/storage.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

Hugging Face's Xet team explains how content-defined chunking is being adapted for production to accelerate uploads and downloads on the Hub. The article describes the trade-offs of fine-grained deduplication, including network, infrastructure, metadata, and storage costs, and introduces a Rust-based chunk-oriented integration designed to improve experimentation and collaboration on models and datasets.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## From Files to Chunks: Improving HF Storage Efficiency

DevFeed: [From Files to Chunks: Improving HF Storage Efficiency](<https://devfeed.tech/articles/from-files-to-chunks-improving-hf-storage-efficiency-7206.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/from-files-to-chunks>)

Author: Jared Sulzdorf; Ann Huang

Published: 2024-11-20T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [xet](<https://devfeed.tech/topics/xet.md>), [content defined chunking](<https://devfeed.tech/topics/content-defined-chunking.md>), [content addressed store](<https://devfeed.tech/topics/content-addressed-store.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Git](<https://devfeed.tech/topics/git.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [content-addressed-store](<https://devfeed.tech/tags/content-addressed-store.md>), [content-defined-chunking](<https://devfeed.tech/tags/content-defined-chunking.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [dedupe](<https://devfeed.tech/tags/dedupe.md>), [git](<https://devfeed.tech/tags/git.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [storage](<https://devfeed.tech/tags/storage.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

Hugging Face's Xet team describes a storage system that splits files into variable-sized chunks using content-defined chunking and a rolling hash. Chunks are stored in a content-addressed store with deduplication, so updates upload only new chunks. The article reports a consistent 50% improvement in storage and transfer performance compared with Git LFS across three iterative development use cases, including the CORD-19 dataset.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Share your open ML datasets on Hugging Face Hub!

DevFeed: [Share your open ML datasets on Hugging Face Hub!](<https://devfeed.tech/articles/share-your-open-ml-datasets-on-hugging-face-hub-7457.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/researcher-dataset-sharing>)

Author: Daniel van Strien; Caleb Fahlgren; Quentin Lhoest; Ann Huang

Published: 2024-11-12T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [hosting](<https://devfeed.tech/topics/hosting.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [CSV](<https://devfeed.tech/topics/csv.md>), [JSON](<https://devfeed.tech/topics/json.md>), [xet](<https://devfeed.tech/topics/xet.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [community](<https://devfeed.tech/tags/community.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [guide](<https://devfeed.tech/tags/guide.md>), [hosting](<https://devfeed.tech/tags/hosting.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [json](<https://devfeed.tech/tags/json.md>), [ml](<https://devfeed.tech/tags/ml.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [research](<https://devfeed.tech/tags/research.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

This article explains how to share open machine learning datasets through Hugging Face Hub. It describes large-scale hosting, dataset uploads and downloads with the Datasets library, streaming for resource-constrained users, browser-based exploration, full-text search, sorting, supported modalities and file formats, compression, and planned Xet improvements to storage and transfer limits.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## XetHub is joining Hugging Face!

DevFeed: [XetHub is joining Hugging Face!](<https://devfeed.tech/articles/xethub-is-joining-hugging-face-7571.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/xethub-joins-hf>)

Author: yuchenglow; Julien Chaumond

Published: 2024-08-08T00:00:00Z

Content type: news

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [xet](<https://devfeed.tech/topics/xet.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Git](<https://devfeed.tech/topics/git.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [parquet](<https://devfeed.tech/topics/parquet.md>), [llama](<https://devfeed.tech/topics/llama.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai-development](<https://devfeed.tech/tags/ai-development.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [git](<https://devfeed.tech/tags/git.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [llama](<https://devfeed.tech/tags/llama.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [parquet](<https://devfeed.tech/tags/parquet.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

Hugging Face announces that XetHub is joining the organization to improve storage and versioning for large AI datasets and models. XetHub's technology uses chunking and deduplication so updates can upload only changed portions instead of entire large files, while helping Git scale to very large repositories.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.