# optimum

Published articles for optimum.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Scaling a Vespa Application: Feeding Fast and Furiously

DevFeed: [Scaling a Vespa Application: Feeding Fast and Furiously](<https://devfeed.tech/articles/scaling-a-vespa-application-feeding-fast-and-furiously-12797.md>)

Original publisher: [Read original article](<https://blog.vespa.ai/scaling-a-vespa-application-feeding-fast-and-furiously/>)

Author: Kai Borgen

Published: 2026-04-28T00:00:00Z

Content type: tutorial

Language: en

Sources: [Vespa Blog](<https://devfeed.tech/sources/vespa-blog.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [information retrieval](<https://devfeed.tech/topics/information-retrieval.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Homebrew](<https://devfeed.tech/topics/homebrew.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [onnx](<https://devfeed.tech/topics/onnx.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [XML](<https://devfeed.tech/topics/xml.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-search](<https://devfeed.tech/tags/ai-search.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [genai](<https://devfeed.tech/tags/genai.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [information-retrieval](<https://devfeed.tech/tags/information-retrieval.md>), [install](<https://devfeed.tech/tags/install.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [onnx](<https://devfeed.tech/tags/onnx.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [performance](<https://devfeed.tech/tags/performance.md>), [rag](<https://devfeed.tech/tags/rag.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [tensors](<https://devfeed.tech/tags/tensors.md>), [tutorial](<https://devfeed.tech/tags/tutorial.md>)

### AI overview

This tutorial demonstrates how to scale a Vespa application while feeding the full MS_marco passages dataset. It covers preparing the dataset, configuring access, deploying a sample application, and using scaling and metrics to improve feed throughput and performance.

### Source excerpt

A tutorial on how to scale the resources in a Vespa application to increase feed throughput. Using the metrics dashboard for informed and optimised scaling.

## Get your VLM running in 3 simple steps on Intel CPUs

DevFeed: [Get your VLM running in 3 simple steps on Intel CPUs](<https://devfeed.tech/articles/get-your-vlm-running-in-3-simple-steps-on-intel-cpus-7431.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/openvino-vlm>)

Author: Ezequiel Lanza; Helena; Nikita; Ella Charlaix; Ilyas Moutawwakil

Published: 2025-10-15T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [smolvlm](<https://devfeed.tech/topics/smolvlm.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [intel](<https://devfeed.tech/topics/intel.md>)

Tags: [inference](<https://devfeed.tech/tags/inference.md>), [intel](<https://devfeed.tech/tags/intel.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [smolvlm](<https://devfeed.tech/tags/smolvlm.md>), [vlm](<https://devfeed.tech/tags/vlm.md>)

### AI overview

A tutorial explains how to run SmolVLM locally with Optimum Intel and OpenVINO, then optimize it for lower memory use and faster inference through quantization.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Introducing the AMD 5th Gen EPYC™ CPU

DevFeed: [Introducing the AMD 5th Gen EPYC™ CPU](<https://devfeed.tech/articles/introducing-the-amd-5th-gen-epyctm-cpu-7246.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/huggingface-amd-turin>)

Author: Mohit Sharma; Morgan Funtowicz

Published: 2024-10-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [cpu](<https://devfeed.tech/topics/cpu.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Keras](<https://devfeed.tech/topics/keras.md>)

Tags: [amd](<https://devfeed.tech/tags/amd.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [ecosystem](<https://devfeed.tech/tags/ecosystem.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [meta](<https://devfeed.tech/tags/meta.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [speed](<https://devfeed.tech/tags/speed.md>), [text-generation](<https://devfeed.tech/tags/text-generation.md>), [tokens](<https://devfeed.tech/tags/tokens.md>)

### AI overview

Hugging Face evaluates AMD's 5th Gen EPYC CPUs, Turin and Genoa, for large language model and RAG workloads. The article presents benchmark results using PyTorch inference optimizations and multiple Meta LLaMA 3.1 8B instances across summarization, chatbot, translation, essay writing, and live captioning, comparing decode throughput at different batch sizes.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Optimize and deploy with Optimum-Intel and OpenVINO GenAI

DevFeed: [Optimize and deploy with Optimum-Intel and OpenVINO GenAI](<https://devfeed.tech/articles/optimize-and-deploy-with-optimum-intel-and-openvino-genai-7167.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/deploy-with-openvino>)

Author: Alexander; Yury Gorbachev; Ekaterina Aidova; Ilya Lavrenov; Raymond Lo (NVIDIA); Helena; Ella Charlaix

Published: 2024-09-20T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [C++](<https://devfeed.tech/topics/c-plus-plus.md>)

Tags: [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [api](<https://devfeed.tech/tags/api.md>), [c-plus-plus](<https://devfeed.tech/tags/c-plus-plus.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [edge](<https://devfeed.tech/tags/edge.md>), [inference](<https://devfeed.tech/tags/inference.md>), [intel](<https://devfeed.tech/tags/intel.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [python](<https://devfeed.tech/tags/python.md>), [quantization](<https://devfeed.tech/tags/quantization.md>)

### AI overview

A tutorial on exporting Transformer models to OpenVINO IR, optimizing LLMs with weight-only quantization, and deploying them through the OpenVINO GenAI API for edge and client devices.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Memory-efficient Diffusion Transformers with Quanto and Diffusers

DevFeed: [Memory-efficient Diffusion Transformers with Quanto and Diffusers](<https://devfeed.tech/articles/memory-efficient-diffusion-transformers-with-quanto-and-diffusers-7450.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/quanto-diffusers>)

Author: Sayak Paul; David Corvoysier

Published: 2024-07-30T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [diffusers](<https://devfeed.tech/topics/diffusers.md>), [diffusion-transformers](<https://devfeed.tech/topics/diffusion-transformers.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>), [stable-diffusion](<https://devfeed.tech/topics/stable-diffusion.md>), [text-to-image](<https://devfeed.tech/topics/text-to-image.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>)

Tags: [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [diffusers](<https://devfeed.tech/tags/diffusers.md>), [diffusion-transformers](<https://devfeed.tech/tags/diffusion-transformers.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [quality](<https://devfeed.tech/tags/quality.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [sd3](<https://devfeed.tech/tags/sd3.md>), [text-to-image](<https://devfeed.tech/tags/text-to-image.md>)

### AI overview

This tutorial explains how to reduce the memory requirements of Transformer-based diffusion pipelines using Quanto quantization utilities from the Diffusers library. It benchmarks FP8 quantization on PixArt-Sigma, Stable Diffusion 3, and Aura Flow, reporting memory savings with slightly higher latency and little quality degradation.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Google Cloud TPUs made available to Hugging Face users

DevFeed: [Google Cloud TPUs made available to Hugging Face users](<https://devfeed.tech/articles/google-cloud-tpus-made-available-to-hugging-face-users-7524.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tpu-inference-endpoints-spaces>)

Author: Simon Pagezy; Michelle Habonneau; Philipp Schmid; Alvaro Moran

Published: 2024-07-09T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Sovereign AI](<https://devfeed.tech/topics/sovereign-ai.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [spaces](<https://devfeed.tech/topics/spaces.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [gemma](<https://devfeed.tech/topics/gemma.md>), [llama](<https://devfeed.tech/topics/llama.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>)

Tags: [gcp](<https://devfeed.tech/tags/gcp.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llama](<https://devfeed.tech/tags/llama.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [performance](<https://devfeed.tech/tags/performance.md>), [spaces](<https://devfeed.tech/tags/spaces.md>), [tgi](<https://devfeed.tech/tags/tgi.md>), [tpu](<https://devfeed.tech/tags/tpu.md>)

### AI overview

Hugging Face announces that Google Cloud TPUs are available for Inference Endpoints and Spaces. Google TPU v5e configurations can deploy supported models through managed infrastructure, while Optimum TPU and Text Generation Inference help train and serve models on TPUs.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Accelerating Protein Language Model ProtST on Intel Gaudi 2

DevFeed: [Accelerating Protein Language Model ProtST on Intel Gaudi 2](<https://devfeed.tech/articles/accelerating-protein-language-model-protst-on-intel-gaudi-2-7291.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/intel-protein-language-model-protst>)

Author: Julien Simon; Jiqing.Feng; Santiago Miret; Xinyu Yuan; Yi Wang; Matrix Yao; Minghao Xu; Ke Ding

Published: 2024-07-03T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [intel](<https://devfeed.tech/topics/intel.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [accelerators](<https://devfeed.tech/tags/accelerators.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [batch](<https://devfeed.tech/tags/batch.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [intel](<https://devfeed.tech/tags/intel.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llm](<https://devfeed.tech/tags/llm.md>), [model](<https://devfeed.tech/tags/model.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [pcie](<https://devfeed.tech/tags/pcie.md>), [precision](<https://devfeed.tech/tags/precision.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

This tutorial explains how to run inference and fine-tune ProtST, a multimodal protein language model, using Intel Gaudi 2 accelerators and the Optimum for Intel Gaudi open-source library. It compares ProtST inference on NVIDIA A100 and Gaudi 2, reporting identical accuracy and 1.76x faster inference on Gaudi 2.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Faster assisted generation support for Intel Gaudi

DevFeed: [Faster assisted generation support for Intel Gaudi](<https://devfeed.tech/articles/faster-assisted-generation-support-for-intel-gaudi-7108.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/assisted-generation-support-gaudi>)

Author: Haim Barad; Neha Raste; Tien Pei Chou

Published: 2024-06-04T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [intel](<https://devfeed.tech/topics/intel.md>), [text-generation](<https://devfeed.tech/topics/text-generation.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [Transformer](<https://devfeed.tech/topics/transformer.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>)

Tags: [accelerate](<https://devfeed.tech/tags/accelerate.md>), [ai](<https://devfeed.tech/tags/ai.md>), [diffusers](<https://devfeed.tech/tags/diffusers.md>), [gaudi](<https://devfeed.tech/tags/gaudi.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [intel](<https://devfeed.tech/tags/intel.md>), [latency](<https://devfeed.tech/tags/latency.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [techniques](<https://devfeed.tech/tags/techniques.md>), [text-generation](<https://devfeed.tech/tags/text-generation.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

The article explains how assisted decoding and speculative sampling were adapted and optimized for Intel Gaudi processors. Integrated into Optimum Habana, these techniques use draft and target models, KV caching, and quantized models to accelerate text generation while preserving the target model's sampling quality.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Deploy models on AWS Inferentia2 from Hugging Face

DevFeed: [Deploy models on AWS Inferentia2 from Hugging Face](<https://devfeed.tech/articles/deploy-models-on-aws-inferentia2-from-hugging-face-7285.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/inferentia-inference-endpoints>)

Author: Jeff Boudier; Philipp Schmid

Published: 2024-05-22T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AWS AI chips](<https://devfeed.tech/topics/aws-ai-chips.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Amazon SageMaker](<https://devfeed.tech/topics/amazon-sagemaker.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [MLOps](<https://devfeed.tech/topics/mlops.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-trainium](<https://devfeed.tech/tags/aws-trainium.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [llama3](<https://devfeed.tech/tags/llama3.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [model](<https://devfeed.tech/tags/model.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [tgi](<https://devfeed.tech/tags/tgi.md>)

### AI overview

Hugging Face announces support for deploying models on AWS Inferentia2 through Amazon SageMaker and Hugging Face Inference Endpoints. The update enables scalable inference for supported models, including Meta Llama 3, with multiple instance sizes, managed features, and autoscaling.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Hugging Face on AMD Instinct MI300 GPU

DevFeed: [Hugging Face on AMD Instinct MI300 GPU](<https://devfeed.tech/articles/hugging-face-on-amd-instinct-mi300-gpu-7245.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/huggingface-amd-mi300>)

Author: Félix Marty; Mohit Sharma; seungrok jung; Morgan Funtowicz

Published: 2024-05-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [CI/CD](<https://devfeed.tech/topics/cicd.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Kubernetes](<https://devfeed.tech/topics/kubernetes.md>), [text-generation](<https://devfeed.tech/topics/text-generation.md>), [migration](<https://devfeed.tech/topics/migration.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [amd](<https://devfeed.tech/tags/amd.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [azure](<https://devfeed.tech/tags/azure.md>), [ci-cd](<https://devfeed.tech/tags/ci-cd.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [instinct](<https://devfeed.tech/tags/instinct.md>), [kubernetes](<https://devfeed.tech/tags/kubernetes.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [migration](<https://devfeed.tech/tags/migration.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [performance](<https://devfeed.tech/tags/performance.md>), [rocm](<https://devfeed.tech/tags/rocm.md>), [text-generation](<https://devfeed.tech/tags/text-generation.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

Hugging Face and AMD describe first-class integration for AMD Instinct MI300 GPU servers across the Hugging Face Platform. The article covers deployment from local development to Azure ND MI300x V5 VMs, compatibility with existing libraries and products, and CI/CD testing on managed Kubernetes infrastructure.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## From cloud to developers: Hugging Face and Microsoft Deepen Collaboration

DevFeed: [From cloud to developers: Hugging Face and Microsoft Deepen Collaboration](<https://devfeed.tech/articles/from-cloud-to-developers-hugging-face-and-microsoft-deepen-collaboration-7348.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/microsoft-collaboration>)

Author: Jeff Boudier; Philipp Schmid

Published: 2024-05-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Microsoft](<https://devfeed.tech/topics/microsoft.md>), [Azure](<https://devfeed.tech/topics/azure.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [rocm](<https://devfeed.tech/topics/rocm.md>), [llama](<https://devfeed.tech/topics/llama.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [azure](<https://devfeed.tech/tags/azure.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cohere](<https://devfeed.tech/tags/cohere.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [meta](<https://devfeed.tech/tags/meta.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [rocm](<https://devfeed.tech/tags/rocm.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Hugging Face and Microsoft expand their collaboration to make open AI models easier to deploy and run on Azure. The article highlights one-click deployment through Azure AI Studio, support for models such as Llama 3, Command R Plus, Qwen 1.5 110B, and Phi-3, and optimization for AMD Instinct MI300X virtual machines using Optimum-AMD and ROCm.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Blazing Fast SetFit Inference with 🤗 Optimum Intel on Xeon

DevFeed: [Blazing Fast SetFit Inference with 🤗 Optimum Intel on Xeon](<https://devfeed.tech/articles/blazing-fast-setfit-inference-with-optimum-intel-on-xeon-7473.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/setfit-optimum-intel>)

Author: Daniel Korat; Tom Aarsen; Oren Pereg; Moshe Wasserblat; Ella Charlaix; Abirami Prabhakaran

Published: 2024-04-03T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Post-training optimization](<https://devfeed.tech/topics/post-training-optimization.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [community](<https://devfeed.tech/tags/community.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [inference](<https://devfeed.tech/tags/inference.md>), [intel](<https://devfeed.tech/tags/intel.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-collab](<https://devfeed.tech/tags/open-source-collab.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [quantization](<https://devfeed.tech/tags/quantization.md>)

### AI overview

A tutorial on accelerating SetFit inference on Intel Xeon CPUs with Optimum Intel and post-training quantization.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.