# inference-endpoints

Published articles for inference-endpoints.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## AI-Ready Private Cloud with Cisco and VMware

DevFeed: [AI-Ready Private Cloud with Cisco and VMware](<https://devfeed.tech/articles/ai-ready-private-cloud-with-cisco-and-vmware-12808.md>)

Original publisher: [Read original article](<https://blogs.vmware.com/cloud-foundation/2026/09/08/ai-ready-private-cloud-with-cisco-and-vmware/>)

Author: sabina anja

Published: 2026-09-08T15:33:39Z

Content type: article

Language: en

Sources: [VMware Blogs](<https://devfeed.tech/sources/vmware-blogs.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Network](<https://devfeed.tech/topics/network.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [networking](<https://devfeed.tech/topics/networking.md>), [cloud-infrastructure](<https://devfeed.tech/topics/cloud-infrastructure.md>), [distributed-training](<https://devfeed.tech/topics/distributed-training.md>), [datacenter](<https://devfeed.tech/topics/datacenter.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [cisco](<https://devfeed.tech/tags/cisco.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cloud-infrastructure](<https://devfeed.tech/tags/cloud-infrastructure.md>), [cloud-platform](<https://devfeed.tech/tags/cloud-platform.md>), [data-center](<https://devfeed.tech/tags/data-center.md>), [distributed-training](<https://devfeed.tech/tags/distributed-training.md>), [fabric](<https://devfeed.tech/tags/fabric.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [home-page](<https://devfeed.tech/tags/home-page.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [latency](<https://devfeed.tech/tags/latency.md>), [networking](<https://devfeed.tech/tags/networking.md>), [object-storage](<https://devfeed.tech/tags/object-storage.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [private-cloud](<https://devfeed.tech/tags/private-cloud.md>), [telemetry](<https://devfeed.tech/tags/telemetry.md>), [vcf-9-1](<https://devfeed.tech/tags/vcf-9-1.md>), [vcf-networking](<https://devfeed.tech/tags/vcf-networking.md>), [vmware](<https://devfeed.tech/tags/vmware.md>), [vmware-cloud-foundation](<https://devfeed.tech/tags/vmware-cloud-foundation.md>)

### AI overview

This article explains why an AI-ready private cloud requires more than adding GPUs. It focuses on how Broadcom and Cisco are integrating VMware Cloud Foundation with Cisco Nexus One Fabric to address AI workload networking, including bandwidth-intensive east-west traffic, bursty north-south traffic, latency, congestion management, and telemetry across virtual and physical infrastructure.

### Source excerpt

An AI-ready private cloud is not simply a private cloud with GPUs added to it. What determines whether a private cloud platform can actually serve AI workloads effectively is everything built around them: how the fabric carries traffic, how the tenancy model lets teams consume capacity, and how policy and telemetry stay coherent across the ... Continued The post AI-Ready Private Cloud with Cisco and VMware appeared first on VMware Blogs.

## How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

DevFeed: [How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code](<https://devfeed.tech/articles/how-hugging-face-inference-endpoints-jobs-and-buckets-power-search-on-papers-with-code-7447.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/pwc-search>)

Author: Niels Rogge

Published: 2026-08-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI search](<https://devfeed.tech/topics/ai-search.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [jobs](<https://devfeed.tech/topics/jobs.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>)

Tags: [database](<https://devfeed.tech/tags/database.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [latency](<https://devfeed.tech/tags/latency.md>), [rag](<https://devfeed.tech/tags/rag.md>), [research](<https://devfeed.tech/tags/research.md>), [search](<https://devfeed.tech/tags/search.md>), [storage](<https://devfeed.tech/tags/storage.md>)

### AI overview

This article explains how Papers with Code uses hybrid search to find research papers through exact keyword matching and semantic vector search. The production system combines PostgreSQL full-text search, pgvector embeddings, reciprocal rank fusion, and Hugging Face Jobs, Storage Buckets, and Inference Endpoints.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Welcome Inkling by Thinking Machines

DevFeed: [Welcome Inkling by Thinking Machines](<https://devfeed.tech/articles/welcome-inkling-by-thinking-machines-7502.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/thinkingmachines-inkling>)

Author: ben burtenshaw; merve; Pedro Cuenca; Aritra Roy Gosthipaty; Andres Marafioti

Published: 2026-07-15T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Mixture of Experts (MoE)](<https://devfeed.tech/topics/mixture-of-experts-moe.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [sglang](<https://devfeed.tech/topics/sglang.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [audio](<https://devfeed.tech/tags/audio.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generation](<https://devfeed.tech/tags/generation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mixture-of-experts](<https://devfeed.tech/tags/mixture-of-experts.md>), [mixture-of-experts-moe](<https://devfeed.tech/tags/mixture-of-experts-moe.md>), [moe](<https://devfeed.tech/tags/moe.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

Thinking Machines Lab's Inkling is presented as a large open multimodal language model that accepts image, text, and audio inputs. The article covers its mixture-of-experts architecture, million-token context window, reasoning across modalities, fine-tuning use cases, model variants, and deployment through Hugging Face Inference Endpoints and inference frameworks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## DigitalOcean Evaluations: Production Model and Router Testing for the Inference Stack

DevFeed: [DigitalOcean Evaluations: Production Model and Router Testing for the Inference Stack](<https://devfeed.tech/articles/digitalocean-evaluations-production-model-and-router-testing-for-the-inference-stack-19920.md>)

Original publisher: [Read original article](<https://www.digitalocean.com/blog/now-available-evaluations>)

Author: Grace Morgan

Published: 2026-07-01T15:41:47Z

Content type: release

Language: en

Sources: [DigitalOcean](<https://devfeed.tech/sources/digitalocean.md>)

Topics: [Digital Ocean](<https://devfeed.tech/topics/digital-ocean.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Testing](<https://devfeed.tech/topics/testing.md>)

Tags: [configuration](<https://devfeed.tech/tags/configuration.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data](<https://devfeed.tech/tags/data.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [digitalocean](<https://devfeed.tech/tags/digitalocean.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [metric](<https://devfeed.tech/tags/metric.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [pii](<https://devfeed.tech/tags/pii.md>), [product-updates](<https://devfeed.tech/tags/product-updates.md>), [production](<https://devfeed.tech/tags/production.md>), [quality](<https://devfeed.tech/tags/quality.md>), [router](<https://devfeed.tech/tags/router.md>), [testing](<https://devfeed.tech/tags/testing.md>), [token](<https://devfeed.tech/tags/token.md>)

### AI overview

DigitalOcean Evaluations adds production testing for models and inference router configurations in the DigitalOcean Inference Engine. Teams can run LLM-as-a-Judge evaluations on their own prompts and data, compare quality, latency, and cost, and use built-in or custom rubrics across models, imports, and router setups.

### Source excerpt

Choosing the right model or inference router for production means more than reading a leaderboard. It means validating any model or routing configuration on your own data using your prompts and your evaluation criteria before it ever reaches production, and comparing quality, latency, and cost in one place. Evaluations, now available on the DigitalOcean Inference Engine, lets teams validate any model or inference router configuration on their own data before production. Run structured LLM-as-a-Judge evaluations across catalog models, fine-tuned models, BYOM imports, and router setups without stitching together a separate evaluation stack. DigitalOcean Evaluations Capabilities Evaluations provide everything teams need to validate model and router performance before production. LLM-as-a-Judge scoring runs across any candidate in your inference stack and returns per-item scores with judge rationale, plus latency, token, and cost tracking per run. Six pre-built metrics cover the most common evaluation needs out of the box. For teams that need full control: custom rubrics, reusable presets, MCP support, and full dataset management -- all in the same platform as the inference endpoints you use in production. View YouTube video Pre-Built and Custom Rubrics: Score Against Criteria That Match Your Domain The six pre-built metrics, correctness, completeness, faithfulness, PII, toxicity, and bias, cover common evaluation needs. For specialized domains, custom rubrics let teams define their own judge instructions and scoring criteria directly in the judge prompt. The judge evaluates responses against these criteria and returns per-item scores with rationale. Custom rubrics can also adapt the built-in correctness metric to different data formats instead of relying on a default interpretation. Evaluation Presets: Save Configurations and Re-Run Without Rebuilding Without saved configurations, every re-run becomes a rebuild with different judge models, parameters, or prompts, making

## Run a vLLM Server on HF Jobs in One Command

DevFeed: [Run a vLLM Server on HF Jobs in One Command](<https://devfeed.tech/articles/run-a-vllm-server-on-hf-jobs-in-one-command-7559.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/vllm-jobs>)

Author: Quentin Gallouédec

Published: 2026-06-26T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [jobs](<https://devfeed.tech/topics/jobs.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [API](<https://devfeed.tech/topics/api.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [cURL](<https://devfeed.tech/topics/curl.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [api](<https://devfeed.tech/tags/api.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [infrastructure](<https://devfeed.tech/tags/infrastructure.md>), [jobs](<https://devfeed.tech/tags/jobs.md>), [json](<https://devfeed.tech/tags/json.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [python](<https://devfeed.tech/tags/python.md>), [vllm](<https://devfeed.tech/tags/vllm.md>)

### AI overview

A practical guide to launching a vLLM model server on Hugging Face Jobs with a single command, querying it through the OpenAI-compatible API, authenticating requests with an HF token, managing costs, and scaling to larger multi-GPU models. It also contrasts ephemeral Jobs with managed Inference Endpoints.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## NVIDIA brings agents to life with DGX Spark and Reachy Mini

DevFeed: [NVIDIA brings agents to life with DGX Spark and Reachy Mini](<https://devfeed.tech/articles/nvidia-brings-agents-to-life-with-dgx-spark-and-reachy-mini-7371.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia-reachy-mini>)

Author: Jeff Boudier; Nader Khalil; Alec Fong

Published: 2026-01-05T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [DGX Spark](<https://devfeed.tech/topics/dgx-spark.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [reachy](<https://devfeed.tech/topics/reachy.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Orchestration](<https://devfeed.tech/topics/orchestration.md>), [Python](<https://devfeed.tech/topics/python.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Serverless](<https://devfeed.tech/topics/serverless.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [code](<https://devfeed.tech/tags/code.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [orchestration](<https://devfeed.tech/tags/orchestration.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [python](<https://devfeed.tech/tags/python.md>), [reachy](<https://devfeed.tech/tags/reachy.md>), [robotics](<https://devfeed.tech/tags/robotics.md>), [serverless](<https://devfeed.tech/tags/serverless.md>)

### AI overview

A step-by-step guide to building a desk-based AI companion with NVIDIA DGX Spark and Reachy Mini. It combines open reasoning and vision models, text-to-speech, Python, agent orchestration, and tool handlers, with options for local, cloud, and serverless deployment.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Building for an Open Future - our new partnership with Google Cloud

DevFeed: [Building for an Open Future - our new partnership with Google Cloud](<https://devfeed.tech/articles/building-for-an-open-future-our-new-partnership-with-google-cloud-7218.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/google-cloud>)

Author: Jeff Boudier; Simon Pagezy

Published: 2025-11-13T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Cloud Run](<https://devfeed.tech/topics/cloud-run.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Low-Latency Inference](<https://devfeed.tech/topics/low-latency-inference.md>), [Cache](<https://devfeed.tech/topics/cache.md>), [xet](<https://devfeed.tech/topics/xet.md>), [data](<https://devfeed.tech/topics/data.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cache](<https://devfeed.tech/tags/cache.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [google](<https://devfeed.tech/tags/google.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [networking](<https://devfeed.tech/tags/networking.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [storage](<https://devfeed.tech/tags/storage.md>), [vertex](<https://devfeed.tech/tags/vertex.md>), [vertex-ai](<https://devfeed.tech/tags/vertex-ai.md>), [xet](<https://devfeed.tech/tags/xet.md>)

### AI overview

Hugging Face and Google Cloud announce a strategic partnership focused on making open AI models easier to use, customize, deploy, and govern. The article describes integrations across Vertex AI, GKE AI/ML, Cloud Run GPUs, and other Google Cloud infrastructure, plus a planned CDN Gateway using Hugging Face Xet and Google Cloud storage and networking to accelerate model and dataset downloads and improve supply-chain robustness.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Blazingly fast whisper transcriptions with Inference Endpoints

DevFeed: [Blazingly fast whisper transcriptions with Inference Endpoints](<https://devfeed.tech/articles/blazingly-fast-whisper-transcriptions-with-inference-endpoints-7192.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/fast-whisper-endpoints>)

Author: Morgan Funtowicz; Freddy Boulton; Steven Zheng; Vaibhav Srivastav; Erik Kaunismäki; Michelle Habonneau

Published: 2025-05-13T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [audio](<https://devfeed.tech/tags/audio.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [guide](<https://devfeed.tech/tags/guide.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

Hugging Face introduces an optimized Whisper inference endpoint powered by vLLM. It targets newer NVIDIA GPUs and combines PyTorch compilation, CUDA graphs, and float8 KV-cache quantization to improve transcription inference speed and memory efficiency.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## The New and Fresh analytics in Inference Endpoints

DevFeed: [The New and Fresh analytics in Inference Endpoints](<https://devfeed.tech/articles/the-new-and-fresh-analytics-in-inference-endpoints-7181.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/endpoint-analytics>)

Author: Erik Kaunismäki; Thibault Goehringer; Remy; Corentin Regal; Michelle Habonneau

Published: 2025-03-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Monitoring](<https://devfeed.tech/topics/monitoring.md>), [debugging](<https://devfeed.tech/topics/debugging.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [real-time](<https://devfeed.tech/topics/real-time.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [debug](<https://devfeed.tech/tags/debug.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [guide](<https://devfeed.tech/tags/guide.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [latency](<https://devfeed.tech/tags/latency.md>), [lifecycle](<https://devfeed.tech/tags/lifecycle.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [monitoring](<https://devfeed.tech/tags/monitoring.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [state](<https://devfeed.tech/tags/state.md>)

### AI overview

Hugging Face has refreshed the analytics dashboard for Inference Endpoints with real-time metrics, faster data loading, customizable time ranges, auto-refresh, and a replica lifecycle view. The updates help users monitor performance, investigate issues, and understand replica state transitions.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Remote VAEs for decoding with Inference Endpoints 🤗

DevFeed: [Remote VAEs for decoding with Inference Endpoints 🤗](<https://devfeed.tech/articles/remote-vaes-for-decoding-with-inference-endpoints-7456.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/remote_vae>)

Author: hlky; Sayak Paul

Published: 2025-02-24T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [VAE](<https://devfeed.tech/topics/vae.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [diffusers](<https://devfeed.tech/topics/diffusers.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Inference Performance](<https://devfeed.tech/topics/inference-performance.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Deadlock](<https://devfeed.tech/topics/deadlock.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [concurrency](<https://devfeed.tech/tags/concurrency.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [data](<https://devfeed.tech/tags/data.md>), [diffusers](<https://devfeed.tech/tags/diffusers.md>), [diffusion](<https://devfeed.tech/tags/diffusion.md>), [generate](<https://devfeed.tech/tags/generate.md>), [generation](<https://devfeed.tech/tags/generation.md>), [getting-started](<https://devfeed.tech/tags/getting-started.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hdd](<https://devfeed.tech/tags/hdd.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [nvme](<https://devfeed.tech/tags/nvme.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [queue](<https://devfeed.tech/tags/queue.md>), [remote](<https://devfeed.tech/tags/remote.md>), [tensors](<https://devfeed.tech/tags/tensors.md>), [time](<https://devfeed.tech/tags/time.md>), [vae](<https://devfeed.tech/tags/vae.md>)

### AI overview

This article presents an experimental approach for decoding latent-space diffusion outputs with remote VAEs hosted on Inference Endpoints. It explains how remote decoding can reduce consumer GPU memory pressure, avoid the quality loss associated with tiled decoding, and improve concurrency by queueing generation requests. The article includes setup and usage examples for random tensors and pipelines involving SD v1.5, Flux, and HunyuanVideo.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## 1 Billion Classifications

DevFeed: [1 Billion Classifications](<https://devfeed.tech/articles/1-billion-classifications-7132.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/billion-classifications>)

Author: Derek Thomas

Published: 2025-02-13T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Inference](<https://devfeed.tech/topics/inference.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [k6](<https://devfeed.tech/topics/k6.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [Grafana](<https://devfeed.tech/topics/grafana.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>), [Testing](<https://devfeed.tech/topics/testing.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [JavaScript](<https://devfeed.tech/topics/javascript.md>), [.env](<https://devfeed.tech/topics/dotenv.md>)

Tags: [classification](<https://devfeed.tech/tags/classification.md>), [cost](<https://devfeed.tech/tags/cost.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [embedding](<https://devfeed.tech/tags/embedding.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [grafana](<https://devfeed.tech/tags/grafana.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [k6](<https://devfeed.tech/tags/k6.md>), [latency](<https://devfeed.tech/tags/latency.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [performance](<https://devfeed.tech/tags/performance.md>), [python](<https://devfeed.tech/tags/python.md>), [scale](<https://devfeed.tech/tags/scale.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

This article presents a methodology for calculating cost and latency when running large-scale classification and embedding workloads. It examines model architectures, hardware options, deployment, load testing, and inference servers, with a focus on processing 1 billion inputs and balancing batch-inference cost against latency.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## How to deploy and fine-tune DeepSeek models on AWS

DevFeed: [How to deploy and fine-tune DeepSeek models on AWS](<https://devfeed.tech/articles/how-to-deploy-and-fine-tune-deepseek-models-on-aws-7161.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/deepseek-r1-aws>)

Author: Simon Pagezy; Jeff Boudier; David Corvoysier

Published: 2025-01-30T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [deepseek](<https://devfeed.tech/topics/deepseek.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Amazon Bedrock](<https://devfeed.tech/topics/amazon-bedrock.md>), [Amazon SageMaker AI](<https://devfeed.tech/topics/amazon-sagemaker-ai.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>)

Tags: [amazon-bedrock](<https://devfeed.tech/tags/amazon-bedrock.md>), [amazon-sagemaker](<https://devfeed.tech/tags/amazon-sagemaker.md>), [amazon-sagemaker-ai](<https://devfeed.tech/tags/amazon-sagemaker-ai.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [aws](<https://devfeed.tech/tags/aws.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [llms](<https://devfeed.tech/tags/llms.md>), [models](<https://devfeed.tech/tags/models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>)

### AI overview

This tutorial explains how to deploy and fine-tune DeepSeek-R1 and its distilled models on AWS using Hugging Face Inference Endpoints. It covers production deployment with dedicated compute, autoscaling, scale-to-zero, security, optimized hardware, and deployment options through Amazon Bedrock and Amazon SageMaker AI.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Investing in Performance: Fine-tune small models with LLM insights - a CFM case study

DevFeed: [Investing in Performance: Fine-tune small models with LLM insights - a CFM case study](<https://devfeed.tech/articles/investing-in-performance-fine-tune-small-models-with-llm-insights-a-cfm-case-study-7139.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/cfm-case-study>)

Author: Oussama Ahouzi; champonnois; Jérémy L'Hour; Pirashanth Ratnamogan; Bérengère Patault; Morgane Goibert

Published: 2024-12-03T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [argilla](<https://devfeed.tech/topics/argilla.md>), [llama](<https://devfeed.tech/topics/llama.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Finance](<https://devfeed.tech/topics/finance.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Scalability](<https://devfeed.tech/topics/scalability.md>)

Tags: [argilla](<https://devfeed.tech/tags/argilla.md>), [case-studies](<https://devfeed.tech/tags/case-studies.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [expert-support](<https://devfeed.tech/tags/expert-support.md>), [expert-support-program](<https://devfeed.tech/tags/expert-support-program.md>), [financial-applications](<https://devfeed.tech/tags/financial-applications.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [scalability](<https://devfeed.tech/tags/scalability.md>)

### AI overview

A case study of Capital Fund Management's use of open-source LLMs to improve financial named entity recognition. It covers LLM-assisted labeling, fine-tuning smaller models with curated datasets, and deployment on Hugging Face Inference Endpoints to balance accuracy, cost, and scalability.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Deploying Speech-to-Speech on Hugging Face

DevFeed: [Deploying Speech-to-Speech on Hugging Face](<https://devfeed.tech/articles/deploying-speech-to-speech-on-hugging-face-7462.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/s2s_endpoint>)

Author: Andres Marafioti; Derek Thomas; Diego Maniloff; Eustache Le Bihan

Published: 2024-10-22T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [audio](<https://devfeed.tech/tags/audio.md>), [docker](<https://devfeed.tech/tags/docker.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>)

### AI overview

A step-by-step guide to packaging and deploying Hugging Face's Speech-to-Speech pipeline on Inference Endpoints using a custom Docker image.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Llama can now see and run on your device - welcome Llama 3.2

DevFeed: [Llama can now see and run on your device - welcome Llama 3.2](<https://devfeed.tech/articles/llama-can-now-see-and-run-on-your-device-welcome-llama-3-2-7337.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/llama32>)

Author: merve; Philipp Schmid; Omar Sanseviero; Vaibhav Srivastav; Lewis Tunstall; Aritra Roy Gosthipaty; Pedro Cuenca

Published: 2024-09-25T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [llama](<https://devfeed.tech/topics/llama.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [meta](<https://devfeed.tech/tags/meta.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [tgi](<https://devfeed.tech/tags/tgi.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

Meta's Llama 3.2 release introduces multimodal Vision models in 11B and 90B sizes, smaller text-only 1B and 3B models for on-device use, and vision-enabled Llama Guard 3. The article describes their capabilities, architecture, supported languages, inference examples, and integrations with Hugging Face Transformers, TGI, Inference Endpoints, Google Cloud, Amazon SageMaker, and DELL Enterprise Hub.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Google Cloud TPUs made available to Hugging Face users

DevFeed: [Google Cloud TPUs made available to Hugging Face users](<https://devfeed.tech/articles/google-cloud-tpus-made-available-to-hugging-face-users-7524.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tpu-inference-endpoints-spaces>)

Author: Simon Pagezy; Michelle Habonneau; Philipp Schmid; Alvaro Moran

Published: 2024-07-09T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Sovereign AI](<https://devfeed.tech/topics/sovereign-ai.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [spaces](<https://devfeed.tech/topics/spaces.md>), [optimum](<https://devfeed.tech/topics/optimum.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [gemma](<https://devfeed.tech/topics/gemma.md>), [llama](<https://devfeed.tech/topics/llama.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [Hardware](<https://devfeed.tech/topics/hardware.md>)

Tags: [gcp](<https://devfeed.tech/tags/gcp.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llama](<https://devfeed.tech/tags/llama.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [performance](<https://devfeed.tech/tags/performance.md>), [spaces](<https://devfeed.tech/tags/spaces.md>), [tgi](<https://devfeed.tech/tags/tgi.md>), [tpu](<https://devfeed.tech/tags/tpu.md>)

### AI overview

Hugging Face announces that Google Cloud TPUs are available for Inference Endpoints and Spaces. Google TPU v5e configurations can deploy supported models through managed infrastructure, while Optimum TPU and Text Generation Inference help train and serve models on TPUs.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Deploy models on AWS Inferentia2 from Hugging Face

DevFeed: [Deploy models on AWS Inferentia2 from Hugging Face](<https://devfeed.tech/articles/deploy-models-on-aws-inferentia2-from-hugging-face-7285.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/inferentia-inference-endpoints>)

Author: Jeff Boudier; Philipp Schmid

Published: 2024-05-22T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AWS AI chips](<https://devfeed.tech/topics/aws-ai-chips.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Amazon SageMaker](<https://devfeed.tech/topics/amazon-sagemaker.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [autoscaling](<https://devfeed.tech/topics/autoscaling.md>), [Amazon Web Services](<https://devfeed.tech/topics/aws.md>), [MLOps](<https://devfeed.tech/topics/mlops.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [autoscaling](<https://devfeed.tech/tags/autoscaling.md>), [aws](<https://devfeed.tech/tags/aws.md>), [aws-trainium](<https://devfeed.tech/tags/aws-trainium.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [llama3](<https://devfeed.tech/tags/llama3.md>), [mlops](<https://devfeed.tech/tags/mlops.md>), [model](<https://devfeed.tech/tags/model.md>), [optimum](<https://devfeed.tech/tags/optimum.md>), [partnership](<https://devfeed.tech/tags/partnership.md>), [tgi](<https://devfeed.tech/tags/tgi.md>)

### AI overview

Hugging Face announces support for deploying models on AWS Inferentia2 through Amazon SageMaker and Hugging Face Inference Endpoints. The update enables scalable inference for supported models, including Meta Llama 3, with multiple instance sizes, managed features, and autoscaling.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints

DevFeed: [Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints](<https://devfeed.tech/articles/powerful-asr-diarization-speculative-decoding-with-hugging-face-inference-endpoints-7106.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/asr-diarization>)

Author: Sergei Petrov; Vaibhav Srivastav; Pedro Cuenca; Philipp Schmid

Published: 2024-05-01T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [flash-attention-2](<https://devfeed.tech/tags/flash-attention-2.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

This article explains how to build a custom inference handler for Automatic Speech Recognition, speaker diarization, and speculative decoding on Hugging Face Inference Endpoints. It covers modular pipeline design, repository files, Pyannote-based diarization, PyTorch SDPA with Flash Attention 2, and constraints on speculative decoding such as batch size one and compatible decoder architectures.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Welcome Llama 3 - Meta's new open LLM

DevFeed: [Welcome Llama 3 - Meta's new open LLM](<https://devfeed.tech/articles/welcome-llama-3-meta-s-new-open-llm-7334.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/llama3>)

Author: Philipp Schmid; Omar Sanseviero; Pedro Cuenca; Younes B; Leandro von Werra

Published: 2024-04-18T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [llama](<https://devfeed.tech/topics/llama.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>), [trl](<https://devfeed.tech/topics/trl.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Amazon SageMaker](<https://devfeed.tech/topics/amazon-sagemaker.md>), [Google Cloud Platform (GCP)](<https://devfeed.tech/topics/google-cloud.md>)

Tags: [amazon-sagemaker](<https://devfeed.tech/tags/amazon-sagemaker.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [community](<https://devfeed.tech/tags/community.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llm](<https://devfeed.tech/tags/llm.md>), [meta](<https://devfeed.tech/tags/meta.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [trl](<https://devfeed.tech/tags/trl.md>)

### AI overview

Meta has released Llama 3, an open-access LLM family available through Hugging Face. The release includes 8B and 70B base and instruction-tuned models, plus Llama Guard 2 for classifying potentially unsafe LLM inputs and responses. It also adds integrations with Transformers, Hugging Chat, Inference Endpoints, Google Cloud, Amazon SageMaker, and TRL.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## CodeGemma - an official Google release for code LLMs

DevFeed: [CodeGemma - an official Google release for code LLMs](<https://devfeed.tech/articles/codegemma-an-official-google-release-for-code-llms-7145.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/codegemma>)

Author: Pedro Cuenca; Omar Sanseviero; Vaibhav Srivastav; Philipp Schmid; Mishig ᠮᠢᠰᠾᠢᠭ; Loubna Ben Allal

Published: 2024-04-09T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [gemma](<https://devfeed.tech/topics/gemma.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [code-completion](<https://devfeed.tech/topics/code-completion.md>), [Google](<https://devfeed.tech/topics/google.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Java](<https://devfeed.tech/topics/java.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [c-plus-plus](<https://devfeed.tech/tags/c-plus-plus.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [code-completion](<https://devfeed.tech/tags/code-completion.md>), [community](<https://devfeed.tech/tags/community.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gcp](<https://devfeed.tech/tags/gcp.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google](<https://devfeed.tech/tags/google.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [java](<https://devfeed.tech/tags/java.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

Google's CodeGemma release introduces open-access 2B and 7B code-specialist LLMs, including base and instruction-tuned variants. The models support code infilling, completion, generation, code understanding, conversational use, and mathematical reasoning, with integrations across the Hugging Face ecosystem and Google Cloud.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.