# VLMs

Published articles for VLMs.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Building Federated Multimodal AI Workflows with NVIDIA FLARE

DevFeed: [Building Federated Multimodal AI Workflows with NVIDIA FLARE](<https://devfeed.tech/articles/building-federated-multimodal-ai-workflows-with-nvidia-flare-6776.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/building-federated-multimodal-ai-workflows-with-nvidia-flare/>)

Author: Tanya Lenz

Published: 2026-08-19T17:50:47Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [edge-computing](<https://devfeed.tech/tags/edge-computing.md>), [featured](<https://devfeed.tech/tags/featured.md>), [federated-learning](<https://devfeed.tech/tags/federated-learning.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [multimodal-ai](<https://devfeed.tech/tags/multimodal-ai.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [python](<https://devfeed.tech/tags/python.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [training](<https://devfeed.tech/tags/training.md>), [training-ai-models](<https://devfeed.tech/tags/training-ai-models.md>), [updates](<https://devfeed.tech/tags/updates.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

The article explains how NVIDIA FLARE supports federated training for multimodal and vision-language models when data remains distributed across sites. It focuses on deciding which model state to exchange and on efficiently transferring and aggregating large updates through externalization, tensor streaming, and disk-backed aggregation.

### Source excerpt

Modern vision-language models (VLMs) can support tasks such as visual question answering, captioning, and image-text reasoning. In practice, however, the data...

## MindTopo reveals VLMs' spatial reasoning abilities

DevFeed: [MindTopo reveals VLMs' spatial reasoning abilities](<https://devfeed.tech/articles/mindtopo-reveals-vlms-spatial-reasoning-abilities-6802.md>)

Original publisher: [Read original article](<https://www.microsoft.com/en-us/research/blog/mindtopo-reveals-vlms-spatial-reasoning-abilities/>)

Author: Yunfei Ge, Anbang Liu, Qineng Wang, Johnalbert Garnica, Zihan Wang, Reuben Tan, Jianfeng Gao, Ruohan Zhang, Yining Hong, Jiajun Wu, Manling Li

Published: 2026-08-12T16:00:00Z

Content type: article

Language: en

Sources: [Microsoft Research](<https://devfeed.tech/sources/microsoft-research.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [multimodal-ai](<https://devfeed.tech/topics/multimodal-ai.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Robotics](<https://devfeed.tech/topics/robotics.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [research-blog](<https://devfeed.tech/tags/research-blog.md>), [testing](<https://devfeed.tech/tags/testing.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

MindTopo is a benchmark for evaluating whether multimodal large language models can understand and manipulate topological relationships such as connectivity, enclosure, order, separation, and knots. It compares static recognition with interactive planning and finds that current models often lose track of structural relationships during sequences of actions.

### Source excerpt

A path, a fence, a knot. MindTopo sets a new benchmark for testing how AI understands topological relationships and highlights new opportunities to strengthen spatial reasoning and planning. The post MindTopo reveals VLMs' spatial reasoning abilities appeared first on Microsoft Research.

## Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

DevFeed: [Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement](<https://devfeed.tech/articles/introducing-care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement-6800.md>)

Original publisher: [Read original article](<https://www.microsoft.com/en-us/research/blog/introducing-care-x-towards-clinically-useful-radiology-vlms-with-auxiliary-supervision-reward-aligned-learning-and-tool-augmented-measurement/>)

Author: Mercy Ranjit, Nikhilesh E, Dr. Abhyuday Kumara Swamy, Tanuja Ganu

Published: 2026-08-11T16:00:00Z

Content type: article

Language: en

Sources: [Microsoft Research](<https://devfeed.tech/sources/microsoft-research.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Reinforcement learning](<https://devfeed.tech/topics/reinforcement-learning.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [data](<https://devfeed.tech/tags/data.md>), [generation](<https://devfeed.tech/tags/generation.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [research-blog](<https://devfeed.tech/tags/research-blog.md>), [tools](<https://devfeed.tech/tags/tools.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

CARE-X is a research chest X-ray vision-language model that combines free-text report generation, structured diagnostic prediction, and reinforcement learning for multi-task clinical interpretation. The article also describes a separate experiment using deterministic measurement tools with Qwen3-VL-4B-Instruct and reports validation on real-world Indian clinical data, while emphasizing that CARE-X is not approved for clinical use.

### Source excerpt

Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.

## Meta is back with Muse Glimmer: local, agentic, multimodal, and open source

DevFeed: [Meta is back with Muse Glimmer: local, agentic, multimodal, and open source](<https://devfeed.tech/articles/meta-is-back-with-muse-glimmer-local-agentic-multimodal-and-open-source-7362.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/muse-glimmer>)

Author: Pedro Cuenca; merve; ben burtenshaw; Aritra Roy Gosthipaty

Published: 2026-08-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [coding](<https://devfeed.tech/topics/coding.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [images](<https://devfeed.tech/tags/images.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llms](<https://devfeed.tech/tags/llms.md>), [local](<https://devfeed.tech/tags/local.md>), [meta](<https://devfeed.tech/tags/meta.md>), [model](<https://devfeed.tech/tags/model.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [muse](<https://devfeed.tech/tags/muse.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [videos](<https://devfeed.tech/tags/videos.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

Hugging Face presents Muse Glimmer, a local, agentic, multimodal, open-source 30B-parameter vision-language model developed with Meta. The article outlines its vision and language architecture, benchmark context, optional speculative decoding for faster generation, and support for both images and videos.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning

DevFeed: [NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning](<https://devfeed.tech/articles/nvidia-ising-enables-fully-automated-quantum-computer-calibration-with-enhanced-in-context-learning-6895.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/nvidia-ising-enables-fully-automated-quantum-computer-calibration-with-enhanced-in-context-learning/>)

Author: Tanya Lenz

Published: 2026-07-27T16:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [Ising](<https://devfeed.tech/topics/ising.md>), [Quantum Computing](<https://devfeed.tech/topics/quantum-computing.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [NVFP4](<https://devfeed.tech/topics/nvfp4.md>), [DGX Spark](<https://devfeed.tech/topics/dgx-spark.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [data-science](<https://devfeed.tech/tags/data-science.md>), [dgx-spark](<https://devfeed.tech/tags/dgx-spark.md>), [featured](<https://devfeed.tech/tags/featured.md>), [ising](<https://devfeed.tech/tags/ising.md>), [nvfp4](<https://devfeed.tech/tags/nvfp4.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [quantum](<https://devfeed.tech/tags/quantum.md>), [quantum-computing](<https://devfeed.tech/tags/quantum-computing.md>), [simulation-modeling-design](<https://devfeed.tech/tags/simulation-modeling-design.md>), [training-ai-models](<https://devfeed.tech/tags/training-ai-models.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

NVIDIA Ising Calibration 1.5 is an open-source vision-language model for interpreting quantum-processor diagnostics and recommending calibration actions. The article highlights zero-shot and in-context learning evaluation on QCalEval, plus an NVFP4-quantized version for local deployment.

### Source excerpt

NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they...

## How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo

DevFeed: [How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo](<https://devfeed.tech/articles/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo-6853.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/>)

Author: Tanya Lenz

Published: 2026-07-14T16:00:00Z

Content type: tutorial

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>), [debugging](<https://devfeed.tech/topics/debugging.md>)

Tags: [agent-skill](<https://devfeed.tech/tags/agent-skill.md>), [agent-skills](<https://devfeed.tech/tags/agent-skills.md>), [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [debugging](<https://devfeed.tech/tags/debugging.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [how-to](<https://devfeed.tech/tags/how-to.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [python](<https://devfeed.tech/tags/python.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [research](<https://devfeed.tech/tags/research.md>), [rl](<https://devfeed.tech/tags/rl.md>), [synthetic-data-generation](<https://devfeed.tech/tags/synthetic-data-generation.md>), [training](<https://devfeed.tech/tags/training.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

A tutorial on running a skill-based autoresearch workflow in which coding AI agents set up, debug, run, monitor, and iterate on reinforcement-learning experiments using NVIDIA NeMo RL and NeMo Gym.

### Source excerpt

Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes, resolve...

## AI Model Co-Design: Hardware-Friendly LLM Design

DevFeed: [AI Model Co-Design: Hardware-Friendly LLM Design](<https://devfeed.tech/articles/ai-model-co-design-hardware-friendly-llm-design-6762.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/ai-model-co-design-hardware-friendly-llm-design/>)

Author: Elizabeth Goodman

Published: 2026-07-10T16:36:02Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-inference](<https://devfeed.tech/tags/ai-inference.md>), [data-center-cloud](<https://devfeed.tech/tags/data-center-cloud.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [design](<https://devfeed.tech/tags/design.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [developers](<https://devfeed.tech/tags/developers.md>), [featured](<https://devfeed.tech/tags/featured.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference-performance](<https://devfeed.tech/tags/inference-performance.md>), [latency](<https://devfeed.tech/tags/latency.md>), [llm](<https://devfeed.tech/tags/llm.md>), [llms](<https://devfeed.tech/tags/llms.md>), [model](<https://devfeed.tech/tags/model.md>), [performance](<https://devfeed.tech/tags/performance.md>), [training](<https://devfeed.tech/tags/training.md>), [training-ai-models](<https://devfeed.tech/tags/training-ai-models.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

A practical primer on designing LLMs for modern hardware while balancing accuracy, throughput, and responsiveness. It explains how context length and latency or throughput goals change the importance of attention, feed-forward layers, and parallelism.

### Source excerpt

AI performance comes down to three dimensions: Accuracy: How well the model reasons and produces outputs Throughput: How many tokens per second a...

## ScreenSuite - The most comprehensive evaluation suite for GUI Agents!

DevFeed: [ScreenSuite - The most comprehensive evaluation suite for GUI Agents!](<https://devfeed.tech/articles/screensuite-the-most-comprehensive-evaluation-suite-for-gui-agents-7468.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/screensuite>)

Author: Amir Mahla; Aymeric Roucher; Thomas Wolf

Published: 2025-06-06T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [screensuite](<https://devfeed.tech/topics/screensuite.md>), [gui-agents](<https://devfeed.tech/topics/gui-agents.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>), [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [docker](<https://devfeed.tech/tags/docker.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gui](<https://devfeed.tech/tags/gui.md>), [gui-agents](<https://devfeed.tech/tags/gui-agents.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [screensuite](<https://devfeed.tech/tags/screensuite.md>), [smolagents](<https://devfeed.tech/tags/smolagents.md>), [virtual-machines](<https://devfeed.tech/tags/virtual-machines.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

ScreenSuite is an open benchmarking and evaluation suite for GUI agents. It unifies 13 benchmarks covering perception, grounding, single-step actions, and multi-step agentic capabilities for Vision Language Models, with support for remote desktop sandboxes and Ubuntu or Android virtual machines launched in Docker.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H

DevFeed: [Holo1: New family of GUI automation VLMs powering GUI agent Surfer-H](<https://devfeed.tech/articles/holo1-new-family-of-gui-automation-vlms-powering-gui-agent-surfer-h-7002.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Hcompany/holo1>)

Author: Mats L Richter; Pierre-Louis Cedoz

Published: 2025-06-03T13:27:59Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [browsers](<https://devfeed.tech/topics/browsers.md>), [Website](<https://devfeed.tech/topics/website.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Transformers](<https://devfeed.tech/topics/transformers.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Localization (l10n)](<https://devfeed.tech/topics/localization.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [automation](<https://devfeed.tech/tags/automation.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [browsers](<https://devfeed.tech/tags/browsers.md>), [cost](<https://devfeed.tech/tags/cost.md>), [gui](<https://devfeed.tech/tags/gui.md>), [models](<https://devfeed.tech/tags/models.md>), [modular](<https://devfeed.tech/tags/modular.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [web](<https://devfeed.tech/tags/web.md>)

### AI overview

Holo1 is an open-source family of Action VLMs for understanding web interfaces and locating UI elements precisely. Holo1-3B and Holo1-7B are released on Hugging Face with the WebClick benchmark; Holo1-7B reports 76.2% average accuracy on common UI localization benchmarks. The article also describes Surfer-H, a browser-based web automation agent built with separate policy, localization, and validation components.

### Source excerpt

Surfer-H, a web-native agent that interacts with browsers like a human relies on the Holo1. Holo1 is the first family of open-source Action VLMs designed specifically for deep web UI understanding and precise localization. The family includes Holo1-3B and Holo1-7B models, with the latter achieving 76.2% average accuracy on common UI localization benchmarks--the highest among small-size models.

## The Transformers Library: standardizing model definitions

DevFeed: [The Transformers Library: standardizing model definitions](<https://devfeed.tech/articles/the-transformers-library-standardizing-model-definitions-7535.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/transformers-model-definition>)

Author: Lysandre; Arthur Zucker; Pedro Cuenca; Julien Chaumond

Published: 2025-05-15T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Transformer](<https://devfeed.tech/topics/transformer.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [interoperability](<https://devfeed.tech/topics/interoperability.md>), [sglang](<https://devfeed.tech/topics/sglang.md>), [vllm](<https://devfeed.tech/topics/vllm.md>), [llama.cpp](<https://devfeed.tech/topics/llama-cpp.md>), [tgi](<https://devfeed.tech/topics/tgi.md>), [MLX](<https://devfeed.tech/topics/mlx.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [community](<https://devfeed.tech/tags/community.md>), [interoperability](<https://devfeed.tech/tags/interoperability.md>), [library](<https://devfeed.tech/tags/library.md>), [llama-cpp](<https://devfeed.tech/tags/llama-cpp.md>), [llms](<https://devfeed.tech/tags/llms.md>), [mlx](<https://devfeed.tech/tags/mlx.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [sglang](<https://devfeed.tech/tags/sglang.md>), [tgi](<https://devfeed.tech/tags/tgi.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [vllm](<https://devfeed.tech/tags/vllm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

The article presents Transformers as a standard library for defining and supporting machine learning model architectures. It describes its broad ecosystem integrations, including training frameworks and inference engines, and highlights interoperability with vLLM, SGLang, TGI, llama.cpp, and MLX.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Vision Language Models (Better, faster, stronger)

DevFeed: [Vision Language Models (Better, faster, stronger)](<https://devfeed.tech/articles/vision-language-models-better-faster-stronger-7561.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/vlms-2025>)

Author: merve; Sergio Paniego; Aritra Roy Gosthipaty; Pedro Cuenca; Andres Marafioti

Published: 2025-05-12T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [qwen](<https://devfeed.tech/topics/qwen.md>), [Meta](<https://devfeed.tech/topics/meta.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>)

Tags: [architecture](<https://devfeed.tech/tags/architecture.md>), [community](<https://devfeed.tech/tags/community.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [research](<https://devfeed.tech/tags/research.md>), [streaming](<https://devfeed.tech/tags/streaming.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

This article reviews major developments in vision-language models over the past year, including smaller and more capable models, new architectures, reasoning, agency, long-video understanding, multimodal Retrieval Augmented Generation, and multimodal agents. It also explains any-to-any models and discusses examples including Chameleon, Lumina-mGPT, Qwen 2.5 Omni, MiniCPM-o 2.6, and Janus-Pro-7B.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Introducing AutoRound: Intel's Advanced Quantization for LLMs and VLMs

DevFeed: [Introducing AutoRound: Intel's Advanced Quantization for LLMs and VLMs](<https://devfeed.tech/articles/introducing-autoround-intel-s-advanced-quantization-for-llms-and-vlms-7113.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/autoround>)

Author: wenhua cheng; Haihao Shen; weiweiz1; Heng Guo; Huang, Tai; Ke Ding; Ilyas Moutawwakil; Marc Sun; Mohamed Mekkouri

Published: 2025-04-29T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [quantization](<https://devfeed.tech/topics/quantization.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [intel](<https://devfeed.tech/topics/intel.md>), [GPU](<https://devfeed.tech/topics/gpu.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [deepseek](<https://devfeed.tech/topics/deepseek.md>), [llama](<https://devfeed.tech/topics/llama.md>), [qwen](<https://devfeed.tech/topics/qwen.md>)

Tags: [architectures](<https://devfeed.tech/tags/architectures.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [cuda](<https://devfeed.tech/tags/cuda.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [intel](<https://devfeed.tech/tags/intel.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llms](<https://devfeed.tech/tags/llms.md>), [model](<https://devfeed.tech/tags/model.md>), [offline](<https://devfeed.tech/tags/offline.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [post-training](<https://devfeed.tech/tags/post-training.md>), [precision](<https://devfeed.tech/tags/precision.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

This article introduces AutoRound, Intel's weight-only post-training quantization method for LLMs and VLMs. It uses signed gradient descent to optimize weight rounding and clipping ranges, supports low-bit formats from INT2 to INT8, and aims to preserve accuracy with efficient quantization. The article describes its model, hardware, export-format, calibration, and tuning-recipe support, including performance on low-bit benchmarks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## SigLIP 2: A better multilingual vision language encoder

DevFeed: [SigLIP 2: A better multilingual vision language encoder](<https://devfeed.tech/articles/siglip-2-a-better-multilingual-vision-language-encoder-7474.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/siglip2>)

Author: Aritra Roy Gosthipaty; merve; Pavel Iakubovskii

Published: 2025-02-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [object-detection](<https://devfeed.tech/topics/object-detection.md>), [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [object-detection](<https://devfeed.tech/tags/object-detection.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [train](<https://devfeed.tech/tags/train.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

Google's SigLIP 2 is a multilingual vision-language encoder family that extends SigLIP's sigmoid-loss training with additional objectives for semantic understanding, localization, and dense visual features. The models improve on SigLIP across scales and core capabilities including zero-shot classification, image-text retrieval, and visual representation transfer, with a dynamic-resolution variant for resolution- and aspect-ratio-sensitive tasks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## PaliGemma 2 Mix - New Instruction Vision Language Models by Google

DevFeed: [PaliGemma 2 Mix - New Instruction Vision Language Models by Google](<https://devfeed.tech/articles/paligemma-2-mix-new-instruction-vision-language-models-by-google-7437.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/paligemma2mix>)

Author: merve; Aritra Roy Gosthipaty; Andreas P. Steiner

Published: 2025-02-19T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Google](<https://devfeed.tech/topics/google.md>), [gemma](<https://devfeed.tech/topics/gemma.md>), [object-detection](<https://devfeed.tech/topics/object-detection.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [demo](<https://devfeed.tech/tags/demo.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google](<https://devfeed.tech/tags/google.md>), [llm](<https://devfeed.tech/tags/llm.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [object-detection](<https://devfeed.tech/tags/object-detection.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

Google's PaliGemma 2 mix models are fine-tuned on a mixture of vision-language tasks, including OCR, image captioning, visual question answering, document understanding, object detection, and image segmentation. The article explains how the mix models indicate the performance of pretrained PaliGemma 2 checkpoints after fine-tuning and describes their prompting approach.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## We now support VLMs in smolagents!

DevFeed: [We now support VLMs in smolagents!](<https://devfeed.tech/articles/we-now-support-vlms-in-smolagents-7478.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/smolagents-can-see>)

Author: Aymeric Roucher; merve; Albert Villanova del Moral

Published: 2025-01-24T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [smolagents](<https://devfeed.tech/topics/smolagents.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [AI Development](<https://devfeed.tech/topics/ai-development.md>), [web browser](<https://devfeed.tech/topics/web-browser.md>), [document ai](<https://devfeed.tech/topics/document-ai.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [agents](<https://devfeed.tech/tags/agents.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [document-ai](<https://devfeed.tech/tags/document-ai.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [smolagents](<https://devfeed.tech/tags/smolagents.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [web-browser](<https://devfeed.tech/tags/web-browser.md>)

### AI overview

smolagents now supports vision-language models natively in agentic pipelines. The article explains how agents can receive images at initialization or dynamically through memory callbacks, enabling visual web browsing, Document AI workflows, and processing long PDFs with visual elements.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Visual Document Retrieval Goes Multilingual

DevFeed: [Visual Document Retrieval Goes Multilingual](<https://devfeed.tech/articles/visual-document-retrieval-goes-multilingual-7550.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/vdr-2b-multilingual>)

Author: Marco Cimolai; Logan Markewich

Published: 2025-01-10T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [datasets](<https://devfeed.tech/topics/datasets.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [llamaindex](<https://devfeed.tech/topics/llamaindex.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [data](<https://devfeed.tech/topics/data.md>), [Query (disambiguation)](<https://devfeed.tech/topics/query.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [cv](<https://devfeed.tech/tags/cv.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [integrations](<https://devfeed.tech/tags/integrations.md>), [llamaindex](<https://devfeed.tech/tags/llamaindex.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [open-source-collab](<https://devfeed.tech/tags/open-source-collab.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [synthetic](<https://devfeed.tech/tags/synthetic.md>), [vector](<https://devfeed.tech/tags/vector.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

The article introduces a multilingual embedding model for visual document retrieval and its English-only counterpart. The models encode document page screenshots into dense single-vector representations, enabling visual search across languages without OCR or document-chunking pipelines. The article also presents a 500,000-sample open-source multilingual synthetic dataset, reports faster inference and lower VRAM usage, and describes cross-lingual retrieval and Matryoshka Representation Learning capabilities.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Introducing TextImage Augmentation for Document Images

DevFeed: [Introducing TextImage Augmentation for Document Images](<https://devfeed.tech/articles/introducing-textimage-augmentation-for-document-images-7172.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/doc_aug_hf_alb>)

Author: Dana Aubakirova; Pablo Montalvo; Vladimir Iglovikov

Published: 2024-08-06T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [data augmentation](<https://devfeed.tech/topics/data-augmentation.md>), [albumentations](<https://devfeed.tech/topics/albumentations.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [albumentations](<https://devfeed.tech/tags/albumentations.md>), [data-augmentation](<https://devfeed.tech/tags/data-augmentation.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [document-ai](<https://devfeed.tech/tags/document-ai.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [research](<https://devfeed.tech/tags/research.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [text-generation](<https://devfeed.tech/tags/text-generation.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

The article introduces a multimodal data augmentation pipeline for document images used in Vision Language Model fine-tuning. Developed with Albumentations AI, it modifies document images and their text annotations together while aiming to preserve text quality. The methods include text insertion, deletion, swapping, and stopword replacement, followed by image masking and inpainting.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?

DevFeed: [LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?](<https://devfeed.tech/articles/lave-zero-shot-vqa-evaluation-on-docmatix-with-llms-do-we-still-need-fine-tuning-7574.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/zero-shot-vqa-docmatix>)

Author: Dana Aubakirova; Andres Marafioti

Published: 2024-07-25T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [LLMs](<https://devfeed.tech/topics/llms.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [llms](<https://devfeed.tech/tags/llms.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [research](<https://devfeed.tech/tags/research.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [vqa](<https://devfeed.tech/tags/vqa.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

This article examines zero-shot visual question answering evaluation on the synthetic Docmatix dataset using large language models and vision-language models. It explains why traditional VQA Accuracy can undervalue semantically correct answers in out-of-distribution settings and discusses the trade-off between fine-tuning models and developing metrics that better reflect human judgment.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Preference Optimization for Vision Language Models

DevFeed: [Preference Optimization for Vision Language Models](<https://devfeed.tech/articles/preference-optimization-for-vision-language-models-7175.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/dpo_vlm>)

Author: Quentin Gallouédec; Shengyi Costa Huang; merve; Kashif Rasul

Published: 2024-07-10T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [data](<https://devfeed.tech/tags/data.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [images](<https://devfeed.tech/tags/images.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [rlhf](<https://devfeed.tech/tags/rlhf.md>), [training](<https://devfeed.tech/tags/training.md>), [trl](<https://devfeed.tech/tags/trl.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

A tutorial on training vision-language models with TRL's direct preference optimization support, covering preference data formatting and memory considerations.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Going multimodal: How Prezi is leveraging the Hub and the Expert Support Program to accelerate their ML roadmap

DevFeed: [Going multimodal: How Prezi is leveraging the Hub and the Expert Support Program to accelerate their ML roadmap](<https://devfeed.tech/articles/going-multimodal-how-prezi-is-leveraging-the-hub-and-the-expert-support-program-to-accelerate-their-ml-roadmap-7444.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/prezi-case-study>)

Author: Violette; Jeff Boudier; Moritz Borrett-Laurer (formerly Laurer); Máté Börcsök

Published: 2024-06-19T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [AI search](<https://devfeed.tech/topics/ai-search.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [case-studies](<https://devfeed.tech/tags/case-studies.md>), [expert-support](<https://devfeed.tech/tags/expert-support.md>), [images](<https://devfeed.tech/tags/images.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [llm](<https://devfeed.tech/tags/llm.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [search](<https://devfeed.tech/tags/search.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlms](<https://devfeed.tech/tags/vlms.md>)

### AI overview

Prezi describes using Hugging Face expert support to improve its multimodal machine-learning workflow for generating presentation drafts. The workflow combines vision, text, and vision-language models, and uses an open-source re-ranker to select presentation assets.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.