# vqa

Published articles for vqa.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Visual Salamandra: Pushing the Boundaries of Multimodal Understanding

DevFeed: [Visual Salamandra: Pushing the Boundaries of Multimodal Understanding](<https://devfeed.tech/articles/visual-salamandra-pushing-the-boundaries-of-multimodal-understanding-6990.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/BSC-LT/visualsalamandra7b>)

Author: Iñigo Pikabea; Jaume Lozano

Published: 2025-04-11T14:21:56Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [multimodal-ai](<https://devfeed.tech/topics/multimodal-ai.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [mlp](<https://devfeed.tech/topics/mlp.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [blog](<https://devfeed.tech/tags/blog.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [grounding](<https://devfeed.tech/tags/grounding.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [language](<https://devfeed.tech/tags/language.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mlp](<https://devfeed.tech/tags/mlp.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [multimodal-ai](<https://devfeed.tech/tags/multimodal-ai.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [research](<https://devfeed.tech/tags/research.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vqa](<https://devfeed.tech/tags/vqa.md>)

### AI overview

Visual Salamandra is a multilingual multimodal model built by extending the Salamandra Instructed 7B model with Google's SigLIP image encoder, an MLP projector, and late-fusion techniques. It processes text, images, and videos, with training focused on visual grounding, document understanding, mathematical reasoning, OCR, and European-language coverage.

### Source excerpt

A Blog post by Language Technologies Laboratory @ Barcelona Supercomputing Center on Hugging Face

## LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?

DevFeed: [LAVE: Zero-shot VQA Evaluation on Docmatix with LLMs - Do We Still Need Fine-Tuning?](<https://devfeed.tech/articles/lave-zero-shot-vqa-evaluation-on-docmatix-with-llms-do-we-still-need-fine-tuning-7574.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/zero-shot-vqa-docmatix>)

Author: Dana Aubakirova; Andres Marafioti

Published: 2024-07-25T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [LLMs](<https://devfeed.tech/topics/llms.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [llms](<https://devfeed.tech/tags/llms.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [research](<https://devfeed.tech/tags/research.md>), [synthetic-data](<https://devfeed.tech/tags/synthetic-data.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [vlms](<https://devfeed.tech/tags/vlms.md>), [vqa](<https://devfeed.tech/tags/vqa.md>), [zero-shot](<https://devfeed.tech/tags/zero-shot.md>)

### AI overview

This article examines zero-shot visual question answering evaluation on the synthetic Docmatix dataset using large language models and vision-language models. It explains why traditional VQA Accuracy can undervalue semantically correct answers in out-of-distribution settings and discusses the trade-off between fine-tuning models and developing metrics that better reflect human judgment.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models

DevFeed: [Fine-tuning Florence-2 - Microsoft's Cutting-edge Vision Language Models](<https://devfeed.tech/articles/fine-tuning-florence-2-microsoft-s-cutting-edge-vision-language-models-7200.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/finetune-florence2>)

Author: Andres Marafioti; merve; Piotr Skalski

Published: 2024-06-24T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Computer vision](<https://devfeed.tech/topics/computer-vision.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>)

Tags: [collaboration](<https://devfeed.tech/tags/collaboration.md>), [community](<https://devfeed.tech/tags/community.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [object-detection](<https://devfeed.tech/tags/object-detection.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [transformer-architecture](<https://devfeed.tech/tags/transformer-architecture.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vqa](<https://devfeed.tech/tags/vqa.md>)

### AI overview

The article explains how to fine-tune Microsoft's Florence-2 vision-language model for DocVQA. It describes the model's sequence-to-sequence architecture, its large FLD-5B pre-training dataset, prompting experiments, and evaluation using Levenshtein similarity.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.