# ocr

Published articles for ocr.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## \[Aug 2026\] AI Community -- Activity Highlights and Achievements

DevFeed: [\[Aug 2026\] AI Community -- Activity Highlights and Achievements](<https://devfeed.tech/articles/aug-2026-ai-community-activity-highlights-and-achievements-41358.md>)

Original publisher: [Read original article](<https://medium.com/google-developer-experts/aug-2026-ai-community-activity-highlights-and-achievements-25e3b1ee42b1?source=rss----a67bd6fa7d58---4>)

Author: Nari Yoon

Published: 2026-09-17T05:12:15Z

Content type: article

Language: en

Sources: [Google Developer Experts - Medium](<https://devfeed.tech/sources/google-developer-experts-medium.md>)

Topics: [Google AI](<https://devfeed.tech/topics/google-ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [google-antigravity](<https://devfeed.tech/topics/google-antigravity.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [MCP Server](<https://devfeed.tech/topics/mcp-server.md>), [computer-use](<https://devfeed.tech/topics/computer-use.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>)

Tags: [2026](<https://devfeed.tech/tags/2026.md>), [agentic-ai](<https://devfeed.tech/tags/agentic-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-coding-agents](<https://devfeed.tech/tags/ai-coding-agents.md>), [ai-studio](<https://devfeed.tech/tags/ai-studio.md>), [antigravity](<https://devfeed.tech/tags/antigravity.md>), [api](<https://devfeed.tech/tags/api.md>), [automation](<https://devfeed.tech/tags/automation.md>), [community](<https://devfeed.tech/tags/community.md>), [computer-use](<https://devfeed.tech/tags/computer-use.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [google](<https://devfeed.tech/tags/google.md>), [google-ai](<https://devfeed.tech/tags/google-ai.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [mcp](<https://devfeed.tech/tags/mcp.md>), [mcp-server](<https://devfeed.tech/tags/mcp-server.md>), [multi-agent](<https://devfeed.tech/tags/multi-agent.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [paper](<https://devfeed.tech/tags/paper.md>), [pitfalls](<https://devfeed.tech/tags/pitfalls.md>)

### AI overview

A monthly roundup of Google AI community activities and achievements, covering Antigravity prototyping and engineering, AI coding agents, MCP-based remote control, computer-use agent orchestration, earthquake research, and TPU fine-tuning and migration guidance.

### Source excerpt

We love sharing the accomplishments of the Google AI communities over the month. We appreciate all the hard work and dedication of our community members. Without further ado, here are the key highlights by products! Agentic DevelopmentAntigravityPrototype App: OCR and Text Extraction by the author Prototyping and Bringing Ideas to Application Using Google AI Studio and Antigravity 2.0 by AI GDE Joan Santoso (Indonesia) shares a rapid prototyping workflow building an AI-powered Form Extractor using the Gemini API, featuring a lightweight OCR and text extraction workflow. Antigravity Engineering Series by GDE Amulya Bhatia (Germany) focuses on key features of Antigravity 2.0 across 10 articles covering topics such as multi-agent orchestration, safety architecture, and workflow automation, accompanied by source code examples. (image soruce) Remote Control for Google Antigravity: Drive Your AI Coding Agent From Telegram 🛰 by GDE Nicola Guglielmi (Italy) introduces an open-source MCP server that turns Telegram into a remote control surface for AI coding agents. Before the Quake: How Antigravity CLI's AI Agents & IoT Data Predict Earthquakes by GDE Kanshi Tanaike (Japan) introduces the paper establishing Unified LAIC-AGW Theory by integrating ultra-dense IoT weather data with seismic moment tensors. It demonstrates a pre-seismic early warning capability by capturing enthalpy anomalies and acoustic-gravity waves. ADKAI GDE Henry Ruiz (US) and AI GDE Margaret Maynard-Reid (US) AI GDE Henry Ruiz (US) and AI GDE Margaret Maynard-Reid (US) introduced UISurf: An Operator-Centric Multi-Agent Platform for Observable and Cross-Environment UI Automation at the Agentic AI Summit 2026. They highlighted how the model-agnostic framework leverages the Google Cloud and Gemini ecosystems, such as GEAP and ADK, to orchestrate and evaluate computer-use agents across web, desktop, and mobile environments. Frameworks and ResearchTPU Introduction to SFT on TPU with Tunix -- 10 pitfalls until 2

## PereStruct: Modular Pipeline and Dataset for Parsing Historical Newspapers

DevFeed: [PereStruct: Modular Pipeline and Dataset for Parsing Historical Newspapers](<https://devfeed.tech/articles/vlm-perestruct-24889.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1076770/>)

Author: makSShan (Яндекс, Yandex Cloud & Yandex Infrastructure)

Published: 2026-09-01T07:05:21Z

Content type: tutorial

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [яндекс](<https://devfeed.tech/topics/tag-4004cf5948d3.md>), [vlm](<https://devfeed.tech/topics/vlm.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-studio](<https://devfeed.tech/tags/ai-studio.md>), [bleu](<https://devfeed.tech/tags/bleu.md>), [computervision](<https://devfeed.tech/tags/computervision.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [perestruct](<https://devfeed.tech/tags/perestruct.md>), [rouge](<https://devfeed.tech/tags/rouge.md>), [tag-4004cf5948d3](<https://devfeed.tech/tags/tag-4004cf5948d3.md>), [tag-65c8d6d9e736](<https://devfeed.tech/tags/tag-65c8d6d9e736.md>), [tag-9bf5e01ce62e](<https://devfeed.tech/tags/tag-9bf5e01ce62e.md>), [tag-d27a0708d400](<https://devfeed.tech/tags/tag-d27a0708d400.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [yandex-ai-studio](<https://devfeed.tech/tags/yandex-ai-studio.md>), [yolo](<https://devfeed.tech/tags/yolo.md>)

### AI overview

The article presents PereStruct, a modular pipeline for reconstructing articles from historical newspaper scans. It combines YOLO-based layout detection, Yandex Vision OCR, Yandex AI Studio models for error correction, and a semantic model for assembling article blocks; the authors also publish code, an annotated dataset, and a benchmark.

### Source excerpt

Попробуйте открыть скан советской газеты и прочитать одну статью от начала до конца. Человек быстро замечает крупный заголовок, продолжение в соседней колонке и подпись под фотографией. Для алгоритма перед ним -- это выцветшая страница с десятками тесно расположенных прямоугольников, нестандартными шрифтами и неоднозначным порядком чтения. Даже если OCR правильно распознаёт почти все слова, на выходе ещё не получится документ. Нужно понять, какие фрагменты относятся к одной статье, где её начало, в каком порядке соединить блоки и что не следует включать в основной текст. Мы разработали PereStruct -- модульный пайплайн для разбора исторических газет. Он объединяет детектор вёрстки на базе YOLO, Yandex Vision OCR, коррекцию ошибок с помощью моделей Yandex AI Studio и отдельную модель семантической сборки статей. Вместе с кодом мы публикуем размеченный датасет и бенчмарк, чтобы другие команды могли изучать подход и ставить эксперименты на исторических документах. Читать далее

## Paperless-ngx 3.1.0 Adds AI Workflow Actions and Document Versioning

DevFeed: [Paperless-ngx 3.1.0 Adds AI Workflow Actions and Document Versioning](<https://devfeed.tech/articles/paperless-ngx-3-1-0-adds-ai-workflow-actions-and-document-versioning-10726.md>)

Original publisher: [Read original article](<https://selfhostlab.io/paperless-ngx-3-1-0-ai-workflow-versioning/>)

Author: Christian Rakoot

Published: 2026-08-31T06:39:07Z

Content type: news

Language: en

Sources: [Self Host Lab](<https://devfeed.tech/sources/self-host-lab.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [OpenID connect (OIDC)](<https://devfeed.tech/topics/oidc.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Authorization](<https://devfeed.tech/topics/authorization.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [bug](<https://devfeed.tech/tags/bug.md>), [changelog](<https://devfeed.tech/tags/changelog.md>), [compression](<https://devfeed.tech/tags/compression.md>), [identity](<https://devfeed.tech/tags/identity.md>), [news](<https://devfeed.tech/tags/news.md>), [news-personal-cloud](<https://devfeed.tech/tags/news-personal-cloud.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [paperless-ngx](<https://devfeed.tech/tags/paperless-ngx.md>), [personal-cloud](<https://devfeed.tech/tags/personal-cloud.md>), [release](<https://devfeed.tech/tags/release.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [updates](<https://devfeed.tech/tags/updates.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

Paperless-ngx 3.1.0 adds automatic AI-based workflow classification, per-document remote OCR selection, OIDC group synchronization, document versioning, configurable ZIP compression, and smaller fixes.

### Source excerpt

Paperless-ngx 3.1.0 landed on August 27, 2026, adding a workflow action that applies AI-generated tag and correspondent suggestions automatically, per-document selective remote OCR, OIDC group sync mapping identity-provider groups to superuser and staff roles, a new versioning system for merging documents into successive versions, configurable ZIP export compression, and numerous smaller UI bug fixes sitewide.

## \[Hands-on\] Turn Scientific Figures Into Structured Data with Mistral OCR

DevFeed: [\[Hands-on\] Turn Scientific Figures Into Structured Data with Mistral OCR](<https://devfeed.tech/articles/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-18234.md>)

Original publisher: [Read original article](<https://blog.dailydoseofds.com/p/hands-on-turn-scientific-figures>)

Author: Avi Chawla

Published: 2026-08-26T21:06:26Z

Content type: tutorial

Language: en

Sources: [Daily Dose of Data Science](<https://devfeed.tech/sources/daily-dose-of-data-science.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [incident](<https://devfeed.tech/topics/incident.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [FIRST](<https://devfeed.tech/topics/first.md>), [Google](<https://devfeed.tech/topics/google.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [code](<https://devfeed.tech/tags/code.md>), [mistral](<https://devfeed.tech/tags/mistral.md>), [ocr](<https://devfeed.tech/tags/ocr.md>)

### AI overview

The supplied excerpts discuss coordinating multiple agents and people through shared channels. They describe the limits of role-based handoffs, including duplicated work, lost context, and invisible negative results, and present Switch as a way to preserve shared reports and threads. The title and body appear to describe different subjects.

### Source excerpt

A full walkthrough of the extraction schema, with code.

## Scanning bus stop codes with ML Kit and Vision in the GalwayBus Compose Multiplatform app

DevFeed: [Scanning bus stop codes with ML Kit and Vision in the GalwayBus Compose Multiplatform app](<https://devfeed.tech/articles/scanning-bus-stop-codes-with-ml-kit-and-vision-in-the-galwaybus-compose-multiplatform-app-25195.md>)

Original publisher: [Read original article](<https://johnoreilly.dev/posts/galwaybus-scan-stop-kmp/>)

Published: 2026-08-06T23:00:00Z

Content type: tutorial

Language: en

Sources: [John O'Reilly](<https://devfeed.tech/sources/john-o-reilly.md>)

Topics: [compose-multiplatform](<https://devfeed.tech/topics/compose-multiplatform.md>), [ML Kit](<https://devfeed.tech/topics/ml-kit.md>), [Android](<https://devfeed.tech/topics/android.md>), [cameraX](<https://devfeed.tech/topics/camerax.md>), [iOS](<https://devfeed.tech/topics/ios.md>), [Code](<https://devfeed.tech/topics/code.md>), [ui](<https://devfeed.tech/topics/ui.md>)

Tags: [android](<https://devfeed.tech/tags/android.md>), [apple](<https://devfeed.tech/tags/apple.md>), [camera](<https://devfeed.tech/tags/camera.md>), [camerax](<https://devfeed.tech/tags/camerax.md>), [code](<https://devfeed.tech/tags/code.md>), [compose-multiplatform](<https://devfeed.tech/tags/compose-multiplatform.md>), [ios](<https://devfeed.tech/tags/ios.md>), [ml-kit](<https://devfeed.tech/tags/ml-kit.md>), [multiplatform](<https://devfeed.tech/tags/multiplatform.md>), [native](<https://devfeed.tech/tags/native.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [platform](<https://devfeed.tech/tags/platform.md>), [recognition](<https://devfeed.tech/tags/recognition.md>), [ui](<https://devfeed.tech/tags/ui.md>)

### AI overview

This tutorial explains how the GalwayBus app scans six-digit bus stop codes using on-device ML Kit on Android and Apple's Vision framework on iOS. Camera preview, matching logic, and UI are shared through Compose Multiplatform, while platform-specific implementations handle camera access and OCR. Recognized six-digit runs are matched exactly to stops, and overlapping recognition requests are avoided by dropping newer frames.

### Source excerpt

Every bus stop in Galway has a plate with a 6-digit stop code printed on it. We recently added a "Scan" tab to the GalwayBus app that lets you point the camera at that plate and jump straight to the stop's departures. The text recognition runs entirely on device (ML Kit on Android, Apple's Vision framework on iOS), with the camera preview, matching logic and UI all living in the shared Compose Multiplatform code.

## Newer Models, Same Advantage

DevFeed: [Newer Models, Same Advantage](<https://devfeed.tech/articles/newer-models-same-advantage-6998.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Dharma-AI/newer-models-same-advantages>)

Author: Erick Lachmann; Gabriel Pimenta de Freitas Cardoso; Francisco de Almeida Rocha Alves; Victor Gabriel Ferreira Barbosa

Published: 2026-07-16T11:49:48Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>)

Tags: [article](<https://devfeed.tech/tags/article.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data](<https://devfeed.tech/tags/data.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [errors](<https://devfeed.tech/tags/errors.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [generative](<https://devfeed.tech/tags/generative.md>), [inference](<https://devfeed.tech/tags/inference.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open](<https://devfeed.tech/tags/open.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [production](<https://devfeed.tech/tags/production.md>), [technology](<https://devfeed.tech/tags/technology.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

DharmaOCR is presented as a Brazilian Portuguese OCR model that outperformed newer alternatives through domain specialization and targeted training. Its two-stage pipeline combines supervised fine-tuning on Portuguese-language documents with Direct Preference Optimization, improving extraction quality, stability, inference efficiency, and production reliability.

### Source excerpt

Despite newer architectures, DharmaOCR outperformed Mistral OCR4 and Unlimited-OCR on Brazilian Portuguese through domain specialization and targeted training. This article presents the evidence and the mechanism behind that advantage. Three months ago, we published a paper on DharmaOCR and open-sourced one of the models. The objective was specific: optical character recognition engineered for Brazilian Portuguese. The training pipeline was built in two stages.

## Why I needed Durable Execution to read a toy manual

DevFeed: [Why I needed Durable Execution to read a toy manual](<https://devfeed.tech/articles/why-i-needed-durable-execution-to-read-a-toy-manual-36107.md>)

Original publisher: [Read original article](<https://temporal.io/blog/why-i-needed-durable-execution-to-read-a-toy-manual>)

Author: Shy Ruparel

Published: 2026-07-13T00:00:00Z

Content type: opinion

Language: en

Sources: [Temporal Blog](<https://devfeed.tech/sources/temporal-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Software](<https://devfeed.tech/topics/software.md>), [Python](<https://devfeed.tech/topics/python.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [cleanup](<https://devfeed.tech/tags/cleanup.md>), [japanese](<https://devfeed.tech/tags/japanese.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [pipeline](<https://devfeed.tech/tags/pipeline.md>), [retries](<https://devfeed.tech/tags/retries.md>), [temporal-voices](<https://devfeed.tech/tags/temporal-voices.md>), [translation](<https://devfeed.tech/tags/translation.md>)

### AI overview

Shy Ruparel describes building Toku Solutions, an AI pipeline that translates Japanese collectible toy manuals into editable static sites. Durable Execution handles OCR, retries, and cleanup so processing does not have to restart from the beginning.

### Source excerpt

Shy Ruparel built an AI pipeline to translate Japanese toy manuals. Durable Execution keeps OCR, retries, and cleanup from starting over.

## Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron

DevFeed: [Building an Analysis AI Agent for Industrial Alarm Management with NVIDIA Nemotron](<https://devfeed.tech/articles/building-an-analysis-ai-agent-for-industrial-alarm-management-with-nvidia-nemotron-6772.md>)

Original publisher: [Read original article](<https://developer.nvidia.com/blog/building-an-analysis-ai-agent-for-industrial-alarm-management-with-nvidia-nemotron/>)

Author: Tanya Lenz

Published: 2026-07-07T17:00:00Z

Content type: article

Language: en

Sources: [NVIDIA Developer](<https://devfeed.tech/sources/nvidia-developer.md>), [NVIDIA Technical Blog](<https://devfeed.tech/sources/nvidia-technical-blog.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agentic-ai-generative-ai](<https://devfeed.tech/tags/agentic-ai-generative-ai.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agent](<https://devfeed.tech/tags/ai-agent.md>), [build-ai-agents](<https://devfeed.tech/tags/build-ai-agents.md>), [developer-tools-techniques](<https://devfeed.tech/tags/developer-tools-techniques.md>), [featured](<https://devfeed.tech/tags/featured.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [industrial-digitalization-digital-twin](<https://devfeed.tech/tags/industrial-digitalization-digital-twin.md>), [llms](<https://devfeed.tech/tags/llms.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemo-retriever](<https://devfeed.tech/tags/nemo-retriever.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [openshell](<https://devfeed.tech/tags/openshell.md>)

### AI overview

The article describes an NVIDIA-based AI agent for analyzing industrial alarms. It gathers historical and playbook context, runs specialist checks such as anomaly detection and OCR, and returns structured recommendations through an HTTP endpoint.

### Source excerpt

Industrial machinery generates more alarms than technicians can triage. For each important alarm requiring follow-up, the technician pulls historical context,...

## Introducing Mistral OCR 4

DevFeed: [Introducing Mistral OCR 4](<https://devfeed.tech/articles/introducing-mistral-ocr-4-7098.md>)

Original publisher: [Read original article](<https://mistral.ai/news/ocr-4/>)

Published: 2026-06-23T12:00:48Z

Content type: article

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [document ai](<https://devfeed.tech/topics/document-ai.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>), [API](<https://devfeed.tech/topics/api.md>), [Security, Privacy and Abuse Prevention](<https://devfeed.tech/topics/security-privacy-and-abuse-prevention.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [document-ai](<https://devfeed.tech/tags/document-ai.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [privacy](<https://devfeed.tech/tags/privacy.md>), [rag](<https://devfeed.tech/tags/rag.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>)

### AI overview

Mistral OCR 4 is a document AI model for parsing documents with bounding boxes, typed block classification, inline confidence scores, and support for 170 languages. The article describes its benchmark performance, self-hosted deployment in a single container, API and Document AI access, and use in enterprise search, retrieval-augmented generation, and agentic workflows.

### Source excerpt

Mistral OCR 4 delivers enterprise document AI with 170-language support, bounding boxes, and self-hosted deployment.

## Direct Preference Optimization Beyond Chatbots

DevFeed: [Direct Preference Optimization Beyond Chatbots](<https://devfeed.tech/articles/direct-preference-optimization-beyond-chatbots-6992.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/Dharma-AI/direct-preference-optimization-beyond-chatbots>)

Author: Erick Lachmann; Gabriel Pimenta de Freitas Cardoso; Francisco de Almeida Rocha Alves

Published: 2026-06-03T12:55:11Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [dpo](<https://devfeed.tech/topics/dpo.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [cost](<https://devfeed.tech/tags/cost.md>), [dpo](<https://devfeed.tech/tags/dpo.md>), [efficiency](<https://devfeed.tech/tags/efficiency.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [model](<https://devfeed.tech/tags/model.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

The article presents Direct Preference Optimization (DPO) as a second training stage for reducing text degeneration in DharmaOCR, a structured OCR model. Applied after supervised fine-tuning, DPO reduced degeneration across every tested model family, with an average reduction of 59.4% and a best-case reduction of 87.6%.

### Source excerpt

In April, we released DharmaOCR, our specialized structured OCR model (available on Hugging Face) along with a paper detailing the methodology behind it and a benchmark demonstrating its superior quality and cost efficiency. The paper benchmarked leading vision-language model families - both open-source and commercial - on a structured document extraction task: OCR on Brazilian Portuguese text.

## PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend

DevFeed: [PaddleOCR 3.5: Running OCR and Document Parsing Tasks with a Transformers Backend](<https://devfeed.tech/articles/paddleocr-3-5-running-ocr-and-document-parsing-tasks-with-a-transformers-backend-7031.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/PaddlePaddle/paddleocr-transformers>)

Author: AlexZhang; Cuicheng; Jun Zhang; Manhui Lin; Yue Zhang

Published: 2026-05-18T15:12:46Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [document ai](<https://devfeed.tech/topics/document-ai.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>), [Retrieval Augmented Generation (RAG)](<https://devfeed.tech/topics/retrieval-augmented-generation-rag.md>)

Tags: [document-ai](<https://devfeed.tech/tags/document-ai.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [python](<https://devfeed.tech/tags/python.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [rag](<https://devfeed.tech/tags/rag.md>), [rocm](<https://devfeed.tech/tags/rocm.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

PaddleOCR 3.5 adds Transformers as a supported inference backend for OCR and document parsing models, while PaddleOCR continues to manage the underlying pipelines. The release simplifies integration with Hugging Face-centered environments and downstream document workflows such as RAG, search, analytics, and automation.

### Source excerpt

PaddleOCR continues to provide OCR model series such as PP-OCRv5 and document parsing model series such as PaddleOCR-VL 1.5, while Transformers becomes one of the supported backends for running them. Try the live demo on Hugging Face Spaces: PaddleOCR 3.5 introduces a more flexible inference-engine interface. Developers can select the backend through the parameter and pass backend-specific options through .

## Falcon Perception

DevFeed: [Falcon Perception](<https://devfeed.tech/articles/falcon-perception-7510.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tiiuae/falcon-perception>)

Author: Basma Boussaha; FalconPerception

Published: 2026-04-01T07:13:20Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Transformer](<https://devfeed.tech/topics/transformer.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [text-generation](<https://devfeed.tech/topics/text-generation.md>), [long-context](<https://devfeed.tech/topics/long-context.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [blog](<https://devfeed.tech/tags/blog.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [model](<https://devfeed.tech/tags/model.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>), [scale](<https://devfeed.tech/tags/scale.md>), [technology](<https://devfeed.tech/tags/technology.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

Falcon Perception is a 0.6B-parameter early-fusion Transformer for open-vocabulary grounding and segmentation from natural-language prompts. The article also introduces PBench, a diagnostic benchmark for perception capabilities and crowded long-context scenes, and Falcon OCR, a 0.3B-parameter open-source OCR model.

### Source excerpt

A Blog post by Technology Innovation Institute on Hugging Face

## Как маскировать персональные данные на изображениях: наш эксперимент с OCR и NER

DevFeed: [Как маскировать персональные данные на изображениях: наш эксперимент с OCR и NER](<https://devfeed.tech/articles/ocr-ner-23996.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/redmadrobot/articles/1011450/>)

Author: andrivasg (red\_mad\_robot)

Published: 2026-03-17T15:55:37Z

Content type: tutorial

Language: ru

Sources: [Redmadrobot EN](<https://devfeed.tech/sources/redmadrobot-en.md>), [Redmadrobot RU](<https://devfeed.tech/sources/redmadrobot-ru.md>)

Topics: [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [pii](<https://devfeed.tech/topics/pii.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [llm](<https://devfeed.tech/tags/llm.md>), [ner](<https://devfeed.tech/tags/ner.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [pii](<https://devfeed.tech/tags/pii.md>), [red-mad-robot](<https://devfeed.tech/tags/red-mad-robot.md>), [rnd](<https://devfeed.tech/tags/rnd.md>), [tag-601fbc7112a4](<https://devfeed.tech/tags/tag-601fbc7112a4.md>), [tag-7bc388df28ed](<https://devfeed.tech/tags/tag-7bc388df28ed.md>), [tag-9bf5e01ce62e](<https://devfeed.tech/tags/tag-9bf5e01ce62e.md>), [tag-b92bf5906bbd](<https://devfeed.tech/tags/tag-b92bf5906bbd.md>), [tag-ef0b1bf200df](<https://devfeed.tech/tags/tag-ef0b1bf200df.md>)

### AI overview

The article describes red_mad_robot's experiment using OCR combined with a NER model to detect and selectively mask personally identifiable information in images without training specialized visual detectors. On a dataset of 40 annotated images, the pipeline masked 90% of personal data while falsely masking 14% of text polygons; performance declined on difficult real-world photographs.

### Source excerpt

Всем привет! Меня зовут Андрей Иванов, я NLP-исследователь в R&D red_mad_robot. Мы разрабатываем систему Guardrails для защиты персональных данных (PII) и фильтрации небезопасного контента. В этой статье расскажу, как мы решали задачу точечного маскирования PII на картинках без обучения специальных визуальных детекторов. Разберём связку оптического распознавания символов (OCR) с NER-моделью, покажем метрики на реальных данных, раскроем ограничения подхода и наши решения для их преодоления. Читать далее

## Self-Hosted Paperless-ngx + Optional Local AI: Private Documents and Improves OCR & Search (Full Setup)

DevFeed: [Self-Hosted Paperless-ngx + Optional Local AI: Private Documents and Improves OCR & Search (Full Setup)](<https://devfeed.tech/articles/self-hosted-paperless-ngx-optional-local-ai-private-documents-and-improves-ocr-search-full-setup-10621.md>)

Original publisher: [Read original article](<https://technotim.com/posts/paperless-ngx-local-ai/>)

Author: Techno Tim

Published: 2026-01-27T13:00:00Z

Content type: tutorial

Language: en

Sources: [Techno Tim](<https://devfeed.tech/sources/techno-tim.md>)

Topics: [Docker Compose](<https://devfeed.tech/topics/docker-compose.md>), [Local AI](<https://devfeed.tech/topics/local-ai.md>), [Docker](<https://devfeed.tech/topics/docker.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [Database](<https://devfeed.tech/topics/database.md>), [Redis](<https://devfeed.tech/topics/redis.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [Server](<https://devfeed.tech/topics/server.md>), [GPU](<https://devfeed.tech/topics/gpu.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [compose](<https://devfeed.tech/tags/compose.md>), [docker](<https://devfeed.tech/tags/docker.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [homelab](<https://devfeed.tech/tags/homelab.md>), [linux](<https://devfeed.tech/tags/linux.md>), [local](<https://devfeed.tech/tags/local.md>), [local-ai](<https://devfeed.tech/tags/local-ai.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [paperless-ngx](<https://devfeed.tech/tags/paperless-ngx.md>), [redis](<https://devfeed.tech/tags/redis.md>), [search](<https://devfeed.tech/tags/search.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>)

### AI overview

A tutorial for deploying a self-hosted Paperless-ngx document management stack with Docker Compose. It explains the core OCR, indexing, tagging, and search workflow, then adds optional local AI through Ollama, Open WebUI, Paperless-AI, and Paperless-GPT to improve OCR and metadata suggestions without using cloud services.

### Source excerpt

Build a complete Paperless-ngx stack in Docker and take control of your documents. We'll get Paperless running first (works great on its own), then optionally add local AI with Ollama + Open WebUI and upgrade OCR using Paperless-GPT and Paperless-AI for more accurate, searchable text and tags - no cloud required. This post walks through a complete, repeatable Docker Compose for Paperless-ngx, ...

## NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI

DevFeed: [NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI](<https://devfeed.tech/articles/nvidia-cosmos-reason-2-brings-advanced-reasoning-to-physical-ai-7400.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/nvidia-cosmos-reason-2-brings-advanced-reasoning>)

Author: Tsung-Yi Lin; Debraj Sinha

Published: 2026-01-05T22:56:51Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Cosmos](<https://devfeed.tech/topics/cosmos.md>), [Physical AI](<https://devfeed.tech/topics/physical-ai.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Video Analytics](<https://devfeed.tech/topics/video-analytics.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [data](<https://devfeed.tech/topics/data.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [analytics](<https://devfeed.tech/tags/analytics.md>), [cloud](<https://devfeed.tech/tags/cloud.md>), [cosmos](<https://devfeed.tech/tags/cosmos.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [language-models](<https://devfeed.tech/tags/language-models.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [models](<https://devfeed.tech/tags/models.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open](<https://devfeed.tech/tags/open.md>), [performance](<https://devfeed.tech/tags/performance.md>), [physical-ai](<https://devfeed.tech/tags/physical-ai.md>), [physics](<https://devfeed.tech/tags/physics.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

NVIDIA released Cosmos Reason 2, an open reasoning vision-language model for physical AI. The model is designed to help robots and AI agents understand, plan, and act in the physical world, with improved spatio-temporal reasoning, visual perception, OCR, long-context input, and deployment from edge to cloud.

### Source excerpt

NVIDIA today released Cosmos Reason 2, the latest advancement in open, reasoning vision language models for physical AI. Cosmos Reason 2 surpasses its previous version in accuracy and tops the Physical AI Bench and Physical Reasoning leaderboards as the #1 open model for visual understanding. Since their introduction, vision-language models have rapidly improved at tasks like object and pattern recognition in images.

## Introducing Mistral OCR 3

DevFeed: [Introducing Mistral OCR 3](<https://devfeed.tech/articles/introducing-mistral-ocr-3-7070.md>)

Original publisher: [Read original article](<https://mistral.ai/news/mistral-ocr-3/>)

Published: 2025-12-17T15:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [document-ai](<https://devfeed.tech/tags/document-ai.md>), [html](<https://devfeed.tech/tags/html.md>), [json](<https://devfeed.tech/tags/json.md>), [models](<https://devfeed.tech/tags/models.md>), [ocr](<https://devfeed.tech/tags/ocr.md>)

### AI overview

Mistral introduces OCR 3, a document-processing model for extracting text and embedded images, reconstructing tables, and producing Markdown or structured JSON. The article describes benchmark comparisons with ground truth and API integration for developers.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Supercharge your OCR Pipelines with Open Models

DevFeed: [Supercharge your OCR Pipelines with Open Models](<https://devfeed.tech/articles/supercharge-your-ocr-pipelines-with-open-models-7406.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ocr-open-models>)

Author: merve; Aritra Roy Gosthipaty; Daniel van Strien; Hynek Kydlicek; Andres Marafioti; Vaibhav Srivastav; Pedro Cuenca

Published: 2025-10-21T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [document ai](<https://devfeed.tech/topics/document-ai.md>), [Computer vision](<https://devfeed.tech/topics/computer-vision.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [document-ai](<https://devfeed.tech/tags/document-ai.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [html](<https://devfeed.tech/tags/html.md>), [llm](<https://devfeed.tech/tags/llm.md>), [markdown](<https://devfeed.tech/tags/markdown.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pipelines](<https://devfeed.tech/tags/pipelines.md>), [qa](<https://devfeed.tech/tags/qa.md>), [retrieval](<https://devfeed.tech/tags/retrieval.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlm](<https://devfeed.tech/tags/vlm.md>)

### AI overview

This guide surveys open-weight OCR and vision-language models for document AI. It explains their capabilities, output formats, multimodal document retrieval, document question answering, and the tradeoffs between fine-tuning and using models out of the box.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## SOTA OCR with Core ML and dots.ocr

DevFeed: [SOTA OCR with Core ML and dots.ocr](<https://devfeed.tech/articles/sota-ocr-with-core-ml-and-dots-ocr-7174.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/dots-ocr-ne>)

Author: Christopher Fleetwood; Pedro Cuenca

Published: 2025-10-02T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [MLX](<https://devfeed.tech/topics/mlx.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [apple](<https://devfeed.tech/tags/apple.md>), [architecture](<https://devfeed.tech/tags/architecture.md>), [battery](<https://devfeed.tech/tags/battery.md>), [coreml](<https://devfeed.tech/tags/coreml.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [developers](<https://devfeed.tech/tags/developers.md>), [framework](<https://devfeed.tech/tags/framework.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [high-performance](<https://devfeed.tech/tags/high-performance.md>), [images](<https://devfeed.tech/tags/images.md>), [mlx](<https://devfeed.tech/tags/mlx.md>), [model](<https://devfeed.tech/tags/model.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [parameter](<https://devfeed.tech/tags/parameter.md>), [precision](<https://devfeed.tech/tags/precision.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [repo](<https://devfeed.tech/tags/repo.md>), [tools](<https://devfeed.tech/tags/tools.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

This tutorial explains how to convert dots.ocr from PyTorch to Core ML for on-device execution on Apple hardware. It discusses the roles of the Neural Engine, GPU, MLX, and Core ML, then outlines a staged conversion process beginning with GPU execution, FLOAT32 precision, and static shapes.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## 📚 3LM: A Benchmark for Arabic LLMs in STEM and Code

DevFeed: [📚 3LM: A Benchmark for Arabic LLMs in STEM and Code](<https://devfeed.tech/articles/3lm-a-benchmark-for-arabic-llms-in-stem-and-code-7503.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tiiuae/3lm-benchmark>)

Author: Basma Boussaha; Leen AlQadi; Mughaira; Shaikha Alsuwaidi; Giulia Campesan; Ahmed Alzubaidi; Mohammed Alyafeai; Hakim Hacid

Published: 2025-08-01T14:25:21Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Large language models (LLMs)](<https://devfeed.tech/topics/large-language-models-llms.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Code generation](<https://devfeed.tech/topics/code-generation.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [code](<https://devfeed.tech/tags/code.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [llms](<https://devfeed.tech/tags/llms.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [quality-assurance](<https://devfeed.tech/tags/quality-assurance.md>), [stem](<https://devfeed.tech/tags/stem.md>)

### AI overview

The article introduces 3LM, a benchmark for evaluating Arabic Large Language Models on STEM subjects, structured reasoning, formal logic, and code generation. It combines native educational multiple-choice questions, synthetic high-difficulty STEM questions, and translated code-generation tasks.

### Source excerpt

A Blog post by Technology Innovation Institute on Hugging Face

## Welcome the NVIDIA Llama Nemotron Nano VLM to Hugging Face Hub

DevFeed: [Welcome the NVIDIA Llama Nemotron Nano VLM to Hugging Face Hub](<https://devfeed.tech/articles/welcome-the-nvidia-llama-nemotron-nano-vlm-to-hugging-face-hub-7384.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/nvidia/llama-nemotron-nano-vl>)

Author: Amanda Saunders; Amala Sanjay Deshmukh; Kateryna Chumachenko; Annie Surla; Karan; Tuomas Rintamaki; Matthieu Le; Yu Yao; Chen Cui; Timo Roman; Zhiding Yu; Mike Ranzinger

Published: 2025-06-27T21:09:27Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [vlm](<https://devfeed.tech/topics/vlm.md>), [llama](<https://devfeed.tech/topics/llama.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>), [Nvidia](<https://devfeed.tech/topics/nvidia.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [NeMo](<https://devfeed.tech/topics/nemo.md>), [datasets](<https://devfeed.tech/topics/datasets.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [data](<https://devfeed.tech/tags/data.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [idp](<https://devfeed.tech/tags/idp.md>), [llama](<https://devfeed.tech/tags/llama.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nemo](<https://devfeed.tech/tags/nemo.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [performance](<https://devfeed.tech/tags/performance.md>), [recognition](<https://devfeed.tech/tags/recognition.md>), [train](<https://devfeed.tech/tags/train.md>), [use-cases](<https://devfeed.tech/tags/use-cases.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vlm](<https://devfeed.tech/tags/vlm.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

NVIDIA Llama Nemotron Nano VL is an 8B vision-language model for intelligent document processing. Available on Hugging Face, it extracts and interprets text, tables, charts, diagrams, and other information from complex documents.

### Source excerpt

NVIDIA Llama Nemotron Nano VL is a state-of-the-art 8B Vision Language Model (VLM) designed for intelligent document processing, offering high accuracy and multimodal understanding. Available on Hugging Face, it excels in extracting and understanding information from complex documents like invoices, receipts, contracts, and more.

## AI Workflows for Docs: Putting Devin to Work

DevFeed: [AI Workflows for Docs: Putting Devin to Work](<https://devfeed.tech/articles/ai-workflows-for-docs-putting-devin-to-work-4957.md>)

Original publisher: [Read original article](<https://neon.com/blog/ai-workflows-for-docs-devin>)

Author: Daniel Price

Published: 2025-06-23T17:10:34Z

Content type: article

Language: en

Sources: [Blog -- Neon Docs](<https://devfeed.tech/sources/blog-neon-docs.md>)

Topics: [AI Bots](<https://devfeed.tech/topics/ai-bots.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [changelog](<https://devfeed.tech/tags/changelog.md>), [docs](<https://devfeed.tech/tags/docs.md>), [github](<https://devfeed.tech/tags/github.md>), [github-issues](<https://devfeed.tech/tags/github-issues.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [product](<https://devfeed.tech/tags/product.md>), [pull-request](<https://devfeed.tech/tags/pull-request.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

Neon's documentation team describes using Devin, an AI agent that opens pull requests, to automate repetitive documentation workflows. Examples include setting up changelogs, addressing GitHub issues and PRs, and adding a relevant repository image while documenting a feature.

### Source excerpt

At Neon, our docs team does a little bit of everything. We work on technical documentation, sometimes UI copy, changelogs, reviews, and the occasional regex-heavy cleanup across hundreds of pages. It's a lot of small, steady work - often, exactly the kind of work you wish an AI a...

## Finetuning olmOCR to be a faithful OCR-Engine

DevFeed: [Finetuning olmOCR to be a faithful OCR-Engine](<https://devfeed.tech/articles/finetuning-olmocr-to-be-a-faithful-ocr-engine-7515.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tngtech/finetuning-olmocr-to-be-a-faithful-ocr-engine>)

Author: Johannes EsslingerTNG; Innovation Hacking

Published: 2025-04-22T18:33:09Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Training AI Models](<https://devfeed.tech/topics/training-ai-models.md>), [Language models](<https://devfeed.tech/topics/language-models.md>)

Tags: [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gradient-accumulation](<https://devfeed.tech/tags/gradient-accumulation.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>)

### AI overview

The article describes fine-tuning olmOCR to retain header and footer information that its original training data intentionally omitted. The authors generate an 8,000-document dataset with Qwen2.5-VL-72B-Instruct, train using the open-source olmOCR pipeline, and evaluate on a customized dataset containing header and footer content.

### Source excerpt

A Blog post by TNG Technology Consulting GmbH on Hugging Face

## Visual Salamandra: Pushing the Boundaries of Multimodal Understanding

DevFeed: [Visual Salamandra: Pushing the Boundaries of Multimodal Understanding](<https://devfeed.tech/articles/visual-salamandra-pushing-the-boundaries-of-multimodal-understanding-6990.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/BSC-LT/visualsalamandra7b>)

Author: Iñigo Pikabea; Jaume Lozano

Published: 2025-04-11T14:21:56Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [multimodal-ai](<https://devfeed.tech/topics/multimodal-ai.md>), [vlm](<https://devfeed.tech/topics/vlm.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Embeddings](<https://devfeed.tech/topics/embeddings.md>), [mlp](<https://devfeed.tech/topics/mlp.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [blog](<https://devfeed.tech/tags/blog.md>), [blog-post](<https://devfeed.tech/tags/blog-post.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [embeddings](<https://devfeed.tech/tags/embeddings.md>), [grounding](<https://devfeed.tech/tags/grounding.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [language](<https://devfeed.tech/tags/language.md>), [llm](<https://devfeed.tech/tags/llm.md>), [mlp](<https://devfeed.tech/tags/mlp.md>), [models](<https://devfeed.tech/tags/models.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [multimodal-ai](<https://devfeed.tech/tags/multimodal-ai.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [research](<https://devfeed.tech/tags/research.md>), [training](<https://devfeed.tech/tags/training.md>), [vision](<https://devfeed.tech/tags/vision.md>), [vqa](<https://devfeed.tech/tags/vqa.md>)

### AI overview

Visual Salamandra is a multilingual multimodal model built by extending the Salamandra Instructed 7B model with Google's SigLIP image encoder, an MLP projector, and late-fusion techniques. It processes text, images, and videos, with training focused on visual grounding, document understanding, mathematical reasoning, OCR, and European-language coverage.

### Source excerpt

A Blog post by Language Technologies Laboratory @ Barcelona Supercomputing Center on Hugging Face

## Как AI-агенты ускоряют работу девелопера: автоматизация данных и управление знаниями

DevFeed: [Как AI-агенты ускоряют работу девелопера: автоматизация данных и управление знаниями](<https://devfeed.tech/articles/ai-24011.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/redmadrobot/articles/892882/>)

Author: redmadrobot (red\_mad\_robot)

Published: 2025-03-26T12:01:58Z

Content type: article

Language: ru

Sources: [Redmadrobot EN](<https://devfeed.tech/sources/redmadrobot-en.md>), [Redmadrobot RU](<https://devfeed.tech/sources/redmadrobot-ru.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [genai](<https://devfeed.tech/topics/genai.md>), [pdf](<https://devfeed.tech/topics/pdf.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-e2239b5ae8fa](<https://devfeed.tech/tags/ai-e2239b5ae8fa.md>), [genai](<https://devfeed.tech/tags/genai.md>), [ocr](<https://devfeed.tech/tags/ocr.md>), [pdf](<https://devfeed.tech/tags/pdf.md>), [rag](<https://devfeed.tech/tags/rag.md>), [ragas](<https://devfeed.tech/tags/ragas.md>), [red-mad-robot](<https://devfeed.tech/tags/red-mad-robot.md>), [router](<https://devfeed.tech/tags/router.md>), [tag-9d8cf70dc46c](<https://devfeed.tech/tags/tag-9d8cf70dc46c.md>), [tag-b6914c0b0244](<https://devfeed.tech/tags/tag-b6914c0b0244.md>), [tag-e0bf7a3b3c2d](<https://devfeed.tech/tags/tag-e0bf7a3b3c2d.md>), [tag-e95ee8b94dca](<https://devfeed.tech/tags/tag-e95ee8b94dca.md>), [vision](<https://devfeed.tech/tags/vision.md>), [workflow](<https://devfeed.tech/tags/workflow.md>)

### AI overview

The NDT by red_mad_robot team describes a smart platform with two workflow AI agents for the FSK Group. One agent answers counterparties' frequently asked questions, while the other helps employees search internal documents and access the corporate knowledge base. The article discusses document processing, including PDF extraction, OCR, and vision-language models, with an emphasis on reducing incorrect AI responses.

### Source excerpt

Привет! На связи команда NDT by red_mad_robot. Рассказываем, как создавали смарт-платформу с двумя AI-агентами для группы компаний ФСК -- одного из крупнейших российских девелоперов. Система автоматизировала работу с данными и значительно снизила нагрузку на сотрудников технической поддержки и коммерческого департамента. Читать далее

[Next page](<https://devfeed.tech/tags/ocr.md?cursor=WyIyMDI1LTAzLTI2VDEyOjAxOjU4KzAwOjAwIiwgIjk4NjFmNzVmLTY4MTAtNDdmNi05MjBjLTgyNGU2YzYyMjFmMCJd>)