# asr

Published articles for asr.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## The Open ASR Leaderboard Adds Its First Global South Language

DevFeed: [The Open ASR Leaderboard Adds Its First Global South Language](<https://devfeed.tech/articles/the-open-asr-leaderboard-adds-its-first-global-south-language-7411.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard-global-south>)

Author: Eric Bezzam; Shobhit Banga; Manas Dhir; Bhaskar Singh; Manmeet Kaur; Aaditya Pareek; Walecha; Sagar Jain; Hanuman Sidh; Vanshika Chhabra

Published: 2026-08-28T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [contributors](<https://devfeed.tech/tags/contributors.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [devices](<https://devfeed.tech/tags/devices.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>)

### AI overview

The Open ASR Leaderboard introduces Monsoon evaluation sets for Hindi in India, expanding coverage beyond European languages and testing how recognition performance varies across populations and conditions. The sets use public and private splits, speaker-disjoint data, detailed speaker attributes, and variation in geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and transcript validity.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Measuring benchmark optimization in speech recognition

DevFeed: [Measuring benchmark optimization in speech recognition](<https://devfeed.tech/articles/measuring-benchmark-optimization-in-speech-recognition-7104.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/asr-benchmark-optimization>)

Author: Theo Lebryk; Eric Bezzam; Alice; David Ayllon; Jakub Piotr Cłapa; Jens Madsen; Panagiotis Tzirakis

Published: 2026-08-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmark overfitting machine learning](<https://devfeed.tech/topics/benchmark-overfitting-machine-learning.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [errors](<https://devfeed.tech/tags/errors.md>), [measurement](<https://devfeed.tech/tags/measurement.md>), [model](<https://devfeed.tech/tags/model.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

The article examines benchmark optimization, or "benchmaxxing," in speech recognition. It presents three tests and evaluates 11 open-source ASR models, finding that some reproduced benchmark transcripts even when the audio contradicted them. The research also uses model ensembles and human annotations to identify and validate likely benchmark errors.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## AI Model Risk Intelligence Know Which Models You Can Trust Before You Deploy

DevFeed: [AI Model Risk Intelligence Know Which Models You Can Trust Before You Deploy](<https://devfeed.tech/articles/ai-model-risk-intelligence-know-which-models-you-can-trust-before-you-deploy-8253.md>)

Original publisher: [Read original article](<https://snyk.io/blog/why-we-rebuilt-evo-ai-model-risk-scoring/>)

Author: Ranko Cupovic

Published: 2026-08-04T04:00:00Z

Content type: article

Language: en

Sources: [Blog RSS Feed | Snyk](<https://devfeed.tech/sources/blog-rss-feed-snyk.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Adversarial attacks](<https://devfeed.tech/topics/adversarial-attacks.md>), [Security](<https://devfeed.tech/topics/security.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [pii](<https://devfeed.tech/topics/pii.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [application-security](<https://devfeed.tech/tags/application-security.md>), [asr](<https://devfeed.tech/tags/asr.md>), [attacks](<https://devfeed.tech/tags/attacks.md>), [awareness](<https://devfeed.tech/tags/awareness.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [blog](<https://devfeed.tech/tags/blog.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [developer](<https://devfeed.tech/tags/developer.md>), [devops](<https://devfeed.tech/tags/devops.md>), [enablement](<https://devfeed.tech/tags/enablement.md>), [model](<https://devfeed.tech/tags/model.md>), [pii](<https://devfeed.tech/tags/pii.md>), [safety](<https://devfeed.tech/tags/safety.md>), [security](<https://devfeed.tech/tags/security.md>), [snyk-platform](<https://devfeed.tech/tags/snyk-platform.md>), [snyk-security-intel](<https://devfeed.tech/tags/snyk-security-intel.md>)

### AI overview

Evo's AI model risk score combines attack success rate and attack impact into a 0-1000 score. It is based on adversarial testing against standard system-prompt hardening and breaks risk down by attacker goals such as PII extraction, system-prompt extraction, and insecure code generation to support deployment decisions.

### Source excerpt

AI model risk depends on how a model is deployed. Learn how Evo combines adversarial testing, attack impact, and deployment context to help teams compare models and enforce policy.

## OpenAI sells a $230 keypad to run AI agents. I built one for $28 on Linux.

DevFeed: [OpenAI sells a $230 keypad to run AI agents. I built one for $28 on Linux.](<https://devfeed.tech/articles/openai-sells-a-230-keypad-to-run-ai-agents-i-built-one-for-28-on-linux-25179.md>)

Original publisher: [Read original article](<https://www.ivanmorgillo.com/2026/07/29/openai-codex-micro-28-dollar-linux-macropad/>)

Author: Ivan Morgillo

Published: 2026-07-29T10:00:00Z

Content type: tutorial

Language: en

Sources: [Ivan Morgillo](<https://devfeed.tech/sources/ivan-morgillo.md>)

Topics: [codex](<https://devfeed.tech/topics/codex.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [OpenAI](<https://devfeed.tech/topics/openai.md>), [Linux](<https://devfeed.tech/topics/linux.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Nemotron](<https://devfeed.tech/topics/nemotron.md>)

Tags: [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [ai-coding-agents](<https://devfeed.tech/tags/ai-coding-agents.md>), [asr](<https://devfeed.tech/tags/asr.md>), [codex](<https://devfeed.tech/tags/codex.md>), [gpu](<https://devfeed.tech/tags/gpu.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [linux](<https://devfeed.tech/tags/linux.md>), [nemotron](<https://devfeed.tech/tags/nemotron.md>), [openai](<https://devfeed.tech/tags/openai.md>)

### AI overview

The article explains how the author recreated the core workflow of OpenAI's Codex Micro for $28 on Linux using a low-cost macropad. The setup supports voice dictation through a local speech-to-text server and one-press commands for interacting with coding agents.

### Source excerpt

OpenAI's first hardware is a $230 keypad for driving AI coding agents. I rebuilt the part that matters -- voice dictation and one-press agent commands -- for $28, on Linux, with a cheap Chinese macropad.

## Adding new quick commands to a smart speaker without degrading existing commands

DevFeed: [Adding new quick commands to a smart speaker without degrading existing commands](<https://devfeed.tech/articles/article-24871.md>)

Original publisher: [Read original article](<https://habr.com/ru/companies/yandex/articles/1061968/>)

Author: khaymon (Яндекс)

Published: 2026-07-23T09:00:04Z

Content type: tutorial

Language: ru

Sources: [Яндекс - Как мы делаем Яндекс / Статьи](<https://devfeed.tech/sources/source.md>)

Topics: [яндекс](<https://devfeed.tech/topics/tag-4004cf5948d3.md>), [asr](<https://devfeed.tech/topics/asr.md>), [cpu](<https://devfeed.tech/topics/cpu.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [continual-learning](<https://devfeed.tech/tags/continual-learning.md>), [cpu](<https://devfeed.tech/tags/cpu.md>), [deep-learning](<https://devfeed.tech/tags/deep-learning.md>), [keyword-spotting](<https://devfeed.tech/tags/keyword-spotting.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [speech-processing](<https://devfeed.tech/tags/speech-processing.md>), [tag-355bb785df82](<https://devfeed.tech/tags/tag-355bb785df82.md>), [tag-4004cf5948d3](<https://devfeed.tech/tags/tag-4004cf5948d3.md>), [tag-8483d32db3b1](<https://devfeed.tech/tags/tag-8483d32db3b1.md>), [tag-967d8467ce56](<https://devfeed.tech/tags/tag-967d8467ce56.md>), [tag-9c1e53a23032](<https://devfeed.tech/tags/tag-9c1e53a23032.md>), [tag-a1312fd2c7ff](<https://devfeed.tech/tags/tag-a1312fd2c7ff.md>), [tag-d14eb265d33e](<https://devfeed.tech/tags/tag-d14eb265d33e.md>), [tag-d346fb5ae499](<https://devfeed.tech/tags/tag-d346fb5ae499.md>)

### AI overview

The article discusses Yandex's compact on-device neural model for recognizing Alice quick commands. It covers adding track-rating and Bluetooth commands while aiming to preserve performance on existing commands and limit device resource use.

### Source excerpt

Чтобы дать команду умной колонке, не обязательно говорить активационное слово "Алиса": есть быстрые команды -- короткие фразы, с помощью которых можно управлять музыкой, громкостью или умным домом. Например, чтобы переключить трек, достаточно просто сказать "дальше", а чтобы убавить звук -- "тише". Весь список команд можно посмотреть в настройках вашего аккаунта в приложении "Дом с Алисой". Быстрые команды удобнее не только пользователям, но и системе: запросы через слово "Алиса" требуют обращения к модели распознавания речи ASR, которой из-за её размеров необходимы серверные вычислительные ресурсы, а модель быстрых команд устроена гораздо компактнее. Она работает прямо на устройстве, а значит, ограничена вычислительными ресурсами самой колонки -- её CPU и оперативной памятью. Из-за этого модель нельзя сильно увеличить: ей приходится оставаться компактной, зато запрос обрабатывается быстрее. За распознавание быстрых команд отвечает нейросеть. Её архитектура почти полностью совпадает с решением для наушников Яндекс Дропс, которое подробно описал в своей статье Григорий Афанасенко. Разница в основном в масштабе: наша модель весит всего от 0,5 до 1,5 МБ в зависимости от железа конкретного устройства. Со временем перед нами встала задача добавить к базовым командам "лайк" и "дизлайк" для управления треками, а также команды "включи блютус" и "выключи блютус". Особенно это актуально для Станции Стрит, которую часто берут с собой на природу, где нет интернета. Но главным было гарантировать абсолютное отсутствие ухудшения на уже запущенных командах и не слишком сильно увеличивать потребление ресурсов на устройстве. Читать далее

## Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World

DevFeed: [Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World](<https://devfeed.tech/articles/introducing-the-ffasr-leaderboard-benchmarking-asr-in-the-real-world-7196.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ffasr-leaderboard>)

Author: Daniel Gert Nielsen; Shivam Saini; Alessia Milo; Georg Götz; Eric Bezzam

Published: 2026-06-24T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [community](<https://devfeed.tech/tags/community.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [development](<https://devfeed.tech/tags/development.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [performance](<https://devfeed.tech/tags/performance.md>), [speech](<https://devfeed.tech/tags/speech.md>)

### AI overview

Treble Technologies and Hugging Face introduce the open, community-driven FFASR Leaderboard, a benchmark for evaluating automatic speech recognition models in realistic far-field acoustic conditions. It measures performance across factors such as reverberation, background noise, microphone distance, and low signal-to-noise ratios, while also showing the tradeoff between recognition accuracy and speed.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Experimenting with the proposed Cross-Origin Storage API in Transformers.js

DevFeed: [Experimenting with the proposed Cross-Origin Storage API in Transformers.js](<https://devfeed.tech/articles/experimenting-with-the-proposed-cross-origin-storage-api-in-transformers-js-7152.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/cross-origin-storage>)

Author: Thomas Steiner

Published: 2026-06-23T00:00:00Z

Content type: tutorial

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [api](<https://devfeed.tech/tags/api.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [asr](<https://devfeed.tech/tags/asr.md>), [browser](<https://devfeed.tech/tags/browser.md>), [cache](<https://devfeed.tech/tags/cache.md>), [chrome](<https://devfeed.tech/tags/chrome.md>), [chrome-extension](<https://devfeed.tech/tags/chrome-extension.md>), [cos](<https://devfeed.tech/tags/cos.md>), [cross-origin-storage](<https://devfeed.tech/tags/cross-origin-storage.md>), [inference](<https://devfeed.tech/tags/inference.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [on-device](<https://devfeed.tech/tags/on-device.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [storage](<https://devfeed.tech/tags/storage.md>), [transformers](<https://devfeed.tech/tags/transformers.md>), [transformers-js](<https://devfeed.tech/tags/transformers-js.md>), [wasm](<https://devfeed.tech/tags/wasm.md>), [web](<https://devfeed.tech/tags/web.md>), [web-apps](<https://devfeed.tech/tags/web-apps.md>), [web-developers](<https://devfeed.tech/tags/web-developers.md>), [web-standards](<https://devfeed.tech/tags/web-standards.md>), [webassembly](<https://devfeed.tech/tags/webassembly.md>)

### AI overview

A tutorial that uses Transformers.js browser inference examples to examine duplicate model and WebAssembly downloads across origins and the proposed Cross-Origin Storage API.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Adding Benchmaxxer Repellant to the Open ASR Leaderboard

DevFeed: [Adding Benchmaxxer Repellant to the Open ASR Leaderboard](<https://devfeed.tech/articles/adding-benchmaxxer-repellant-to-the-open-asr-leaderboard-7413.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard-private-data>)

Author: Eric Bezzam; Steven Zheng; Eustache Le Bihan; Sergio Bruccoleri; Jeanine Sinanan-Singh; Casey Ford; Guanbo Wang; Yukai Huang; Ke Li; Yufeng Hao

Published: 2026-05-06T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The article announces private English ASR datasets from Appen and DataoceanAI for the Open ASR Leaderboard. Keeping the datasets private is intended to reduce benchmark-specific optimization and test-set contamination while preserving a high-quality evaluation of multiple speech-recognition tasks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## WAXAL: A large-scale open resource for African language speech technology

DevFeed: [WAXAL: A large-scale open resource for African language speech technology](<https://devfeed.tech/articles/waxal-a-large-scale-open-resource-for-african-language-speech-technology-6927.md>)

Original publisher: [Read original article](<https://research.google/blog/waxal-a-large-scale-open-resource-for-african-language-speech-technology/>)

Published: 2026-03-06T20:06:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [data](<https://devfeed.tech/topics/data.md>), [Google](<https://devfeed.tech/topics/google.md>), [AI Research](<https://devfeed.tech/topics/ai-research.md>), [systems](<https://devfeed.tech/topics/systems.md>)

Tags: [africa](<https://devfeed.tech/tags/africa.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-research](<https://devfeed.tech/tags/ai-research.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [google](<https://devfeed.tech/tags/google.md>), [natural-language-processing](<https://devfeed.tech/tags/natural-language-processing.md>), [open-source-models-datasets](<https://devfeed.tech/tags/open-source-models-datasets.md>), [research](<https://devfeed.tech/tags/research.md>), [resources](<https://devfeed.tech/tags/resources.md>), [speech](<https://devfeed.tech/tags/speech.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [voice](<https://devfeed.tech/tags/voice.md>)

### AI overview

Google Research introduces WAXAL, an open-access speech dataset covering 27 Sub-Saharan African languages. The release includes approximately 1,846 hours of transcribed ASR data and more than 565 hours of high-fidelity TTS recordings under a CC-BY-4.0 license.

### Source excerpt

Natural Language Processing

## HarborFM - Self-Hosted Podcast Creator

DevFeed: [HarborFM - Self-Hosted Podcast Creator](<https://devfeed.tech/articles/harborfm-self-hosted-podcast-creator-10731.md>)

Original publisher: [Read original article](<https://noted.lol/harborfm/>)

Author: Logan Rickert

Published: 2026-02-20T18:37:34Z

Content type: article

Language: en

Sources: [Noted](<https://devfeed.tech/sources/noted.md>)

Topics: [Open Source](<https://devfeed.tech/topics/open-source.md>), [App](<https://devfeed.tech/topics/app.md>), [RSS Feed](<https://devfeed.tech/topics/rss-feed.md>), [PWA](<https://devfeed.tech/topics/pwa.md>), [WebRTC](<https://devfeed.tech/topics/webrtc.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [podcast](<https://devfeed.tech/tags/podcast.md>), [rss](<https://devfeed.tech/tags/rss.md>), [self-hosted](<https://devfeed.tech/tags/self-hosted.md>), [self-hosted-media-streaming-audio-streaming](<https://devfeed.tech/tags/self-hosted-media-streaming-audio-streaming.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [vision](<https://devfeed.tech/tags/vision.md>), [webrtc](<https://devfeed.tech/tags/webrtc.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

HarborFM is presented as a free, open-source, self-hosted podcast creation tool. It lets users assemble episodes from reusable audio segments, publish per-show RSS feeds, record remote guests through WebRTC, collaborate on episodes, manage private feeds, and view analytics. It also supports podcast chapters, Whisper-based transcription, and AI summaries.

### Source excerpt

Take back your podcast data! HarborFM is a free and open source platform to create and host your podcasts.

## Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR

DevFeed: [Next generation medical image interpretation with MedGemma 1.5 and medical speech to text with MedASR](<https://devfeed.tech/articles/next-generation-medical-image-interpretation-with-medgemma-1-5-and-medical-speech-to-text-with-medasr-6842.md>)

Original publisher: [Read original article](<https://research.google/blog/next-generation-medical-image-interpretation-with-medgemma-15-and-medical-speech-to-text-with-medasr/>)

Published: 2026-01-13T20:57:16Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [Medical imaging](<https://devfeed.tech/topics/medical-imaging.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Generative AI](<https://devfeed.tech/topics/generative-ai.md>), [Google](<https://devfeed.tech/topics/google.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Cloud](<https://devfeed.tech/topics/cloud.md>), [Kaggle](<https://devfeed.tech/topics/kaggle.md>)

Tags: [ai-models](<https://devfeed.tech/tags/ai-models.md>), [asr](<https://devfeed.tech/tags/asr.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [google](<https://devfeed.tech/tags/google.md>), [google-cloud](<https://devfeed.tech/tags/google-cloud.md>), [health](<https://devfeed.tech/tags/health.md>), [health-bioscience](<https://devfeed.tech/tags/health-bioscience.md>), [healthcare](<https://devfeed.tech/tags/healthcare.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [kaggle](<https://devfeed.tech/tags/kaggle.md>), [machine-intelligence](<https://devfeed.tech/tags/machine-intelligence.md>), [medical-imaging](<https://devfeed.tech/tags/medical-imaging.md>)

### AI overview

Google Research describes MedGemma 1.5 4B, an updated open medical generative AI model with improved support for medical imaging, text, medical records, and 2D images. The article also presents MedASR, an open medical speech-to-text model for dictation that can pair with MedGemma for advanced reasoning. The models are available for research and commercial use through Hugging Face and Vertex AI, with a related medical AI hackathon on Kaggle.

### Source excerpt

Generative AI

## Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks

DevFeed: [Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks](<https://devfeed.tech/articles/open-asr-leaderboard-trends-and-insights-with-new-multilingual-long-form-tracks-7409.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard>)

Author: Eric Bezzam; Steven Zheng; Eustache Le Bihan; Vaibhav Srivastav

Published: 2025-11-21T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Language models](<https://devfeed.tech/topics/language-models.md>), [Fine-tuning](<https://devfeed.tech/topics/fine-tuning.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [fine-tuning](<https://devfeed.tech/tags/fine-tuning.md>), [ibm](<https://devfeed.tech/tags/ibm.md>), [inference](<https://devfeed.tech/tags/inference.md>), [large-language-model](<https://devfeed.tech/tags/large-language-model.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [microsoft](<https://devfeed.tech/tags/microsoft.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [nvidia](<https://devfeed.tech/tags/nvidia.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [qwen](<https://devfeed.tech/tags/qwen.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcription](<https://devfeed.tech/tags/transcription.md>), [trends](<https://devfeed.tech/tags/trends.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The Open ASR Leaderboard expands its evaluation with multilingual and long-form transcription tracks. The article highlights accuracy advantages from combining Conformer encoders with LLM decoders, throughput advantages from CTC and TDT decoders, Whisper as a multilingual baseline, and the effects of fine-tuning on specialized performance.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Real-time speech-to-speech translation

DevFeed: [Real-time speech-to-speech translation](<https://devfeed.tech/articles/real-time-speech-to-speech-translation-6852.md>)

Original publisher: [Read original article](<https://research.google/blog/real-time-speech-to-speech-translation/>)

Published: 2025-11-19T09:59:00Z

Content type: article

Language: en

Sources: [The latest research from Google](<https://devfeed.tech/sources/the-latest-research-from-google.md>)

Topics: [speech-to-speech](<https://devfeed.tech/topics/speech-to-speech.md>), [real-time](<https://devfeed.tech/topics/real-time.md>), [Machine learning](<https://devfeed.tech/topics/machine-learning.md>), [Streaming](<https://devfeed.tech/topics/streaming.md>), [data-processing](<https://devfeed.tech/topics/data-processing.md>), [asr](<https://devfeed.tech/topics/asr.md>), [Google](<https://devfeed.tech/topics/google.md>)

Tags: [algorithms-theory](<https://devfeed.tech/tags/algorithms-theory.md>), [asr](<https://devfeed.tech/tags/asr.md>), [data](<https://devfeed.tech/tags/data.md>), [google](<https://devfeed.tech/tags/google.md>), [ml](<https://devfeed.tech/tags/ml.md>), [personalization](<https://devfeed.tech/tags/personalization.md>), [real-time](<https://devfeed.tech/tags/real-time.md>), [speech](<https://devfeed.tech/tags/speech.md>), [speech-to-speech](<https://devfeed.tech/tags/speech-to-speech.md>), [streaming](<https://devfeed.tech/tags/streaming.md>)

### AI overview

Google researchers introduce an end-to-end speech-to-speech translation model that translates speech in real time while preserving the original speaker's voice, with a two-second delay. The system uses a streaming architecture, time-synchronized training data, and a scalable data-acquisition pipeline to support more languages and improve conversational naturalness.

### Source excerpt

Algorithms & Theory

## Voice Cloning with Consent

DevFeed: [Voice Cloning with Consent](<https://devfeed.tech/articles/voice-cloning-with-consent-7562.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/voice-consent-gate>)

Author: Margaret Mitchell; Lucie-Aimée Kaffee

Published: 2025-10-28T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [AI Development](<https://devfeed.tech/topics/ai-development.md>), [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [asr](<https://devfeed.tech/topics/asr.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [ethics](<https://devfeed.tech/tags/ethics.md>), [guide](<https://devfeed.tech/tags/guide.md>), [speech](<https://devfeed.tech/tags/speech.md>), [systems](<https://devfeed.tech/tags/systems.md>), [text-to-speech](<https://devfeed.tech/tags/text-to-speech.md>), [voice](<https://devfeed.tech/tags/voice.md>), [voice-cloning](<https://devfeed.tech/tags/voice-cloning.md>), [workflows](<https://devfeed.tech/tags/workflows.md>)

### AI overview

This article presents a voice consent gate for voice cloning. The system requires a speaker to generate and speak a consent phrase, uses automatic speech recognition to verify it, and only then permits a text-to-speech voice-cloning model to run. The design makes consent a traceable and auditable prerequisite for AI action while addressing both the risks and beneficial uses of voice cloning.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Swift Transformers Reaches 1.0 - and Looks to the Future

DevFeed: [Swift Transformers Reaches 1.0 - and Looks to the Future](<https://devfeed.tech/articles/swift-transformers-reaches-1-0-and-looks-to-the-future-7496.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/swift-transformers>)

Author: Pedro Cuenca; Christopher Fleetwood; Mattt; Vaibhav Srivastav

Published: 2025-09-26T00:00:00Z

Content type: release

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [swift-transformers](<https://devfeed.tech/topics/swift-transformers.md>), [MLX](<https://devfeed.tech/topics/mlx.md>), [Swift](<https://devfeed.tech/topics/swift.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Local AI](<https://devfeed.tech/topics/local-ai.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [apple](<https://devfeed.tech/tags/apple.md>), [asr](<https://devfeed.tech/tags/asr.md>), [coreml](<https://devfeed.tech/tags/coreml.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [llms](<https://devfeed.tech/tags/llms.md>), [local](<https://devfeed.tech/tags/local.md>), [mlx](<https://devfeed.tech/tags/mlx.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [swift](<https://devfeed.tech/tags/swift.md>), [swift-transformers](<https://devfeed.tech/tags/swift-transformers.md>), [transformers](<https://devfeed.tech/tags/transformers.md>)

### AI overview

swift-transformers reaches version 1.0 as a Swift library for running local models on Apple Silicon, including iPhones. It combines input preparation, Hugging Face Hub access, local caching and offline downloads, and wrappers for running Core ML-converted LLMs. The release establishes a stable foundation for Apple-focused local inference and future MLX and agentic use cases.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Mistral releases Voxtral open speech understanding models

DevFeed: [Mistral releases Voxtral open speech understanding models](<https://devfeed.tech/articles/voxtral-7138.md>)

Original publisher: [Read original article](<https://mistral.ai/news/voxtral/>)

Published: 2025-07-15T12:00:00Z

Content type: release

Language: en

Sources: [Mistral AI Blog](<https://devfeed.tech/sources/mistral-ai-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [release](<https://devfeed.tech/tags/release.md>), [transcription](<https://devfeed.tech/tags/transcription.md>)

### AI overview

Mistral announces Voxtral, a family of open speech understanding models available in 24B and 3B variants. The models support transcription, semantic understanding, audio Q&A and summarization, with deployment options for production, local, and edge use.

### Source excerpt

The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.

## Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints

DevFeed: [Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints](<https://devfeed.tech/articles/powerful-asr-diarization-speculative-decoding-with-hugging-face-inference-endpoints-7106.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/asr-diarization>)

Author: Sergei Petrov; Vaibhav Srivastav; Pedro Cuenca; Philipp Schmid

Published: 2024-05-01T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [inference-endpoints](<https://devfeed.tech/topics/inference-endpoints.md>), [hugging face](<https://devfeed.tech/topics/hugging-face.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [PyTorch](<https://devfeed.tech/topics/pytorch.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [flash-attention-2](<https://devfeed.tech/tags/flash-attention-2.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [inference-endpoints](<https://devfeed.tech/tags/inference-endpoints.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pytorch](<https://devfeed.tech/tags/pytorch.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

This article explains how to build a custom inference handler for Automatic Speech Recognition, speaker diarization, and speculative decoding on Hugging Face Inference Endpoints. It covers modular pipeline design, repository files, Pyannote-based diarization, PyTorch SDPA with Flash Attention 2, and constraints on speculative decoding such as batch size one and compatible decoder architectures.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.