# leaderboard

Published articles for leaderboard.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Measuring and Improving Consistency in Repeated Agent Runs

DevFeed: [Measuring and Improving Consistency in Repeated Agent Runs](<https://devfeed.tech/articles/your-agent-aced-the-task-will-it-do-it-again-26920.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ibm-research/altk-evolve-consistency>)

Author: Evelyn Duesterwald; Lilian Ngweta; Vatche Isahagian; Jayaram Radhakrishnan; Vinod Muthusamy; Gaodan Fang; Ashwath Vaithinathan Aravindan; Punleuk Oum; G Thomas; Merve Unuvar; Ayhan Sebin; Michał Ulewi

Published: 2026-09-15T16:00:44Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [consistency](<https://devfeed.tech/tags/consistency.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [reports](<https://devfeed.tech/tags/reports.md>), [standard](<https://devfeed.tech/tags/standard.md>), [technical](<https://devfeed.tech/tags/technical.md>)

### AI overview

This article presents the Consistency Analyzer, a diagnostic for finding decision points where an agent's behavior may change across repeated runs. It introduces consistency guidelines in ALTK-Evolve and reports that they reduced the consistency gap from 24.4 percentage points to 12.0 points without reducing average accuracy.

### Source excerpt

That is embarrassing onstage. In production, it is a reliability problem: a workflow that succeeded once may fail the next time a user makes the same request. For mission-critical work, such as reconciling a financial transaction or checking a contract for an obligation, that can be a showstopper. Most benchmarks hide this variability behind an average. On AppWorld, a ReAct agent using GPT-4.1 succeeded on 77.4% of runs across five repetitions.

## Claude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests.

DevFeed: [Claude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests.](<https://devfeed.tech/articles/claude-performed-best-on-a-new-benchmark-for-agents-that-build-agents-but-it-passed-fewer-than-a-quarter-of-the-tests-8472.md>)

Original publisher: [Read original article](<https://thenewstack.io/claude-build-agents-benchmark/>)

Author: Paul Sawers

Published: 2026-09-09T20:14:09Z

Content type: news

Language: en

Sources: [The New Stack](<https://devfeed.tech/sources/the-new-stack.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [coding](<https://devfeed.tech/topics/coding.md>), [anthropic](<https://devfeed.tech/topics/anthropic.md>), [Software](<https://devfeed.tech/topics/software.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [ai-engineering](<https://devfeed.tech/tags/ai-engineering.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude](<https://devfeed.tech/tags/claude.md>), [claude-code](<https://devfeed.tech/tags/claude-code.md>), [codex](<https://devfeed.tech/tags/codex.md>), [coding](<https://devfeed.tech/tags/coding.md>), [large-language-models](<https://devfeed.tech/tags/large-language-models.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [testing](<https://devfeed.tech/tags/testing.md>)

### AI overview

Hyper-𝜏-bench evaluates whether AI developer agents can build customer-service agents from simulated business materials. Claude Opus 5 in Claude Code led the six tested configurations at 23.9%, while none exceeded 25%.

### Source excerpt

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests. appeared first on The New Stack.

## Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)

DevFeed: [Agents on Rails: Claude Fable 5.1 and GLM 5.3 Flash (formerly known as ox-alpha)](<https://devfeed.tech/articles/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash-formerly-known-as-ox-alpha-3565.md>)

Original publisher: [Read original article](<https://rubyonrails.org/2026/9/2/agents-on-rails-claude-fable-5-1-and-glm-5-3-flash>)

Author: Svyatoslav Kryukov, Artur Petrov

Published: 2026-09-02T00:00:00Z

Content type: news

Language: en

Sources: [Ruby on Rails: Compress the complexity of modern web apps](<https://devfeed.tech/sources/ruby-on-rails-compress-the-complexity-of-modern-web-apps.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [apis](<https://devfeed.tech/tags/apis.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cost](<https://devfeed.tech/tags/cost.md>), [flash](<https://devfeed.tech/tags/flash.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [security](<https://devfeed.tech/tags/security.md>)

### AI overview

A benchmark update compares Claude Fable 5.1 with Claude Opus 5, Fable 5, and GLM 5.3 Flash on Rails tasks, accuracy, cost, speed, security-task handling, and Rails API recall.

### Source excerpt

Claude Fable 5.1 dropped yesterday. See how it handles real Rails tasks, and where it landed on the leaderboard today. We also finally have a name for the stealth model from the last round.

## The Open ASR Leaderboard Adds Its First Global South Language

DevFeed: [The Open ASR Leaderboard Adds Its First Global South Language](<https://devfeed.tech/articles/the-open-asr-leaderboard-adds-its-first-global-south-language-7411.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard-global-south>)

Author: Eric Bezzam; Shobhit Banga; Manas Dhir; Bhaskar Singh; Manmeet Kaur; Aaditya Pareek; Walecha; Sagar Jain; Hanuman Sidh; Vanshika Chhabra

Published: 2026-08-28T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [contributors](<https://devfeed.tech/tags/contributors.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [devices](<https://devfeed.tech/tags/devices.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [research](<https://devfeed.tech/tags/research.md>), [speech](<https://devfeed.tech/tags/speech.md>)

### AI overview

The Open ASR Leaderboard introduces Monsoon evaluation sets for Hindi in India, expanding coverage beyond European languages and testing how recognition performance varies across populations and conditions. The sets use public and private splits, speaker-disjoint data, detailed speaker attributes, and variation in geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and transcript validity.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Raising machine-checked security benchmarks to advance hash-based SNARKs through agentic collaboration

DevFeed: [Raising machine-checked security benchmarks to advance hash-based SNARKs through agentic collaboration](<https://devfeed.tech/articles/raising-machine-checked-security-benchmarks-to-advance-hash-based-snarks-through-agentic-collaboration-17233.md>)

Original publisher: [Read original article](<https://blog.ethereum.org/en/2026/08/20/better-codes-challenge>)

Author: Ethereum Foundation Formal Verification team

Published: 2026-08-20T00:00:00Z

Content type: article

Language: en

Sources: [Ethereum Foundation Blog](<https://devfeed.tech/sources/ethereum-foundation-blog.md>)

Topics: [Formal verification](<https://devfeed.tech/topics/formal-verification.md>), [Lean](<https://devfeed.tech/topics/lean.md>), [Ethereum](<https://devfeed.tech/topics/ethereum.md>), [AI research agents](<https://devfeed.tech/topics/ai-research-agents.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Library](<https://devfeed.tech/topics/library.md>), [Security](<https://devfeed.tech/topics/security.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [ethereum](<https://devfeed.tech/tags/ethereum.md>), [formal-verification](<https://devfeed.tech/tags/formal-verification.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [paper](<https://devfeed.tech/tags/paper.md>), [research](<https://devfeed.tech/tags/research.md>), [research-development](<https://devfeed.tech/tags/research-development.md>), [security](<https://devfeed.tech/tags/security.md>), [verification](<https://devfeed.tech/tags/verification.md>)

### AI overview

The Ethereum Foundation's Formal Verification team launched better.codes, an open autoresearch challenge focused on raising the machine-checked soundness bound of the Lean-formalized koalaIRS12 Reed-Solomon proximity problem toward a fixed 128-bit target. Submissions are checked by the Lean kernel and promoted proofs are shared publicly.

### Source excerpt

better.codes, an open autoresearch challenge built by the Ethereum Foundation Formal Verification team in collaboration with Yukon and zkSecurity, is now live. better.codes takes a self-contained problem from the Proximity Prize research, formalized in Lean, and puts its soundness bound on a public leaderboard that anyone can push forward....

## Agents on Rails: Grok 4.6, GLM 5.3, Gemini 3.7 Flash, and Opus 4.8

DevFeed: [Agents on Rails: Grok 4.6, GLM 5.3, Gemini 3.7 Flash, and Opus 4.8](<https://devfeed.tech/articles/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8-3558.md>)

Original publisher: [Read original article](<https://rubyonrails.org/2026/8/17/agents-on-rails-grok-4-6-glm-5-3-gemini-3-7-flash-and-opus-4-8>)

Author: Svyatoslav Kryukov, Artur Petrov

Published: 2026-08-17T00:00:00Z

Content type: news

Language: en

Sources: [Ruby on Rails: Compress the complexity of modern web apps](<https://devfeed.tech/sources/ruby-on-rails-compress-the-complexity-of-modern-web-apps.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Rails](<https://devfeed.tech/topics/rails.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [API](<https://devfeed.tech/topics/api.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [bug](<https://devfeed.tech/tags/bug.md>), [claude](<https://devfeed.tech/tags/claude.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [models](<https://devfeed.tech/tags/models.md>), [traces](<https://devfeed.tech/tags/traces.md>)

### AI overview

Agents on Rails reports benchmark results for four newly added models: Grok 4.6, GLM 5.3, Claude Opus 4.8, and Gemini 3.7 Flash. Grok leads the newcomers with 52 of 63 completed runs, while Claude Opus 5 remains the overall leader at 58 of 63. The update also includes refreshed insights, an updated leaderboard, costs and recall rates, and full traces from the first two rounds.

### Source excerpt

Last week we launched Agents on Rails and published the first benchmark report. The response was immediate: suggestions, questions, model requests, and more than a few "but have you tried..." messages. We love the enthusiasm! We want this benchmark to be useful to you, so for this run we added four new models, updated the insights, and uploaded the full traces of the first two rounds (every command, diff, and verdict).

## 10x more capacity for Laguna S 2.1 on AI Gateway

DevFeed: [10x more capacity for Laguna S 2.1 on AI Gateway](<https://devfeed.tech/articles/10x-more-capacity-for-laguna-s-2-1-on-ai-gateway-791.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/10x-more-capacity-for-laguna-s-2-1-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-31T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [api](<https://devfeed.tech/tags/api.md>), [coding](<https://devfeed.tech/tags/coding.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Laguna S 2.1 on AI Gateway now has 10x more capacity for both paid and free model versions. The article explains configuring the model through the AI SDK or a coding agent, and describes AI Gateway features including usage tracking, retries, failover, and a model-usage leaderboard.

### Source excerpt

Laguna S 2.1 from Poolside now has 10x more capacity on AI Gateway. The increase applies to the paid version, poolside/laguna-s-2.1, and the free version, poolside/laguna-s-2.1-free, so you can send far more requests, good for high-volume agentic coding and long-running tasks. To use Laguna S 2.1, set model to poolside/laguna-s-2.1-free or poolside/laguna-s-2.1 in the AI SDK: To run it in a coding agent, use vercel ai-gateway coding-agents setup to connect your agents to AI Gateway, then select poolside/laguna-s-2.1 or poolside/laguna-s-2.1-free in the agent's model configuration. See the coding agents guide. AI Gateway gives you one API to hundreds of models, with usage tracking, retries, failover, and higher-than-provider uptime built in. It reflects provider pricing with no markup and no platform fee, including on Bring Your Own Key requests. Try Laguna S 2.1 in the model playground. Read more

## Inkling Small from Thinking Machines is now available on AI Gateway

DevFeed: [Inkling Small from Thinking Machines is now available on AI Gateway](<https://devfeed.tech/articles/inkling-small-from-thinking-machines-is-now-available-on-ai-gateway-984.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/inkling-small-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-30T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agentic-coding](<https://devfeed.tech/tags/agentic-coding.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [audio](<https://devfeed.tech/tags/audio.md>), [cost](<https://devfeed.tech/tags/cost.md>), [images](<https://devfeed.tech/tags/images.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [tool](<https://devfeed.tech/tags/tool.md>)

### AI overview

Inkling Small is now available through AI Gateway. The release highlights multimodal reasoning, configurable thinking effort, visual inspection capabilities, coding and tool-use workflows, Zero Data Retention support, and AI SDK integration.

### Source excerpt

Inkling Small from Thinking Machines is now available on AI Gateway. Inkling Small reaches performance comparable to the larger Inkling model at about a quarter of the size, using much less compute per task. It is a broad generalist with native reasoning over audio and images, and it holds up well on reasoning, agentic coding, and tool use. Controllable thinking effort lets you trade quality against cost and latency, from minimal to maximum reasoning. For visual tasks, it can crop, zoom, and inspect images programmatically, which helps on documents and charts where the relevant detail is small. To use Inkling, set model to thinkingmachines/inkling-small in the AI SDK: Inkling-Small is compatible with Zero Data Retention. Turn it on team-wide from the dashboard, or per request with zeroDataRetention: true, and AI Gateway routes only to providers that delete prompts and responses after each request. Inkling-Small is also a cost-efficient choice for coding and tool-use workflows. Run vercel ai-gateway coding-agents setup to connect your coding agents to AI Gateway, then select thinkingmachines/inkling-small in the agent's model configuration. See the coding agents guide. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Inkling Small in the model playground. Read more

## Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway

DevFeed: [Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway](<https://devfeed.tech/articles/gemini-3-6-flash-and-gemini-3-5-flash-lite-are-now-available-on-ai-gateway-945.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/gemini-3-6-flash-3-5-flash-lite-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-21T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Web Development](<https://devfeed.tech/topics/web-development.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [api](<https://devfeed.tech/tags/api.md>), [api-keys](<https://devfeed.tech/tags/api-keys.md>), [coding](<https://devfeed.tech/tags/coding.md>), [cost](<https://devfeed.tech/tags/cost.md>), [flash](<https://devfeed.tech/tags/flash.md>), [gemini](<https://devfeed.tech/tags/gemini.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [routing](<https://devfeed.tech/tags/routing.md>), [web-development](<https://devfeed.tech/tags/web-development.md>)

### AI overview

Vercel AI Gateway now offers Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. The release highlights coding and agentic-task improvements, AI SDK model identifiers, and Gateway features for model access, usage tracking, routing, and reliability.

### Source excerpt

Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are now available on AI Gateway. Gemini 3.6 Flash improves quality across coding, agentic tasks, and web development while consuming fewer tokens and making fewer model calls. It produces cleaner web and app development output. Gemini 3.5 Flash Lite upgrades the agentic capabilities of the Flash-Lite tier, making it a good fit for subagents that handle scoped parts of a larger task. To use them, set model to google/gemini-3.6-flash or google/gemini-3.5-flash-lite in the AI SDK: AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Gemini 3.6 Flash in the model playground. Read more

## Introducing Real World VoiceEQ: Measuring the human quality of voice AI

DevFeed: [Introducing Real World VoiceEQ: Measuring the human quality of voice AI](<https://devfeed.tech/articles/introducing-real-world-voiceeq-measuring-the-human-quality-of-voice-ai-7454.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/real-world-voiceeq>)

Author: David Ayllon; Alice; Jeff Brooks; Franc Camps Febrer; Jakub Piotr Cłapa; Theo Lebryk; Jens Madsen; Olya Ossipova; Sharath Rao; Hoon Shin

Published: 2026-07-15T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [voice ai](<https://devfeed.tech/topics/voice-ai.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [asr](<https://devfeed.tech/topics/asr.md>), [human feedback](<https://devfeed.tech/topics/human-feedback.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [announcement](<https://devfeed.tech/tags/announcement.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [human-feedback](<https://devfeed.tech/tags/human-feedback.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reinforcement-learning](<https://devfeed.tech/tags/reinforcement-learning.md>), [speech](<https://devfeed.tech/tags/speech.md>), [voice-ai](<https://devfeed.tech/tags/voice-ai.md>)

### AI overview

Real World VoiceEQ is a benchmark for evaluating the human quality of voice AI beyond latency and word error rate. It measures how voice systems recognize, produce, and respond to acoustic information such as tone, emotion, speaker identity, and background context across ASR, TTS, speech-to-speech, and speech understanding. The benchmark covers more than 40 voice models, 15+ evaluation dimensions, and more than 60 metrics, using over 1 million human ratings.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Access and share AI Gateway leaderboard data

DevFeed: [Access and share AI Gateway leaderboard data](<https://devfeed.tech/articles/access-and-share-ai-gateway-leaderboard-data-1036.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/open-data-and-shareable-charts-for-ai-gateway-leaderboards>)

Author: Jerilyn Zheng

Published: 2026-07-14T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [API](<https://devfeed.tech/topics/api.md>), [CSV](<https://devfeed.tech/topics/csv.md>), [inference-providers](<https://devfeed.tech/topics/inference-providers.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [api](<https://devfeed.tech/tags/api.md>), [data](<https://devfeed.tech/tags/data.md>), [images](<https://devfeed.tech/tags/images.md>), [inference-providers](<https://devfeed.tech/tags/inference-providers.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [metrics](<https://devfeed.tech/tags/metrics.md>), [model](<https://devfeed.tech/tags/model.md>), [open](<https://devfeed.tech/tags/open.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [time](<https://devfeed.tech/tags/time.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [videos](<https://devfeed.tech/tags/videos.md>)

### AI overview

Vercel has opened the data behind the AI Gateway leaderboards under the CC BY 4.0 license. The data can be downloaded as CSV or queried through an export API, while charts can be shared as branded PNG images. The leaderboards track daily production usage across models, labs, apps, and inference providers.

### Source excerpt

We are making the data behind the AI Gateway leaderboards open under the CC BY 4.0 license. You can now download or query the data through the leaderboard-export API endpoint and render any chart as a shareable image. The AI Gateway leaderboards show how AI is used in production, ranking traffic for models, labs, apps, and providers. Data is aggregated daily across trillions of tokens, so you can see what gets adopted and how that changes over time. For deeper analysis, see the July AI Gateway Production Index. What's ranked There are four leaderboards, each with its own metrics: Leaderboard Ranks Metrics Models Individual models Requests, token volume, spend, images or videos generated Labs Model labs Requests, token volume, spend, images or videos generated Apps Opted-in apps built on the AI Gateway Token volume, spend Providers Inference providers Token volume, spend Models and labs can be filtered by modality (text, image, video) and show a daily percentage share over time; apps and providers are aggregated across all modalities and show a ranked top list. Open data The data behind the leaderboards is open, published under Creative Commons Attribution 4.0 (CC BY 4.0). You are free to use, share, and adapt it, including commercially, as long as you give credit, link to the license, and indicate if changes were made. Every chart and ranked list has a download button that exports the current view as a CSV. For programmatic access, use the export endpoint, which returns the same data and is cached for 24 hours: For models and labs, each row is one entity's daily share of a single metric. One response includes rows for requests, tokens, spend, imageCount, and videoCount, so filter on the metric field to pull out the series you want. Share a chart Every chart has a share button that turns the current view into an image. Pick an aspect ratio (landscape, square, or portrait), then download it as a PNG or copy it to your clipboard. The image includes the legend, title, a

## Muse Spark 1.1 is now available on AI Gateway

DevFeed: [Muse Spark 1.1 is now available on AI Gateway](<https://devfeed.tech/articles/muse-spark-1-1-is-now-available-on-ai-gateway-1020.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/muse-spark-1-1-is-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-09T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI Models](<https://devfeed.tech/topics/ai-models.md>), [multimodal](<https://devfeed.tech/topics/multimodal.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Model Context Protocol](<https://devfeed.tech/topics/model-context-protocol.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [API keys](<https://devfeed.tech/topics/api-keys.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [audio](<https://devfeed.tech/tags/audio.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [image](<https://devfeed.tech/tags/image.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [multimodal](<https://devfeed.tech/tags/multimodal.md>), [muse](<https://devfeed.tech/tags/muse.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [routing](<https://devfeed.tech/tags/routing.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [spark](<https://devfeed.tech/tags/spark.md>), [tokens](<https://devfeed.tech/tags/tokens.md>), [tool](<https://devfeed.tech/tags/tool.md>), [tools](<https://devfeed.tech/tags/tools.md>)

### AI overview

Muse Spark 1.1 from Meta is now available through Vercel's AI Gateway. It is a multimodal reasoning model with a 1M-token context window for agentic tasks, supporting multiple input types, tool orchestration, MCP servers, custom skills, parallel tool calls, structured output, and search with citations.

### Source excerpt

Muse Spark 1.1 from Meta is now available on AI Gateway. It is a multimodal reasoning model with a 1M token context window built for agentic tasks, accepting text, image, video, PDF, and audio inputs. Muse Spark 1.1 plans and orchestrates work across tools and services, operating as a main agent or as a subagent, and it works with new tools, MCP servers, and custom skills without examples. The model supports parallel tool calling, structured output, and built-in search with citations. To use Muse Spark 1.1, set model to meta/muse-spark-1.1 in the AI SDK: AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Muse Spark 1.1 in the model playground. Read more

## Don't bring exposed developer credentials to Black Hat

DevFeed: [Don't bring exposed developer credentials to Black Hat](<https://devfeed.tech/articles/don-t-bring-exposed-developer-credentials-to-black-hat-1914.md>)

Original publisher: [Read original article](<https://1password.com/blog/developer-credential-security-black-hat-2026>)

Author: info@1password.com (Eric Eddy)

Published: 2026-07-09T00:00:00Z

Content type: article

Language: en

Sources: [Blog on 1Password Blog](<https://devfeed.tech/sources/blog-on-1password-blog.md>)

Topics: [Cybersecurity](<https://devfeed.tech/topics/cybersecurity.md>), [Security](<https://devfeed.tech/topics/security.md>), [ssh](<https://devfeed.tech/topics/ssh.md>), [Encryption](<https://devfeed.tech/topics/encryption.md>), [passwords](<https://devfeed.tech/topics/passwords.md>), [Command-line interface](<https://devfeed.tech/topics/cli.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>), [mount](<https://devfeed.tech/topics/mount.md>), [Go](<https://devfeed.tech/topics/go.md>), [JavaScript](<https://devfeed.tech/topics/javascript.md>), [Python](<https://devfeed.tech/topics/python.md>), [make](<https://devfeed.tech/topics/make.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [attacks](<https://devfeed.tech/tags/attacks.md>), [black-hat](<https://devfeed.tech/tags/black-hat.md>), [cli](<https://devfeed.tech/tags/cli.md>), [cryptography](<https://devfeed.tech/tags/cryptography.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [developer](<https://devfeed.tech/tags/developer.md>), [developer-tools](<https://devfeed.tech/tags/developer-tools.md>), [developers](<https://devfeed.tech/tags/developers.md>), [events](<https://devfeed.tech/tags/events.md>), [go](<https://devfeed.tech/tags/go.md>), [integrations](<https://devfeed.tech/tags/integrations.md>), [javascript](<https://devfeed.tech/tags/javascript.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [passwords](<https://devfeed.tech/tags/passwords.md>), [python](<https://devfeed.tech/tags/python.md>), [secrets](<https://devfeed.tech/tags/secrets.md>), [security](<https://devfeed.tech/tags/security.md>), [ssh](<https://devfeed.tech/tags/ssh.md>), [supply-chain](<https://devfeed.tech/tags/supply-chain.md>)

### AI overview

This article urges developers attending Black Hat to secure locally stored credentials before the conference. It discusses exposed AWS credentials and SSH keys, plaintext secrets, outdated cryptography, and 1Password tools for discovering, securing, and accessing credentials. It also describes mounted environment variables, CLI and SDK integrations, developer demonstrations, and a credential-sprawl challenge.

### Source excerpt

Black Hat is where the security industry gathers to compare notes on current cybersecurity topics. It brings together a diverse group of security experts, from C-suite executives to black-hat hackers. Some attendees see it as a target-rich environment for testing their latest hacks. Many hackers and supply chain attacks rely on the fact that local credentials are stored in predictable locations with standardized file names, in clear text. For example, AWS credentials usually live in ~/.aws/credentials because the CLI writes them there by default. SSH keys live in ~/.ssh. 1Password developer tools can secure these credentials. It's never a bad time to secure locally-stored developer credentials, but if you're attending Black Hat, this might be an especially good time. Secure your credentials in 1Password before the conference, and find us at the booth to get an exclusive sticker. Find and secure SSH keys Developer watchtower discovers SSH keys that are stored in plaintext or use outdated cryptography. Follow the documentation to discover and secure your local SSH keys, which you can then access from the terminal using biometrics, just the same way you do for your passwords. Secure environment variables (Beta) 1Password Environments make your Environment's variables available via locally mounted .env files, without writing your credentials to disk. You can securely share them with team members and access them programmatically in your terminal via our CLI or via our SDK in Go, JavaScript, or Python integrations. Follow the documentation to secure and mount your environment variables. Find us at booth 4735 Located in the main exhibit hall near the Bayside C escalators. If you're attending the conference please come say hello, pick up some stickers, and ask us all your questions about 1Password developer tools! Play our developer challenge, Credential Sprawl Capture the Flag to win exclusive swag and get your name on our leaderboard. We'll be running live demos and havin

## Grok 4.5 now available on AI Gateway

DevFeed: [Grok 4.5 now available on AI Gateway](<https://devfeed.tech/articles/grok-4-5-now-available-on-ai-gateway-971.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/grok-4-5-now-available-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-07-08T00:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [API](<https://devfeed.tech/topics/api.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [api](<https://devfeed.tech/tags/api.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [ranking](<https://devfeed.tech/tags/ranking.md>), [routing](<https://devfeed.tech/tags/routing.md>)

### AI overview

Grok 4.5 from SpaceXAI is now available through Vercel AI Gateway. The model accepts text and image inputs, offers configurable reasoning levels, and can be selected through the AI SDK or routing rules. AI Gateway provides unified model access, usage and cost tracking, retries, failover, reporting, API key budgets, and performance optimizations without provider-price markup or platform inference fees.

### Source excerpt

Grok 4.5 from SpaceXAI is now available on AI Gateway. Built for coding, knowledge work, and STEM, the model accepts text and image inputs. Grok 4.5 supports low, medium, and high reasoning levels and defaults to high. Set the level with reasoning to balance speed against depth. To use Grok 4.5, set model to xai/grok-4.5 in the AI SDK: You can also set routing rules to switch to Grok 4.5 from other gateway models without touching your code. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, routing rules, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Try Grok 4.5 in the model playground. Read more

## ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

DevFeed: [ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration](<https://devfeed.tech/articles/scarfbench-benchmarking-ai-agents-for-enterprise-java-framework-migration-7269.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ibm-research/scarfbench>)

Author: Raju Pavuluri; Rahul Krishna; Srikanth Govindaraj Tamilselvam; Bridget M; Ashita Saxena; George Safta; Advait Pavuluri; Michele Merler

Published: 2026-06-30T18:32:50Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [migration](<https://devfeed.tech/topics/migration.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Software Engineering](<https://devfeed.tech/topics/software-engineering.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Jakarta EE](<https://devfeed.tech/topics/jakarta-ee.md>), [Quarkus](<https://devfeed.tech/topics/quarkus.md>), [Deployment](<https://devfeed.tech/topics/deployment.md>), [Refactoring](<https://devfeed.tech/topics/refactoring.md>), [configuration](<https://devfeed.tech/topics/configuration.md>), [agent observability](<https://devfeed.tech/topics/agent-observability.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-agents](<https://devfeed.tech/tags/ai-agents.md>), [applications](<https://devfeed.tech/tags/applications.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [coding](<https://devfeed.tech/tags/coding.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [engineering](<https://devfeed.tech/tags/engineering.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [java](<https://devfeed.tech/tags/java.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [migration](<https://devfeed.tech/tags/migration.md>), [modernization](<https://devfeed.tech/tags/modernization.md>), [open](<https://devfeed.tech/tags/open.md>), [software](<https://devfeed.tech/tags/software.md>), [software-engineering](<https://devfeed.tech/tags/software-engineering.md>), [spring](<https://devfeed.tech/tags/spring.md>), [validation](<https://devfeed.tech/tags/validation.md>)

### AI overview

ScarfBench is an open benchmark for evaluating AI agents on enterprise Java framework migrations across Spring, Jakarta EE, and Quarkus. It measures whether migrated applications build, deploy, and preserve behavior, and reports that current agents achieve less than 10% behavioral success on the benchmark.

### Source excerpt

Recent advances in coding agents have sparked excitement around AI-assisted modernization. But an important question remains: Can AI agents reliably modernize real-world enterprise applications? Existing software engineering benchmarks have demonstrated impressive progress in bug fixing and code generation, but framework migration presents a fundamentally different challenge.

## Claude Sonnet 5 now available on Vercel AI Gateway

DevFeed: [Claude Sonnet 5 now available on Vercel AI Gateway](<https://devfeed.tech/articles/claude-sonnet-5-now-available-on-vercel-ai-gateway-869.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/claude-sonnet-5-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-06-30T00:00:00Z

Content type: news

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI, ML & Data Engineering](<https://devfeed.tech/topics/ai-ml-data-engineering.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [SDKs](<https://devfeed.tech/topics/sdks.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [ai-models](<https://devfeed.tech/tags/ai-models.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [anthropic-claude](<https://devfeed.tech/tags/anthropic-claude.md>), [api](<https://devfeed.tech/tags/api.md>), [api-keys](<https://devfeed.tech/tags/api-keys.md>), [claude](<https://devfeed.tech/tags/claude.md>), [coding](<https://devfeed.tech/tags/coding.md>), [cost](<https://devfeed.tech/tags/cost.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [memory](<https://devfeed.tech/tags/memory.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [performance](<https://devfeed.tech/tags/performance.md>), [playground](<https://devfeed.tech/tags/playground.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Claude Sonnet 5 is available through Vercel AI Gateway, with claimed improvements in coding, agentic work, document parsing, and long-context memory. The announcement lists launch and standard pricing and describes Gateway features for model access, usage tracking, cost controls, retries, and failover.

### Source excerpt

Claude Sonnet 5 from Anthropic is now available on AI Gateway. Sonnet 5 improves on Sonnet 4.6 across coding and agentic work, reaching outcomes on many tasks that previously needed an Opus model, at Sonnet pricing. The model is more agentic and follows instructions more closely. Document parsing and long-context memory use are also stronger. Sonnet 5 also uses an updated tokenizer, like the recent Opus models, which can map the same input to more tokens. Launch pricing of $2 per million input tokens and $10 per million output tokens runs through August 31, 2026. Standard list price will be $3/M input tokens, $15/M output tokens. To use Sonnet 5, set model to anthropic/claude-sonnet-5 in the AI SDK: You can also try Sonnet 5 in the model playground. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, budgets for API keys, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Read more

## Featuring Every Eval Ever Results on Hugging Face Model Pages

DevFeed: [Featuring Every Eval Ever Results on Hugging Face Model Pages](<https://devfeed.tech/articles/featuring-every-eval-ever-results-on-hugging-face-model-pages-7179.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/eee-community-evals>)

Author: Sree Harsha Nelaturu; Avijit Ghosh; Nathan Habib; Jan Batzner; Leshem Choshen; Irene Solaiman; Julien Chaumond

Published: 2026-06-30T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [data](<https://devfeed.tech/tags/data.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [integration](<https://devfeed.tech/tags/integration.md>), [json](<https://devfeed.tech/tags/json.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

The article presents EEE, a JSON schema and datastore for standardizing AI evaluation results, and its integration with Hugging Face Community Evals through a converter that produces YAML files. It aims to make model benchmark results more comparable by recording evaluation context and preserving results from varied reporting formats.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World

DevFeed: [Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World](<https://devfeed.tech/articles/introducing-the-ffasr-leaderboard-benchmarking-asr-in-the-real-world-7196.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/ffasr-leaderboard>)

Author: Daniel Gert Nielsen; Shivam Saini; Alessia Milo; Georg Götz; Eric Bezzam

Published: 2026-06-24T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [announcement](<https://devfeed.tech/tags/announcement.md>), [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [community](<https://devfeed.tech/tags/community.md>), [deployment](<https://devfeed.tech/tags/deployment.md>), [development](<https://devfeed.tech/tags/development.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [performance](<https://devfeed.tech/tags/performance.md>), [speech](<https://devfeed.tech/tags/speech.md>)

### AI overview

Treble Technologies and Hugging Face introduce the open, community-driven FFASR Leaderboard, a benchmark for evaluating automatic speech recognition models in realistic far-field acoustic conditions. It measures performance across factors such as reverberation, background noise, microphone distance, and low signal-to-noise ratios, while also showing the tradeoff between recognition accuracy and speed.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Gamification 2.0. Beyond Points and Badges: Designing for Players, Not Metrics. Chapter 5: Implementation

DevFeed: [Gamification 2.0. Beyond Points and Badges: Designing for Players, Not Metrics. Chapter 5: Implementation](<https://devfeed.tech/articles/gamification-2-0-beyond-points-and-badges-designing-for-players-not-metrics-chapter-5-implementation-9080.md>)

Original publisher: [Read original article](<https://uxmag.com/articles/gamification-2-0-beyond-points-and-badges-designing-for-players-not-metrics-chapter-5-implementation>)

Author: Montgomery Singman

Published: 2026-06-09T03:26:13Z

Content type: tutorial

Language: en

Sources: [UX Magazine](<https://devfeed.tech/sources/ux-magazine.md>)

Topics: [implementation](<https://devfeed.tech/topics/implementation.md>), [User experience (UX)](<https://devfeed.tech/topics/ux.md>), [App](<https://devfeed.tech/topics/app.md>)

Tags: [app](<https://devfeed.tech/tags/app.md>), [building](<https://devfeed.tech/tags/building.md>), [complex-systems](<https://devfeed.tech/tags/complex-systems.md>), [core](<https://devfeed.tech/tags/core.md>), [creative](<https://devfeed.tech/tags/creative.md>), [design](<https://devfeed.tech/tags/design.md>), [expression](<https://devfeed.tech/tags/expression.md>), [games](<https://devfeed.tech/tags/games.md>), [implementation](<https://devfeed.tech/tags/implementation.md>), [interface](<https://devfeed.tech/tags/interface.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [loops](<https://devfeed.tech/tags/loops.md>), [management](<https://devfeed.tech/tags/management.md>), [multiplayer](<https://devfeed.tech/tags/multiplayer.md>), [puzzle](<https://devfeed.tech/tags/puzzle.md>), [rpg](<https://devfeed.tech/tags/rpg.md>), [sandbox](<https://devfeed.tech/tags/sandbox.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [strategy](<https://devfeed.tech/tags/strategy.md>), [ux](<https://devfeed.tech/tags/ux.md>)

### AI overview

This chapter presents a practical framework for implementing Gamification 2.0. It recommends choosing a dominant game genre that matches an app's core activities, aligning that genre with user psychology, and designing satisfying intrinsic interaction loops before adding points, badges, or other extrinsic rewards.

### Source excerpt

Part 5 of the "Gamification Series." A framework for developers: from theory to practice Everything I've outlined so far is meaningless if you can't apply it. So let me give you a practical framework for actually implementing Gamification 2.0. Step 1: Stop copying mechanics; choose a genre Your first question isn't "What gamification mechanics should The post Gamification 2.0. Beyond Points and Badges: Designing for Players, Not Metrics. Chapter 5: Implementation appeared first on UX Magazine.

## Opus 4.8 on AI Gateway

DevFeed: [Opus 4.8 on AI Gateway](<https://devfeed.tech/articles/opus-4-8-on-ai-gateway-1039.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/opus-4-8-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-05-28T07:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Claude](<https://devfeed.tech/topics/claude.md>), [Anthropic Claude](<https://devfeed.tech/topics/anthropic-claude.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [API](<https://devfeed.tech/topics/api.md>), [coding](<https://devfeed.tech/topics/coding.md>), [Latency](<https://devfeed.tech/topics/latency.md>), [model-deployment](<https://devfeed.tech/topics/model-deployment.md>)

Tags: [agentic](<https://devfeed.tech/tags/agentic.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [anthropic-claude](<https://devfeed.tech/tags/anthropic-claude.md>), [api](<https://devfeed.tech/tags/api.md>), [claude](<https://devfeed.tech/tags/claude.md>), [coding](<https://devfeed.tech/tags/coding.md>), [cost](<https://devfeed.tech/tags/cost.md>), [latency](<https://devfeed.tech/tags/latency.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [performance](<https://devfeed.tech/tags/performance.md>), [playground](<https://devfeed.tech/tags/playground.md>), [pricing](<https://devfeed.tech/tags/pricing.md>), [retention](<https://devfeed.tech/tags/retention.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [uptime](<https://devfeed.tech/tags/uptime.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Claude Opus 4.8 is now available through Vercel AI Gateway and the AI SDK. It is designed for long-horizon agentic execution, complex multi-step coding tasks, and knowledge work. AI Gateway offers unified model access, usage and cost tracking, retries, failover, performance optimization, reporting, Zero Data Retention support, and provider selection based on latency and cost without adding inference fees.

### Source excerpt

Claude Opus 4.8 is now available on Vercel AI Gateway. Claude Opus 4.8 is built for long-horizon agentic execution and handles complex, multi-step coding tasks like refactors that previously required human correction mid-task. The model also produces clearer, less hedgy prose for knowledge work like drafting documents, analyzing data, and building presentations. To use Opus 4.8, set model to anthropic/claude-opus-4.8 in the AI SDK. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, Zero Data Retention support, dynamic provider sorting by latency & cost, and more. AI Gateway reflects provider pricing with no markup and does not charge a platform fee on inference, including on Bring Your Own Key (BYOK) requests. Learn more about AI Gateway, view the AI Gateway model leaderboard or try it in our model playground. Read more

## What Parameter Golf taught us about AI-assisted research

DevFeed: [What Parameter Golf taught us about AI-assisted research](<https://devfeed.tech/articles/what-parameter-golf-taught-us-about-ai-assisted-research-6717.md>)

Original publisher: [Read original article](<https://openai.com/index/what-parameter-golf-taught-us>)

Published: 2026-05-12T00:00:00Z

Content type: article

Language: en

Sources: [OpenAI News](<https://devfeed.tech/sources/openai-news.md>)

Topics: [machine learning research](<https://devfeed.tech/topics/machine-learning-research.md>), [AI-assisted coding](<https://devfeed.tech/topics/ai-assisted-coding.md>), [Optimization](<https://devfeed.tech/topics/optimization.md>), [quantization](<https://devfeed.tech/topics/quantization.md>), [Compression](<https://devfeed.tech/topics/compression.md>), [Code](<https://devfeed.tech/topics/code.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [GitHub](<https://devfeed.tech/topics/github.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-coding](<https://devfeed.tech/tags/ai-coding.md>), [code](<https://devfeed.tech/tags/code.md>), [coding](<https://devfeed.tech/tags/coding.md>), [compression](<https://devfeed.tech/tags/compression.md>), [data](<https://devfeed.tech/tags/data.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [experimentation](<https://devfeed.tech/tags/experimentation.md>), [github](<https://devfeed.tech/tags/github.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [machine-learning-research](<https://devfeed.tech/tags/machine-learning-research.md>), [open](<https://devfeed.tech/tags/open.md>), [optimization](<https://devfeed.tech/tags/optimization.md>), [quantization](<https://devfeed.tech/tags/quantization.md>), [research](<https://devfeed.tech/tags/research.md>), [training](<https://devfeed.tech/tags/training.md>)

### AI overview

Parameter Golf was a machine learning research challenge with strict limits on artifact size, training time, and held-out loss. The article examines lessons from more than 2,000 submissions, including optimizer tuning, quantization, evaluation strategies, new modeling ideas, and the growing use of AI coding agents.

### Source excerpt

Parameter Golf brought together 1,000+ participants and 2,000+ submissions to explore AI-assisted machine learning research, coding agents, quantization, and novel model design under strict constraints.

## Adding Benchmaxxer Repellant to the Open ASR Leaderboard

DevFeed: [Adding Benchmaxxer Repellant to the Open ASR Leaderboard](<https://devfeed.tech/articles/adding-benchmaxxer-repellant-to-the-open-asr-leaderboard-7413.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/open-asr-leaderboard-private-data>)

Author: Eric Bezzam; Steven Zheng; Eustache Le Bihan; Sergio Bruccoleri; Jeanine Sinanan-Singh; Casey Ford; Guanbo Wang; Yukai Huang; Ke Li; Yufeng Hao

Published: 2026-05-06T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [asr](<https://devfeed.tech/topics/asr.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [benchmarking](<https://devfeed.tech/topics/benchmarking.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Whisper](<https://devfeed.tech/topics/whisper.md>)

Tags: [asr](<https://devfeed.tech/tags/asr.md>), [audio](<https://devfeed.tech/tags/audio.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [partnerships](<https://devfeed.tech/tags/partnerships.md>), [speech](<https://devfeed.tech/tags/speech.md>), [transcripts](<https://devfeed.tech/tags/transcripts.md>), [whisper](<https://devfeed.tech/tags/whisper.md>)

### AI overview

The article announces private English ASR datasets from Appen and DataoceanAI for the Open ASR Leaderboard. Keeping the datasets private is intended to reduce benchmark-specific optimization and test-set contamination while preserving a high-quality evaluation of multiple speech-recognition tasks.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Grok 4.3 on AI Gateway

DevFeed: [Grok 4.3 on AI Gateway](<https://devfeed.tech/articles/grok-4-3-on-ai-gateway-970.md>)

Original publisher: [Read original article](<https://vercel.com/changelog/grok-4-3-on-ai-gateway>)

Author: Jerilyn Zheng

Published: 2026-04-30T07:00:00Z

Content type: release

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [vercel ai sdk](<https://devfeed.tech/topics/vercel-ai-sdk.md>), [API](<https://devfeed.tech/topics/api.md>), [context window](<https://devfeed.tech/topics/context-window.md>), [long-context](<https://devfeed.tech/topics/long-context.md>), [Routing (disambiguation)](<https://devfeed.tech/topics/routing.md>), [Vercel](<https://devfeed.tech/topics/vercel.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [api](<https://devfeed.tech/tags/api.md>), [context-window](<https://devfeed.tech/tags/context-window.md>), [cost](<https://devfeed.tech/tags/cost.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [model](<https://devfeed.tech/tags/model.md>), [models](<https://devfeed.tech/tags/models.md>), [observability](<https://devfeed.tech/tags/observability.md>), [performance](<https://devfeed.tech/tags/performance.md>), [playground](<https://devfeed.tech/tags/playground.md>), [routing](<https://devfeed.tech/tags/routing.md>), [sdk](<https://devfeed.tech/tags/sdk.md>), [tool](<https://devfeed.tech/tags/tool.md>), [uptime](<https://devfeed.tech/tags/uptime.md>), [vercel](<https://devfeed.tech/tags/vercel.md>)

### AI overview

Grok 4.3 is now available through Vercel AI Gateway, offering a 1M-token context window and improvements in accuracy, tool calling, and instruction following. The article explains that AI Gateway provides a unified API, usage and cost tracking, retries, failover, provider routing, observability, custom reporting, and Bring Your Own Key support.

### Source excerpt

Grok 4.3 is now available on Vercel AI Gateway. The model has a 1M token context window and improvements in accuracy, tool calling, and instruction following. To use Grok 4.3, set model to xai/grok-4.3 in the AI SDK. AI Gateway provides a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime. It includes built-in custom reporting, observability, Bring Your Own Key support, and intelligent provider routing with automatic retries. Learn more about AI Gateway, view the AI Gateway model leaderboard or try it in our model playground. Read more

## QIMMA قِمّة ⛰: A Quality-First Arabic LLM Leaderboard

DevFeed: [QIMMA قِمّة ⛰: A Quality-First Arabic LLM Leaderboard](<https://devfeed.tech/articles/qimma-a-quality-first-arabic-llm-leaderboard-7511.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/tiiuae/qimma-arabic-leaderboard>)

Author: Leen AlQadi; Ahmed Alzubaidi; Mohammed Alyafeai; Maitha Alhammadi; Shaikha Alsuwaidi; Omar saif alkaabi; Basma Boussaha; Hakim Hacid

Published: 2026-04-21T10:09:58Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Natural language processing](<https://devfeed.tech/topics/nlp.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [Inference](<https://devfeed.tech/topics/inference.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [Code](<https://devfeed.tech/topics/code.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [blog](<https://devfeed.tech/tags/blog.md>), [code](<https://devfeed.tech/tags/code.md>), [data](<https://devfeed.tech/tags/data.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [llm](<https://devfeed.tech/tags/llm.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>)

### AI overview

QIMMA is a quality-first Arabic LLM leaderboard that validates benchmark data before evaluating models. It addresses translation issues, annotation errors, encoding problems, cultural bias, reproducibility gaps, and fragmented task coverage. The platform combines native Arabic content, systematic validation, code evaluation, and public per-sample inference outputs across 109 subsets from 14 benchmarks and more than 52,000 samples.

### Source excerpt

A Blog post by Technology Innovation Institute on Hugging Face

[Next page](<https://devfeed.tech/tags/leaderboard.md?cursor=WyIyMDI2LTA0LTIxVDEwOjA5OjU4KzAwOjAwIiwgImVjOTIyMTczLTQxZWMtNGEyZS1iZmEzLTJlMDI5OTcxYTdhZCJd>)