# leaderboards

Published articles for leaderboards.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Benchmaxxing: When the Benchmark Becomes the Target

DevFeed: [Benchmaxxing: When the Benchmark Becomes the Target](<https://devfeed.tech/articles/benchmaxxing-when-the-benchmark-becomes-the-target-8302.md>)

Original publisher: [Read original article](<https://www.crowdstrike.com/en-us/blog/benchmaxxing-when-benchmark-becomes-the-target/>)

Author: Nathan Danneman

Published: 2026-09-12T11:17:51.295154Z

Content type: article

Language: en

Sources: [Blog](<https://devfeed.tech/sources/blog.md>)

Topics: [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Cybersecurity](<https://devfeed.tech/topics/cybersecurity.md>), [Ground truth / benchmark quality](<https://devfeed.tech/topics/ground-truth-benchmark-quality.md>), [Detection engineering](<https://devfeed.tech/topics/detection-engineering.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-and-cybersecurity](<https://devfeed.tech/tags/ai-and-cybersecurity.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [benchmarks](<https://devfeed.tech/tags/benchmarks.md>), [blog](<https://devfeed.tech/tags/blog.md>), [cybersecurity](<https://devfeed.tech/tags/cybersecurity.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [securing-ai](<https://devfeed.tech/tags/securing-ai.md>), [security](<https://devfeed.tech/tags/security.md>), [validation](<https://devfeed.tech/tags/validation.md>)

### AI overview

The article explains how public AI and cybersecurity benchmarks can become targets for optimization, a practice it calls "benchmaxxing." It argues that gaming, ceiling effects, data leakage, binary scoring, omitted costs, and aggregate scores can make benchmark results poor proxies for real-world defensive capability. The article proposes task-coupled internal benchmarks intended to evaluate end-to-end cyber agents and support rigorous science rather than visibility-driven score optimization.

### Source excerpt

The more attention a benchmark receives, the stronger the incentive to optimize for it. In AI and cybersecurity, this can have significant consequences.

## Squingle Arcade Is Available Now For Free On PC VR

DevFeed: [Squingle Arcade Is Available Now For Free On PC VR](<https://devfeed.tech/articles/squingle-arcade-is-available-now-for-free-on-pc-vr-17308.md>)

Original publisher: [Read original article](<https://www.uploadvr.com/squingle-arcade-available-now-on-pcvr/>)

Author: Mike Johnson

Published: 2026-09-10T00:14:39Z

Content type: news

Language: en

Sources: [UploadVR](<https://devfeed.tech/sources/uploadvr.md>)

Topics: [arcade](<https://devfeed.tech/topics/arcade.md>), [Mazes](<https://devfeed.tech/topics/maze.md>), [pc](<https://devfeed.tech/topics/pc.md>)

Tags: [free](<https://devfeed.tech/tags/free.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [multiplayer](<https://devfeed.tech/tags/multiplayer.md>), [orbs](<https://devfeed.tech/tags/orbs.md>), [pc](<https://devfeed.tech/tags/pc.md>), [puzzle](<https://devfeed.tech/tags/puzzle.md>), [vr-gaming](<https://devfeed.tech/tags/vr-gaming.md>)

### AI overview

Squingle Arcade, a free-to-play quickplay version of the fluid-maze puzzle game Squingle, is now available on SteamVR and Quest. It includes more than 60 levels, asynchronous multiplayer, ghost racing, leaderboards, and scalable mazes.

### Source excerpt

Squingle Arcade, the follow up to the original fluid-maze-navigating puzzle game Squingle, is out now for PC VR.

## DeepSeek overtakes Google on volume, cost per token falls 13.6%

DevFeed: [DeepSeek overtakes Google on volume, cost per token falls 13.6%](<https://devfeed.tech/articles/deepseek-overtakes-google-on-volume-cost-per-token-falls-13-6-731.md>)

Original publisher: [Read original article](<https://vercel.com/blog/deepseek-overtakes-google-on-volume-cost-per-token-falls>)

Author: Harpreet Arora

Published: 2026-08-11T04:00:00Z

Content type: news

Language: en

Sources: [Vercel News](<https://devfeed.tech/sources/vercel-news.md>)

Topics: [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [ai-gateway](<https://devfeed.tech/tags/ai-gateway.md>), [anthropic](<https://devfeed.tech/tags/anthropic.md>), [applications](<https://devfeed.tech/tags/applications.md>), [claude](<https://devfeed.tech/tags/claude.md>), [cost](<https://devfeed.tech/tags/cost.md>), [data](<https://devfeed.tech/tags/data.md>), [deepseek](<https://devfeed.tech/tags/deepseek.md>), [enterprise](<https://devfeed.tech/tags/enterprise.md>), [google](<https://devfeed.tech/tags/google.md>), [gpt](<https://devfeed.tech/tags/gpt.md>), [launch](<https://devfeed.tech/tags/launch.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [media](<https://devfeed.tech/tags/media.md>), [models](<https://devfeed.tech/tags/models.md>), [openai](<https://devfeed.tech/tags/openai.md>), [production](<https://devfeed.tech/tags/production.md>), [video](<https://devfeed.tech/tags/video.md>), [work](<https://devfeed.tech/tags/work.md>)

### AI overview

The August AI Gateway Production Index reports that token prices fell 13.6% in July while token use and spending grew. DeepSeek surpassed Google for second place by token volume, while Kimi K3 and other open-weight models gained usage and spending.

### Source excerpt

AI Gateway Production Index -- August 2026 Every month, AI Gateway routes tens of trillions of tokens between production applications and AI labs. That traffic gives us a view of what AI usage actually looks like in today's enterprise, and we publish it here monthly. See the Production Index reports from May, June, and July. August 2026 Summary The August index reports on AI Gateway data collected through July 2026. The average price paid per token fell 13.6% in July, after rising almost 20% in May and holding steady in June. Token consumption grew so quickly that even with 37% growth in spend, cost per token saw a double-digit drop. DeepSeek became the second-largest lab by token volume, now running more than twice Google's volume. Anthropic collected 65% of gateway spending on 30% of token volume, at 4.4 times the average price of every other lab's tokens. Both media leaderboards changed hands. Google's Nano Banana took the lead in image volume from OpenAI's GPT Image, and ByteDance's Seedance led video in both volume and dollars. Kimi K3 launched into agent work Moonshot released Kimi K3 on July 16. Like Z.ai's GLM 5.2 released in June, it is built for long-horizon agent work, so its usage was heavy from the start at about twelve times the tokens per request of its predecessor, K2.5. It scaled quickly. K3's daily volume tripled between launch week and the final week of July, and by month end its requests were as heavy as Claude Opus 4.8's. On the last full day of July it ranked eighth on the gateway by token volume. The demand was new rather than diverted. K3 processed nearly two-thirds of all Kimi tokens within two weeks and 82% by the final week, while the rest of the family's volume fell only slightly. Open weight's share of gateway spend more than doubled in July to 8.6%. More than 90% of that growth is Moonshot and Z.ai. Moonshot's share of total gateway spend quadrupled, to 2.3%. Cheap open-weight models have been taking volume for months without taking reve

## Building Internal Flashboards in Slack with AI

DevFeed: [Building Internal Flashboards in Slack with AI](<https://devfeed.tech/articles/building-internal-flashboards-in-slack-with-ai-16051.md>)

Original publisher: [Read original article](<https://workos.com/blog/reporting-tool-with-no-editor>)

Author: WorkOS

Published: 2026-08-10T16:15:41Z

Content type: article

Language: en

Sources: [WorkOS Blog](<https://devfeed.tech/sources/workos-blog.md>)

Topics: [Slack](<https://devfeed.tech/topics/slack.md>), [dashboards](<https://devfeed.tech/topics/dashboards.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [Claude](<https://devfeed.tech/topics/claude.md>), [data](<https://devfeed.tech/topics/data.md>), [API](<https://devfeed.tech/topics/api.md>), [HTML](<https://devfeed.tech/topics/html.md>)

Tags: [agent](<https://devfeed.tech/tags/agent.md>), [ai](<https://devfeed.tech/tags/ai.md>), [api](<https://devfeed.tech/tags/api.md>), [claude](<https://devfeed.tech/tags/claude.md>), [code-generation](<https://devfeed.tech/tags/code-generation.md>), [html](<https://devfeed.tech/tags/html.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [reporting](<https://devfeed.tech/tags/reporting.md>), [slack](<https://devfeed.tech/tags/slack.md>)

### AI overview

WorkOS describes Flashboards, an internal reporting tool built as documents with live data connections and edited through AI agents in Slack. The article explains its HTML-based architecture, read-only data access, employee authentication, natural-language creation workflow, interactive filters, and versioned iteration.

### Source excerpt

Flashboards is WorkOS's internal reporting tool: a doc with a live data connection, and agents as the only editing interface. Zero to 55 weekly readers in seven weeks.

## Community Evals: Because we're done trusting black-box leaderboards over the community

DevFeed: [Community Evals: Because we're done trusting black-box leaderboards over the community](<https://devfeed.tech/articles/community-evals-because-we-re-done-trusting-black-box-leaderboards-over-the-community-7146.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/community-evals>)

Author: ben burtenshaw; Nathan Habib; Bertrand Chevrier; merve; Daniel van Strien; Niels Rogge; Julien Chaumond

Published: 2026-02-04T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>)

Tags: [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [git](<https://devfeed.tech/tags/git.md>), [hub](<https://devfeed.tech/tags/hub.md>), [hugging-face](<https://devfeed.tech/tags/hugging-face.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [models](<https://devfeed.tech/tags/models.md>), [open](<https://devfeed.tech/tags/open.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [pull-requests](<https://devfeed.tech/tags/pull-requests.md>)

### AI overview

Hugging Face introduces community-reported benchmark evaluations and dataset-hosted leaderboards, with reproducible specifications, pull-request submissions, and verification badges.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Gaia2 and ARE: Empowering the community to study agents

DevFeed: [Gaia2 and ARE: Empowering the community to study agents](<https://devfeed.tech/articles/gaia2-and-are-empowering-the-community-to-study-agents-7209.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/gaia2>)

Author: Clémentine Fourrier; Grégoire Mialon; Maxime Lecanu; Pierre Andrews; Adrien Carreira; frere thibaud; Avijit Ghosh; Romain Froger; Dheeraj Mekala; Caroline Pascal

Published: 2025-09-22T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [gaia](<https://devfeed.tech/topics/gaia.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Framework](<https://devfeed.tech/topics/framework.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>), [debugging](<https://devfeed.tech/topics/debugging.md>)

Tags: [agents](<https://devfeed.tech/tags/agents.md>), [apis](<https://devfeed.tech/tags/apis.md>), [artificial-intelligence](<https://devfeed.tech/tags/artificial-intelligence.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [complexity](<https://devfeed.tech/tags/complexity.md>), [dataset](<https://devfeed.tech/tags/dataset.md>), [debug](<https://devfeed.tech/tags/debug.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gaia](<https://devfeed.tech/tags/gaia.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [simulation](<https://devfeed.tech/tags/simulation.md>), [time](<https://devfeed.tech/tags/time.md>)

### AI overview

The article introduces Gaia2, a harder follow-up to the GAIA benchmark for evaluating interactive AI agents. Gaia2 expands evaluation from read-only information retrieval to read-and-write tasks involving tool use, web browsing, ambiguous and time-sensitive instructions, controlled failures, adaptability, and agent-to-agent collaboration. It is released with the open Meta Agents Research Environments (ARE) framework, which supports running, debugging, and evaluating agents in customizable simulated real-world conditions.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More

DevFeed: [Arabic Leaderboards: Introducing Arabic Instruction Following, Updating AraGen, and More](<https://devfeed.tech/articles/arabic-leaderboards-introducing-arabic-instruction-following-updating-aragen-and-more-7309.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/leaderboard-3c3h-aragen-ifeval>)

Author: Ali El Filali; Sarah AlBarri; Abouelseoud; samta kamboj; Neha Sengupta; Preslav Nakov

Published: 2025-04-08T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [AI Models](<https://devfeed.tech/topics/ai-models.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [arabic](<https://devfeed.tech/tags/arabic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [community](<https://devfeed.tech/tags/community.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [llm](<https://devfeed.tech/tags/llm.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [reasoning](<https://devfeed.tech/tags/reasoning.md>), [research](<https://devfeed.tech/tags/research.md>), [safety](<https://devfeed.tech/tags/safety.md>)

### AI overview

This article announces Arabic-Leaderboards, a unified platform for Arabic AI evaluations, and introduces Arabic Instruction Following powered by the Arabic IFEval Benchmark. It also presents an updated AraGen release with a larger 340-pair evaluation dataset and publicly releases the AraGen-12-24 benchmark and model responses evaluated by Claude-3.5-Sonnet under the 3C3H guidelines.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Building real-time leaderboards with Tinybird

DevFeed: [Building real-time leaderboards with Tinybird](<https://devfeed.tech/articles/building-real-time-leaderboards-with-tinybird-18410.md>)

Original publisher: [Read original article](<https://www.tinybird.co/blog/building-real-time-leaderboards-with-tinybird>)

Author: Jim Moffitt

Published: 2024-06-06T00:00:00Z

Content type: article

Language: en

Sources: [Tinybird](<https://devfeed.tech/sources/tinybird.md>)

Topics: [real-time](<https://devfeed.tech/topics/real-time.md>), [data](<https://devfeed.tech/topics/data.md>)

Tags: [analytics](<https://devfeed.tech/tags/analytics.md>), [building](<https://devfeed.tech/tags/building.md>), [data](<https://devfeed.tech/tags/data.md>), [i-built-this](<https://devfeed.tech/tags/i-built-this.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [real-time](<https://devfeed.tech/tags/real-time.md>)

### AI overview

A tutorial-style article about building real-time leaderboards with Tinybird, using live data and real-time analytics that update in milliseconds.

### Source excerpt

Build real-time leaderboards using Tinybird. Engage users with live data and real-time analytics that update in milliseconds.

## Improving Prompt Consistency with Structured Generations

DevFeed: [Improving Prompt Consistency with Structured Generations](<https://devfeed.tech/articles/improving-prompt-consistency-with-structured-generations-7188.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/evaluation-structured-outputs>)

Author: Will Kurt; Remi Louf; Clémentine Fourrier

Published: 2024-04-30T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [Prompt Engineering](<https://devfeed.tech/topics/prompt-engineering.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [tokenization](<https://devfeed.tech/topics/tokenization.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>)

Tags: [benchmark](<https://devfeed.tech/tags/benchmark.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [compare](<https://devfeed.tech/tags/compare.md>), [consistency](<https://devfeed.tech/tags/consistency.md>), [evals](<https://devfeed.tech/tags/evals.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [prompt](<https://devfeed.tech/tags/prompt.md>), [research](<https://devfeed.tech/tags/research.md>), [tokenization](<https://devfeed.tech/tags/tokenization.md>)

### AI overview

The article examines how superficial changes to prompt formatting can substantially affect large language model benchmark scores and model rankings. Using MMLU evaluations across multiple prompt formats and models, it shows that prompt structure introduces significant variance, sometimes because of tokenizer issues, and motivates structured generations to improve consistency.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Roboflow.com choose Supabase to power Paint.wtf leaderboard

DevFeed: [Roboflow.com choose Supabase to power Paint.wtf leaderboard](<https://devfeed.tech/articles/roboflow-com-choose-supabase-to-power-paint-wtf-leaderboard-334.md>)

Original publisher: [Read original article](<https://supabase.com/blog/case-study-roboflow>)

Author: Rory Wilding

Published: 2021-02-09T07:00:00Z

Content type: article

Language: en

Sources: [Supabase Blog](<https://devfeed.tech/sources/supabase-blog.md>)

Topics: [Supabase](<https://devfeed.tech/topics/supabase.md>), [Artificial Intelligence](<https://devfeed.tech/topics/ai.md>), [PostgreSQL](<https://devfeed.tech/topics/postgresql.md>), [BigQuery](<https://devfeed.tech/topics/bigquery.md>), [Firebase](<https://devfeed.tech/topics/firebase.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [applications](<https://devfeed.tech/tags/applications.md>), [computer-vision](<https://devfeed.tech/tags/computer-vision.md>), [data](<https://devfeed.tech/tags/data.md>), [developers](<https://devfeed.tech/tags/developers.md>), [experiment](<https://devfeed.tech/tags/experiment.md>), [image](<https://devfeed.tech/tags/image.md>), [images](<https://devfeed.tech/tags/images.md>), [launch](<https://devfeed.tech/tags/launch.md>), [leaderboards](<https://devfeed.tech/tags/leaderboards.md>), [news](<https://devfeed.tech/tags/news.md>), [openai](<https://devfeed.tech/tags/openai.md>), [performance](<https://devfeed.tech/tags/performance.md>), [postgresql](<https://devfeed.tech/tags/postgresql.md>), [product](<https://devfeed.tech/tags/product.md>), [startup](<https://devfeed.tech/tags/startup.md>), [use-cases](<https://devfeed.tech/tags/use-cases.md>)

### AI overview

Roboflow used Supabase and PostgreSQL to build the leaderboard for Paint.wtf, a weekend project that uses OpenAI's CLIP model to score users' drawings against text prompts. The product handled more than 100,000 drawing submissions in 24 hours after gaining attention on Hacker News, Reddit, and Product Hunt.

### Source excerpt

Learn how Roboflow.com used Supabase to build their Paint.wtf leaderboard