# open-llm-leaderboard

Published articles for open-llm-leaderboard.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Fixing Open LLM Leaderboard with Math-Verify

DevFeed: [Fixing Open LLM Leaderboard with Math-Verify](<https://devfeed.tech/articles/fixing-open-llm-leaderboard-with-math-verify-7345.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/math_verify_leaderboard>)

Author: Hynek Kydlicek; Alina Lozovskaya; Nathan Habib; Clémentine Fourrier

Published: 2025-02-14T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [math-verify](<https://devfeed.tech/topics/math-verify.md>), [open-llm-leaderboard](<https://devfeed.tech/topics/open-llm-leaderboard.md>), [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Large Language Model](<https://devfeed.tech/topics/llm.md>), [math](<https://devfeed.tech/topics/math.md>), [Parser](<https://devfeed.tech/topics/parser.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [evals](<https://devfeed.tech/tags/evals.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [llm](<https://devfeed.tech/tags/llm.md>), [math-verify](<https://devfeed.tech/tags/math-verify.md>), [open-llm-leaderboard](<https://devfeed.tech/tags/open-llm-leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [science](<https://devfeed.tech/tags/science.md>)

### AI overview

The article explains how Math-Verify was used to re-evaluate 3,751 models submitted to the Open LLM Leaderboard. It addresses answer-format, symbolic parsing, and comparison problems in the previous MATH-Hard evaluator, leading to substantially revised leaderboard results.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## CO₂ Emissions and Models Performance: Insights from the Open LLM Leaderboard

DevFeed: [CO₂ Emissions and Models Performance: Insights from the Open LLM Leaderboard](<https://devfeed.tech/articles/co2-emissions-and-models-performance-insights-from-the-open-llm-leaderboard-7316.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/leaderboard-emissions-analysis>)

Author: Alina Lozovskaya; Nathan Habib; Albert Villanova del Moral; Clémentine Fourrier

Published: 2025-01-09T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [open-llm-leaderboard](<https://devfeed.tech/topics/open-llm-leaderboard.md>), [AI Inference](<https://devfeed.tech/topics/ai-inference.md>), [LLMs](<https://devfeed.tech/topics/llms.md>), [Open Source](<https://devfeed.tech/topics/open-source.md>)

Tags: [architectures](<https://devfeed.tech/tags/architectures.md>), [energy](<https://devfeed.tech/tags/energy.md>), [energy-efficiency](<https://devfeed.tech/tags/energy-efficiency.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [hardware](<https://devfeed.tech/tags/hardware.md>), [inference](<https://devfeed.tech/tags/inference.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [llm](<https://devfeed.tech/tags/llm.md>), [models](<https://devfeed.tech/tags/models.md>), [open-llm-leaderboard](<https://devfeed.tech/tags/open-llm-leaderboard.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [performance](<https://devfeed.tech/tags/performance.md>)

### AI overview

The article analyzes CO₂ emissions from evaluating large language models on the Open LLM Leaderboard. It explains a heuristic based on evaluation time, hardware energy use, and electricity carbon intensity, then examines emissions and performance trends across 2,742 models and several model families.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.

## Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard

DevFeed: [Rethinking LLM Evaluation with 3C3H: AraGen Benchmark and Leaderboard](<https://devfeed.tech/articles/rethinking-llm-evaluation-with-3c3h-aragen-benchmark-and-leaderboard-7308.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/leaderboard-3c3h-aragen>)

Author: Ali El Filali; Neha Sengupta; Abouelseoud; Preslav Nakov; Clémentine Fourrier

Published: 2024-12-04T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [LLM evaluation / benchmarking](<https://devfeed.tech/topics/llm-evaluation-benchmarking.md>), [Benchmark](<https://devfeed.tech/topics/benchmark.md>), [Human-AI evaluation](<https://devfeed.tech/topics/human-ai-evaluation.md>), [datasets](<https://devfeed.tech/topics/datasets.md>), [open-llm-leaderboard](<https://devfeed.tech/topics/open-llm-leaderboard.md>)

Tags: [arabic](<https://devfeed.tech/tags/arabic.md>), [benchmark](<https://devfeed.tech/tags/benchmark.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [llm](<https://devfeed.tech/tags/llm.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-llm-leaderboard](<https://devfeed.tech/tags/open-llm-leaderboard.md>), [performance](<https://devfeed.tech/tags/performance.md>), [research](<https://devfeed.tech/tags/research.md>)

### AI overview

The article introduces the AraGen Benchmark and Leaderboard for evaluating Arabic large language models. It presents the 3C3H Measure, which uses LLM-as-judge assessment across correctness, completeness, conciseness, helpfulness, honesty, and harmlessness. AraGen also uses private three-month blind testing cycles and a multi-turn and single-turn Arabic evaluation dataset to reduce data contamination and assess both factuality and practical usability.

### Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.