# Judge Arena: Benchmarking LLMs as Evaluators

DevFeed: [Judge Arena: Benchmarking LLMs as Evaluators](<https://devfeed.tech/articles/judge-arena-benchmarking-llms-as-evaluators-7099.md>)

Original publisher: [Read original article](<https://huggingface.co/blog/arena-atla>)

Author: kyle; Maurice; Roman Engeler; Max Bartolo; Clémentine Fourrier; Toby Drane; Mathias Leys; Jake Golden

Published: 2024-11-19T00:00:00Z

Content type: article

Language: en

Sources: [Hugging Face - Blog](<https://devfeed.tech/sources/hugging-face-blog.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [anthropic](<https://devfeed.tech/tags/anthropic.md>), [api](<https://devfeed.tech/tags/api.md>), [arena](<https://devfeed.tech/tags/arena.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [collaboration](<https://devfeed.tech/tags/collaboration.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [gemma](<https://devfeed.tech/tags/gemma.md>), [google](<https://devfeed.tech/tags/google.md>), [human-feedback](<https://devfeed.tech/tags/human-feedback.md>), [leaderboard](<https://devfeed.tech/tags/leaderboard.md>), [llama](<https://devfeed.tech/tags/llama.md>), [llms](<https://devfeed.tech/tags/llms.md>), [nlp](<https://devfeed.tech/tags/nlp.md>), [open-source](<https://devfeed.tech/tags/open-source.md>), [openai](<https://devfeed.tech/tags/openai.md>), [qwen](<https://devfeed.tech/tags/qwen.md>)

## AI overview

Judge Arena is a crowdsourced platform for comparing LLMs used as evaluators. Users review two judges' scores and critiques of a response, vote for the evaluation they prefer, and contribute to a leaderboard.

## Source excerpt

We're on a journey to advance and democratize artificial intelligence through open source and open science.