# Fact verification benchmark

Published articles for Fact verification benchmark.

This is one page of public article previews, not the complete archive. Follow Next page to continue. Summaries are not the original full articles.

## Ground truth is a process, not a dataset

DevFeed: [Ground truth is a process, not a dataset](<https://devfeed.tech/articles/ground-truth-is-a-process-not-a-dataset-7600.md>)

Original publisher: [Read original article](<https://www.amazon.science/blog/ground-truth-is-a-process-not-a-dataset>)

Author: Venkatesh Saligrama

Published: 2026-06-03T15:56:57Z

Content type: article

Language: en

Sources: [Amazon Science homepage](<https://devfeed.tech/sources/amazon-science-homepage.md>)

Topics: [ai observability](<https://devfeed.tech/topics/ai-observability.md>), [AI Chat](<https://devfeed.tech/topics/ai-chat.md>)

Tags: [ai](<https://devfeed.tech/tags/ai.md>), [ai-fact-checking](<https://devfeed.tech/tags/ai-fact-checking.md>), [ai-generated-research-reports](<https://devfeed.tech/tags/ai-generated-research-reports.md>), [audit-then-score](<https://devfeed.tech/tags/audit-then-score.md>), [benchmarking](<https://devfeed.tech/tags/benchmarking.md>), [conversational-ai](<https://devfeed.tech/tags/conversational-ai.md>), [datasets](<https://devfeed.tech/tags/datasets.md>), [deep-research-verification](<https://devfeed.tech/tags/deep-research-verification.md>), [deepfact-bench](<https://devfeed.tech/tags/deepfact-bench.md>), [evaluation](<https://devfeed.tech/tags/evaluation.md>), [fact-verification](<https://devfeed.tech/tags/fact-verification.md>), [fact-verification-benchmark](<https://devfeed.tech/tags/fact-verification-benchmark.md>), [generative-ai](<https://devfeed.tech/tags/generative-ai.md>), [ground-truth-benchmark-quality](<https://devfeed.tech/tags/ground-truth-benchmark-quality.md>), [hallucination-detection](<https://devfeed.tech/tags/hallucination-detection.md>), [hallucinations](<https://devfeed.tech/tags/hallucinations.md>), [human-ai-evaluation](<https://devfeed.tech/tags/human-ai-evaluation.md>), [large-language-models-llms](<https://devfeed.tech/tags/large-language-models-llms.md>), [llm-evaluation-benchmarking](<https://devfeed.tech/tags/llm-evaluation-benchmarking.md>), [long-context](<https://devfeed.tech/tags/long-context.md>), [machine-learning](<https://devfeed.tech/tags/machine-learning.md>), [responsible-ai](<https://devfeed.tech/tags/responsible-ai.md>), [retrieval-augmented-generation-rag](<https://devfeed.tech/tags/retrieval-augmented-generation-rag.md>)

### AI overview

The article argues that evaluating factuality in long AI-generated research reports requires a process-based approach to ground truth. It introduces audit-then-score and accompanying datasets for benchmarking AI fact checkers.

### Source excerpt

Automatically fact-checking long, AI-generated research reports poses new challenges -- including benchmarking.