ListQA

A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models

Mayank Kulkarni*, Rakesh Vaideeswaran*

Amazon AGI

*Equal contribution

NeurIPS 2026 · Evaluations & Datasets Track
9,045
human-curated, cross-validated questions
3,765
unique Wikipedia source tables
8
topical categories
60+
countries spanned by the entities in questions

Overview

Many real information needs have not one correct answer but an entire set: a traveler asks which direct routes an airline launched since 2022, and a researcher asks which countries ratified a treaty before a given year. The user needs the complete set of qualifying facts, without fabricated additions, organized with the requested attributes. ListQA has two subsets: SimpleListQA (flat lists) and ComplexListQA (hierarchical lists with sub-attributes).

SimpleListQAsingle-faceted · flat list

Who were the actresses selected for Best Actress nominations at the 8th Academy Awards?

  1. Bette Davis
  2. Elisabeth Bergner
  3. Claudette Colbert
  4. Katharine Hepburn
  5. Miriam Hopkins
  6. Merle Oberon
ComplexListQAmulti-faceted · hierarchical list

List the songs and years in which Bruce Springsteen won the Grammy Award for Best Rock Song between 2000 and 2024.

  1. “The Rising”
    • year: 2003
  2. “Radio Nowhere”
    • year: 2008
  3. “Girls in Their Summer Clothes”
    • year: 2009

Abstract

Many real-world questions demand not a single fact but an organized collection of them, yet existing factual knowledge benchmarks almost exclusively target single-answer retrieval. We introduce ListQA, a benchmark of 9,045 human-curated, cross-validated questions (audited error rate of 1.7%) that require LLMs to recall multiple facts and compose them into structured lists of 3–10 elements. Questions range from flat lists (“Which NCAA teams went undefeated between 2000 and 2024?”) to hierarchical lists with sub-attributes (“…and what was their record and result?”), spanning eight categories across 60+ countries. An evergreen design (explicit temporal constraints or historically immutable facts) keeps ground truth valid without periodic updates. For evaluation, we frame element matching as a linear sum assignment problem: optimal bipartite matching paired with LLM-based semantic grading jointly captures factual accuracy, hallucination, and format compliance. We benchmark 34 models (8 families, 42 configurations across standard and thinking-enabled inference) and find: the best standard-inference model scores below 37% LLM-Judge-F1, multi-faceted questions are harder across the board, and providing the source Wikipedia page lifts scores dramatically, confirming that models can organize facts but struggle to retrieve them from parameters alone. Extended analyses cover base vs. instruction-tuned checkpoints and token efficiency.

What Makes ListQA Different

To our knowledge, ListQA is the first benchmark that requires multi-element retrieval, structured output, and joint evaluation of knowledge and instruction following, all at once.

Multi-element retrieval

Every reference answer is a list of 3 to 8 elements. Models are scored on how completely they recall the qualifying facts and how well they avoid adding fabricated ones.

Structured output

SimpleListQA needs flat lists. ComplexListQA needs hierarchical lists with sub-attributes (e.g. recipient, result). Models must pick the right structure for each question.

Evergreen by design

Each question has an explicit time window (e.g. “from 2000 to 2024”) or targets facts that cannot change, so the gold lists stay valid without periodic updates. Answers are validated as of December 31, 2024.

Grounded & verifiable

Every question comes from a Wikipedia table and ships with its source URL, category, time range, question type, and ordering flag.

Comparison with representative benchmarks

BenchmarkMESOEGK+IF
Single-answer factual QA
TriviaQA✕✕✕✕
Natural Questions✕✕✕✕
PopQA✕✕✕✕
SimpleQA✕✕✓✕
FACT-Bench✕✕✕✕
Multi-answer factual QA
QAMPARI✓✕✕✕
RoMQA✓✕✕✕
Multi-hop & reasoning QA
HotpotQA✕✕✕✕
MuSiQue✕✕✕✕
GPQA✕✕✕✕
BenchmarkMESOEGK+IF
Structured & table QA
WikiTableQuestions◐✕✕✕
HybridQA✕✕✕✕
Temporal & dynamic
FreshQA✕✕◐✕
RecencyQA✕✕◐✕
Long-form factuality
FActScore◐✕✕✕
FACTS Grounding◐✕✕✓
Knowledge calibration
SelfAware✕✕✕✕
ListQA (ours)✓✓✓✓

supported  ·  partial  ·  not supported  ·  ME multi-element retrieval, SO structured output, EG evergreen design, K+IF knowledge + instruction following.

Dataset

The full dataset contains 9,045 question–answer pairs sourced from 3,765 unique Wikipedia tables. Each sample comes with metadata: the source Wikipedia URL, topical category, temporal range, question type, and ordering requirement. That makes the ground truth independently checkable and supports fine-grained analysis along any of these dimensions.

Key statistics
Questions9,045
SimpleListQA / ComplexListQA59.5% / 40.5%
Unique Wikipedia source tables3,765
Elements per answer (median / mean)3–8 (4 / 4.65)
Sub-attributes per element (Complex)1–8
Year-constrained / timeless70.1% / 29.9%
Ordered / unordered14.6% / 85.4%
Countries represented60+ (6 continents)

Four-stage quality assurance

Pilot calibration

Two pilot rounds calibrated how annotators understood the task. Pilot data is not in the final dataset.

Author & cross-validate

Each QA pair was written by one of 44 annotators and independently checked by a second annotator who had no part in writing it (100% coverage).

Flag, fix or discard

Validators flagged issues using a fixed error taxonomy. Samples that needed major rework were discarded (10%). The rest were corrected.

Verify corrections

At least one more reviewer checked every correction, and two reviewers checked 20% of them.

Evaluation

Judging a predicted list as a whole correlates poorly with human evaluation. Instead, ListQA grades each predicted element against its reference counterpart with an LLM judge. To keep that affordable, element matching is framed as a linear sum assignment problem, which cuts the number of judge calls from O(n²) to O(n).

LLM-Judge-F1 pipeline: list element parser, ROUGE linear sum assignment, and pairwise LLM judge evaluation
How LLM-Judge-F1 scores a list answer (illustrative example). (1) A regex parser extracts list elements from both the reference and the prediction. (2) Pairwise ROUGE-1 scores form a cost matrix, and an optimal one-to-one assignment (Jonker–Volgenant) pairs predicted and reference elements. (3) An LLM judge grades each matched pair CORRECT or INCORRECT. Unmatched predicted elements count as INCORRECT. Here one hallucinated element (“Walking Alive”) gives P = R = F1 = 0.75.

Try it: a worked scoring example

Metrics

Knowledge + format · semantic

LLM-Judge-F1 (headline metric)

The headline factual metric. It uses the parse → assign → judge pipeline above, and precision and recall are computed over elements the judge marks CORRECT. Only content inside the prescribed list structure earns credit.

Instruction following

Format-Correctness (headline metric)

Checks list structure only, not facts. Single-faceted questions need a flat list and multi-faceted questions need sub-bullets. It is the fraction of correctly structured elements, and 0 if fewer than 3 elements are produced.

Knowledge · format-agnostic

Relaxed-Recall

The fraction of reference elements (and sub-element values) found by string matching anywhere in the raw model output, whether or not that output is formatted as a list.

Hallucination · string match

Relaxed-Precision

The fraction of predicted list elements whose content matches the reference by string matching. It penalizes fabricated extras.

Knowledge + format · lexical

ROUGE-F1

The same optimal assignment, scored by ROUGE-1 overlap (a pair counts as matched at ≥ 0.5). A cheaper but less discriminative proxy for LLM-Judge-F1.

Sequencing

LLM-Judge-F1-Ordered

For the 1,319 questions (14.6%) that need a specific order. Elements are paired by position instead of by optimal assignment, then judged, so correct elements in the wrong place are penalized.

Dataset Explorer

Filter by subset, category, region, country, time window, list length, or ordering, and search the question text or the reference answers.

Citation

If you use ListQA in your research, please cite:

@inproceedings{kulkarni2026listqa,
  title     = {ListQA: A Benchmark for Evaluating List-Formatted Factual Knowledge Retrieval in Large Language Models},
  author    = {Kulkarni, Mayank and Vaideeswaran, Rakesh},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
  year      = {2026}
}