Elicit: Benchmarks for Scientific Literature Review Performance (public)

Which benchmarks measure performance on scientific literature review tasks?

Benchmarks for scientific literature review tasks include datasets like LitSearch, SciReviewGen, SWIFT‐Review, and RoBBR that measure performance across retrieval, generation, screening, extraction, and meta-evaluation areas.

Abstract

Benchmarks for scientific literature review tasks address five primary areas: retrieval, generation, screening, extraction, and meta-evaluation. Seventeen studies target retrieval using datasets such as LitSearch, CLEF TAR, CSMeD, SIGIR2017-PICO-Collection, FASS‐BSLR, RELISH, and ResearchArena. These studies report performance using metrics such as recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, and Work Saved over Sampling. Eleven studies focus on generation tasks—including review, table, or abstract writing—with benchmarks like SciReviewGen, ARXIV2TABLE, MSLR2022, ScholarQABench, and Manalyzer; performance is measured by ROUGE scores, hallucination rates, semantic coverage, and human ratings. Nine studies on screening tasks (using tools such as SWIFT‐Review and LGAR) report measures such as precision at high recall levels and WSS@95, while five studies on extraction or structured analysis (using RoBBR, EvidenceBench, and ECACT) employ metrics such as macro‐F1, accuracy, and Cohen’s kappa. One study adopts a meta‐evaluation benchmark using Elo ratings and inter‐annotator agreement. Overall, papers combine automated retrieval and overlap metrics with human evaluation methods to capture the multifaceted nature of performance in scientific literature review tasks.

Methods

We analyzed 40 sources from an initial pool of 1000, using 5 screening criteria. Each paper was reviewed for 4 key aspects that mattered most to the research question. More on methods

Papers identified with Elicit search

n = 1000

Papers screened using: Benchmark Focus, Computational Literature Review Systems, Beyond Methods Description Only, Literature Review Specific Application, Computational Component

n = 1000

Papers screened out

n = 960

Papers included for extraction

n = 40

Paper search

Using your research question "Which benchmarks measure performance on scientific literature review tasks?", we searched across over 126 million academic papers from the Semantic Scholar corpus. We retrieved the 1000 papers most relevant to the query.

Screening

We screened in sources based on their abstracts that met these criteria:

We considered all screening questions together and made a holistic judgement about whether to screen in each paper.

Data extraction

We asked a large language model to extract each data column below from each paper. We gave the model the extraction instructions shown below for each column.

Identify and describe the specific type of benchmark created in the study. Look for explicit statements about the benchmark’s purpose, focus, and unique characteristics.

Extraction guidelines:

Examples might include:

If no clear benchmark is described, write “Not applicable” or “No benchmark created”

Extract detailed information about the dataset used to create or evaluate the benchmark.

Look for:

Extraction guidelines:

Format examples:

Identify and describe the specific metrics used to assess benchmark performance.

Extraction guidelines:

Look for metrics such as:

Format examples:

Document the models or systems tested against the benchmark.

Extraction guidelines:

Look for:

Format examples:

Results

Characteristics of Included Studies

Study Study Focus Benchmark/Dataset Name Literature Review Task Type Evaluation Approach Full text retrieved
Zhu et al., 2023 Hierarchical catalogue generation for literature reviews HiCaD Hierarchical catalogue generation CEDS (semantic/structural similarity), CQE (informativeness) Yes
Ajith et al., 2024 Literature search retrieval LitSearch Literature search/retrieval Recall at 5, recall at 20, normalized Discounted Cumulative Gain at 10 (nDCG@10) Yes
Kanoulas et al., 2019 Systematic review retrieval and ranking CLEF 2019 e-Health Technology-Assisted Review (TAR) Retrieval, ranking for systematic reviews No mention found in abstract No
Kasanishi et al., 2023 Automatic literature review generation SciReviewGen Literature review generation (summarization) Recall-Oriented Understudy for Gisting Evaluation (ROUGE), human evaluation (relevance, coherence, informativeness, factuality) Yes
Howard et al., 2016 Automated citation screening SWIFT-Review datasets Citation screening Work Saved over Sampling at 95% (WSS@95%), precision at 95% recall Yes
Wang et al., 2025a Literature review table generation ARXIV2TABLE Table generation for literature reviews Recall, precision, F1 for schema/cell/pairwise overlap Yes
Kusa et al., 2023a Outcome-based evaluation of systematic review automation CLEF TAR 2019 Systematic review automation (retrieval) Mean Average Precision (MAP), Recall at k%, WSS, Area Under the Curve (AUC), outcome difference Yes
Kusa et al., 2023b Automated citation screening CSMeD, CSMeD-FT Citation/full-text screening True Negative Rate at 95% (TNR@95%), normalized Precision at 95% (nP@95%), nDCG@10, MAP, macro-precision/recall/F1 Yes
Scells et al., 2017 Retrieval for systematic reviews SIGIR2017-PICO-Collection Retrieval, screening prioritization Precision-recall, F-beta, WSS, Average Precision (AP), nDCG, MAP Yes
Budau and Ensan, 2024 Automated study search for biomedical systematic literature reviews FASS-BSLR Study search (retrieval, Boolean query generation) Precision, Recall, NDCG, MAP, Recall at 1000 No

Distribution of Literature Review Automation Tasks

Distribution of Evaluation Approaches

Some studies addressed multiple task types or used multiple evaluation approaches. We did not find mention of evaluation approach for 6 studies, and task type information was sometimes ambiguous due to overlapping or unclear descriptions.

Thematic Analysis

Benchmarks for Literature Retrieval and Screening Tasks

Benchmarks in this category focus on identifying, retrieving, and prioritizing relevant literature for systematic reviews or evidence synthesis. Key benchmarks include:

Benchmarks for Literature Summarization and Synthesis Tasks

This theme covers benchmarks designed to evaluate the ability of models to generate summaries, reviews, or structured tables from collections of scientific papers. Notable benchmarks include:

Benchmarks for Structured Literature Analysis Tasks

Structured analysis tasks include risk of bias assessment, data element extraction, compliance and traceability, and meta-evaluation. Key benchmarks:

Evaluation Methodologies Across Benchmarks

Benchmark Category

Primary Metrics Evaluation Method Performance Measurement Focus
Retrieval/Screening Recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, Work Saved over Sampling, Average Precision Automated scoring, comparison to human or baseline models
Summarization/Synthesis ROUGE, human ratings, semantic coverage, hallucination rate Automated metrics, human evaluation
Structured Analysis Macro-F1, accuracy, Cohen’s kappa, composite scores (ECACT) Automated scoring, statistical analysis, human validation
Meta-evaluation Elo ratings, accuracy, agreement Human preference voting, model-based evaluator assessment

Summary of Evaluation Approaches: