Elicit: Benchmarks for Scientific Literature Review Performance (public)
Which benchmarks measure performance on scientific literature review tasks?
Benchmarks for scientific literature review tasks include datasets like LitSearch, SciReviewGen, SWIFT‐Review, and RoBBR that measure performance across retrieval, generation, screening, extraction, and meta-evaluation areas.
Abstract
Benchmarks for scientific literature review tasks address five primary areas: retrieval, generation, screening, extraction, and meta-evaluation. Seventeen studies target retrieval using datasets such as LitSearch, CLEF TAR, CSMeD, SIGIR2017-PICO-Collection, FASS‐BSLR, RELISH, and ResearchArena. These studies report performance using metrics such as recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, and Work Saved over Sampling. Eleven studies focus on generation tasks—including review, table, or abstract writing—with benchmarks like SciReviewGen, ARXIV2TABLE, MSLR2022, ScholarQABench, and Manalyzer; performance is measured by ROUGE scores, hallucination rates, semantic coverage, and human ratings. Nine studies on screening tasks (using tools such as SWIFT‐Review and LGAR) report measures such as precision at high recall levels and WSS@95, while five studies on extraction or structured analysis (using RoBBR, EvidenceBench, and ECACT) employ metrics such as macro‐F1, accuracy, and Cohen’s kappa. One study adopts a meta‐evaluation benchmark using Elo ratings and inter‐annotator agreement. Overall, papers combine automated retrieval and overlap metrics with human evaluation methods to capture the multifaceted nature of performance in scientific literature review tasks.
Methods
We analyzed 40 sources from an initial pool of 1000, using 5 screening criteria. Each paper was reviewed for 4 key aspects that mattered most to the research question.
Screening
We screened in sources based on their abstracts that met these criteria:
- Benchmark Focus: Does this study describe, develop, evaluate, present datasets for, or validate benchmarks/evaluation frameworks specifically designed to measure performance on scientific literature review tasks?
- Computational Literature Review Systems: Does this study focus on computational systems that perform literature search, screening, extraction, synthesis, or other core literature review functions, OR does it synthesize knowledge about literature review benchmarks through systematic/narrative reviews?
- Beyond Methods Description Only: Does this study go beyond only describing literature review methods to actually present, evaluate, or analyze benchmarks or evaluation metrics?
- Literature Review Specific Application: Are the benchmarks or evaluation approaches specifically applied to literature review tasks (rather than focusing solely on general information retrieval or text mining without literature review application)?
- Computational Component: Does this study include computational components rather than focusing exclusively on manual literature review processes?
Data extraction
We asked a large language model to extract each data column below from each paper. We gave the model the extraction instructions shown below for each column.
Benchmark Type: Identify and describe the specific type of benchmark created in the study. Look for explicit statements about the benchmark’s purpose, focus, and unique characteristics.
Benchmark Dataset Characteristics: Extract detailed information about the dataset used to create or evaluate the benchmark.
Evaluation Metrics: Identify and describe the specific metrics used to assess benchmark performance.
Benchmark Models and Comparisons: Document the models or systems tested against the benchmark.
Results
Characteristics of Included Studies
| Study | Study Focus | Benchmark/Dataset Name | Literature Review Task Type | Evaluation Approach | Full text retrieved |
|---|---|---|---|---|---|
| Zhu et al., 2023 | Hierarchical catalogue generation for literature reviews | HiCaD | Hierarchical catalogue generation | CEDS (semantic/structural similarity), CQE (informativeness) | Yes |
| Ajith et al., 2024 | Literature search retrieval | LitSearch | Literature search/retrieval | Recall at 5, recall at 20, normalized Discounted Cumulative Gain at 10 (nDCG@10) | Yes |
| Kanoulas et al., 2019 | Systematic review retrieval and ranking | CLEF 2019 e-Health Technology-Assisted Review (TAR) | Retrieval, ranking for systematic reviews | No mention found in abstract | No |
| Kasanishi et al., 2023 | Automatic literature review generation | SciReviewGen | Literature review generation | Recall-Oriented Understudy for Gisting Evaluation (ROUGE), human evaluation (relevance, coherence, informativeness, factuality) | Yes |
| Howard et al., 2016 | Automated citation screening | SWIFT-Review datasets | Citation screening | Work Saved over Sampling at 95% (WSS@95%), precision at 95% recall | Yes |
| Wang et al., 2025a | Literature review table generation | ARXIV2TABLE | Table generation for literature reviews | Recall, precision, F1 for schema/cell/pairwise overlap | Yes |
| Kusa et al., 2023a | Outcome-based evaluation of systematic review automation | CLEF TAR 2019 | Systematic review automation (retrieval) | Mean Average Precision (MAP), Recall at k%, WSS, Area Under the Curve (AUC), outcome difference | Yes |
| ... | ... | ... | ... | ... | ... |
Distribution of Literature Review Automation Tasks
- Retrieval, search, or ranking tasks: Addressed in 17 studies.
- Generation tasks: Addressed in 11 studies.
- Screening tasks: Addressed in 9 studies.
- Extraction tasks: Addressed in 5 studies.
- Meta-evaluation or meta-analysis: Addressed in 1 study.
- Other tasks: Addressed in 10 studies.
Distribution of Evaluation Approaches
- Information retrieval metrics: Used in 27 studies.
- Human evaluation: Used in 6 studies.
- Overlap, coverage, or consistency metrics: Used in 6 studies.
- Specialized or composite metrics: Used in 7 studies.
- No mention found of evaluation approach: In 6 studies.
Thematic Analysis
Benchmarks for Literature Retrieval and Screening Tasks
Benchmarks in this category focus on identifying, retrieving, and prioritizing relevant literature for systematic reviews or evidence synthesis. Key benchmarks include:
- LitSearch, CLEF Technology-Assisted Review (TAR) (2017–2019), CSMeD, SIGIR2017-PICO-Collection, FASS-BSLR, RELISH, and ResearchArena.
Benchmarks for Literature Summarization and Synthesis Tasks
This theme covers benchmarks designed to evaluate the ability of models to generate summaries, reviews, or structured tables from collections of scientific papers. Notable benchmarks include: SciReviewGen, ARXIV2TABLE, MSLR2022, ScholarQABench, and Manalyzer.
Benchmarks for Structured Literature Analysis Tasks
Structured analysis tasks include risk of bias assessment, data element extraction, compliance and traceability, and meta-evaluation. Key benchmarks: RoBBR, EvidenceBench, ECACT, SciArena-Eval, and internal datasets.
Evaluation Methodologies Across Benchmarks
| Benchmark Category | Primary Metrics | Evaluation Method | Performance Measurement Focus |
|---|---|---|---|
| Retrieval/Screening | Recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, Work Saved over Sampling, Average Precision | Automated scoring, comparison to human or baseline models | Relevance, workload reduction, ranking effectiveness |
| Summarization/Synthesis | ROUGE, human ratings, semantic coverage, hallucination rate | Automated metrics, human evaluation | Informativeness, factuality, coherence, citation accuracy |
| Structured Analysis | Macro-F1, accuracy, Cohen’s kappa, composite scores (ECACT) | Automated scoring, statistical analysis, human validation | Bias, compliance, extraction accuracy, traceability |
| Meta-evaluation | Elo ratings, accuracy, agreement | Human preference voting, model-based evaluator assessment | Alignment with human judgment, evaluator reliability |
References
- Wang et al., 2024 - Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark
- Zhu et al., 2023 - Hierarchical Catalogue Generation for Literature Review: A Benchmark
- Additional references continue similarly.