Elicit: Benchmarks for Scientific Literature Review Performance (public)
Which benchmarks measure performance on scientific literature review tasks?
Benchmarks for scientific literature review tasks include datasets like LitSearch, SciReviewGen, SWIFT‐Review, and RoBBR that measure performance across retrieval, generation, screening, extraction, and meta-evaluation areas.
Abstract
Benchmarks for scientific literature review tasks address five primary areas: retrieval, generation, screening, extraction, and meta-evaluation. Seventeen studies target retrieval using datasets such as LitSearch, CLEF TAR, CSMeD, SIGIR2017-PICO-Collection, FASS‐BSLR, RELISH, and ResearchArena. These studies report performance using metrics such as recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, and Work Saved over Sampling. Eleven studies focus on generation tasks—including review, table, or abstract writing—with benchmarks like SciReviewGen, ARXIV2TABLE, MSLR2022, ScholarQABench, and Manalyzer; performance is measured by ROUGE scores, hallucination rates, semantic coverage, and human ratings. Nine studies on screening tasks (using tools such as SWIFT‐Review and LGAR) report measures such as precision at high recall levels and WSS@95, while five studies on extraction or structured analysis (using RoBBR, EvidenceBench, and ECACT) employ metrics such as macro‐F1, accuracy, and Cohen’s kappa. One study adopts a meta‐evaluation benchmark using Elo ratings and inter‐annotator agreement. Overall, papers combine automated retrieval and overlap metrics with human evaluation methods to capture the multifaceted nature of performance in scientific literature review tasks.
Methods
We analyzed 40 sources from an initial pool of 1000, using 5 screening criteria. Each paper was reviewed for 4 key aspects that mattered most to the research question.
Paper search
Using your research question “Which benchmarks measure performance on scientific literature review tasks?”, we searched across over 126 million academic papers from the Semantic Scholar corpus. We retrieved the 1000 papers most relevant to the query.
Screening
We screened in sources based on their abstracts that met these criteria:
- Benchmark Focus: Does this study describe, develop, evaluate, present datasets for, or validate benchmarks/evaluation frameworks specifically designed to measure performance on scientific literature review tasks?
- Computational Literature Review Systems: Does this study focus on computational systems that perform literature search, screening, extraction, synthesis, or other core literature review functions, OR does it synthesize knowledge about literature review benchmarks through systematic/narrative reviews?
- Beyond Methods Description Only: Does this study go beyond only describing literature review methods to actually present, evaluate, or analyze benchmarks or evaluation metrics?
- Literature Review Specific Application: Are the benchmarks or evaluation approaches specifically applied to literature review tasks (rather than focusing solely on general information retrieval or text mining without literature review application)?
- Computational Component: Does this study include computational components rather than focusing exclusively on manual literature review processes?
Data extraction
We asked a large language model to extract each data column below from each paper. We gave the model the extraction instructions shown below for each column.
Benchmark Type:
- Identify and describe the specific type of benchmark created in the study. Look for explicit statements about the benchmark’s purpose, focus, and unique characteristics.
- Extraction guidelines:
- Locate the benchmark description in the introduction or methods section
- Capture the specific scientific literature review task the benchmark addresses
- If multiple benchmark types are described, list all of them
- Be precise about the benchmark’s scope (e.g., retrieval, generation, search)
- Examples might include:
- Hierarchical catalogue generation benchmark
- Literature search retrieval benchmark
- Information distillation benchmark
Benchmark Dataset Characteristics:
- Extract detailed information about the dataset used to create or evaluate the benchmark.
- Look for:
- Total number of items in the dataset
- Types of items (e.g., research papers, citations, queries)
- Domain or field of research covered
- Data collection method
- Source of data items
- Extraction guidelines:
- Prioritize quantitative details (e.g., “7.6k literature review catalogues and 389k reference papers”)
- Note the specific research domains if mentioned
- If multiple datasets are used, list all with their characteristics
- If dataset details are incomplete, note “Insufficient information”
- Format examples:
- “7,600 literature review catalogues from scientific papers”
- “597 literature search queries across ML and NLP domains”
Evaluation Metrics:
- Identify and describe the specific metrics used to assess benchmark performance.
- Extraction guidelines:
- Locate metrics in results, methods, or discussion sections
- Capture both quantitative metrics and qualitative assessment approaches
- Note the specific performance dimensions being measured
- If multiple metrics are used, list all of them
- Look for metrics such as:
- Recall@5
- Semantic similarity scores
- Informativeness ratings
- Comparative performance against baseline models
Benchmark Models and Comparisons:
- Document the models or systems tested against the benchmark.
- Extraction guidelines:
- List all models/systems evaluated
- Note their performance relative to the benchmark
- Capture any comparative analysis between different approaches
- Include both state-of-the-art and baseline models
- Look for:
- Specific model names (e.g., BART, ChatGPT)
- Performance rankings
- Comparative performance percentages
Results
Characteristics of Included Studies
| Study | Study Focus | Benchmark/Dataset Name | Literature Review Task Type | Evaluation Approach | Full text retrieved |
|---|---|---|---|---|---|
| Zhu et al., 2023 | Hierarchical catalogue generation for literature reviews | HiCaD | Hierarchical catalogue generation | CEDS (semantic/structural similarity), CQE (informativeness) | Yes |
| Ajith et al., 2024 | Literature search retrieval | LitSearch | Literature search/retrieval | Recall at 5, recall at 20, normalized Discounted Cumulative Gain at 10 (nDCG@10) | Yes |
| Kanoulas et al., 2019 | Systematic review retrieval and ranking | CLEF 2019 e-Health Technology-Assisted Review (TAR) | Retrieval, ranking for systematic reviews | No mention found in abstract | No |
| Kasanishi et al., 2023 | Automatic literature review generation | SciReviewGen | Literature review generation (summarization) | Recall-Oriented Understudy for Gisting Evaluation (ROUGE), human evaluation (relevance, coherence, informativeness, factuality) | Yes |
| Howard et al., 2016 | Automated citation screening | SWIFT-Review datasets | Citation screening | Work Saved over Sampling at 95% (WSS@95%), precision at 95% recall | Yes |
| Wang et al., 2025a | Literature review table generation | ARXIV2TABLE | Table generation for literature reviews | Recall, precision, F1 for schema/cell/pairwise overlap | Yes |
| Kusa et al., 2023a | Outcome-based evaluation of systematic review automation | CLEF TAR 2019 | Systematic review automation (retrieval) | Mean Average Precision (MAP), Recall at k%, WSS, Area Under the Curve (AUC), outcome difference | Yes |
| Kusa et al., 2023b | Automated citation screening | CSMeD, CSMeD-FT | Citation/full-text screening | True Negative Rate at 95% (TNR@95%), normalized Precision at 95% (nP@95%), nDCG@10, MAP, macro-precision/recall/F1 | Yes |
| Scells et al., 2017 | Retrieval for systematic reviews | SIGIR2017-PICO-Collection | Retrieval, screening prioritization | Precision-recall, F-beta, WSS, Average Precision (AP), nDCG, MAP | Yes |
| Budau and Ensan, 2024 | Automated study search for biomedical systematic literature reviews | FASS-BSLR | Study search (retrieval, Boolean query generation) | Precision, Recall, NDCG, MAP, Recall at 1000 | No |
Distribution of Literature Review Automation Tasks
- Retrieval, search, or ranking tasks: Addressed in 17 studies. These focus on identifying and prioritizing relevant literature for systematic reviews or evidence synthesis.
- Generation tasks: Addressed in 11 studies. These include review, table, summary, reference, or abstract generation.
- Screening tasks: Addressed in 9 studies. These include citation, abstract, or full-text screening.
- Extraction tasks: Addressed in 5 studies. These include data, evidence, risk of bias, or meta-analysis.
- Meta-evaluation or meta-analysis: Addressed in 1 study.
- Other tasks: Addressed in 10 studies. These include query generation, compliance, information distillation, academic survey, trial design, systematic review-style reasoning, or multiple systematic literature review tasks.
Distribution of Evaluation Approaches
- Information retrieval metrics: Used in 27 studies. Metrics include precision, recall, F1, Mean Average Precision (MAP), normalized Discounted Cumulative Gain (nDCG), Mean Reciprocal Rank (MRR), Average Precision (AP), and others.
- Human evaluation: Used in 6 studies. Includes inter-annotator agreement, Cohen’s kappa, or explicit human ratings.
- Overlap, coverage, or consistency metrics: Used in 6 studies. Includes schema/cell overlap, coverage, factual consistency, or hallucination rate.
- Specialized or composite metrics: Used in 7 studies. Includes Elo ratings, CEDS, CQE, ECACT, or outcome difference.
- No mention found of evaluation approach: In 6 studies.
Thematic Analysis
Benchmarks for Literature Retrieval and Screening Tasks
Benchmarks in this category focus on identifying, retrieving, and prioritizing relevant literature for systematic reviews or evidence synthesis. Key benchmarks include:
- LitSearch, CLEF Technology-Assisted Review (TAR) (2017–2019), CSMeD, SIGIR2017-PICO-Collection, FASS-BSLR, RELISH, and ResearchArena: These benchmarks use large-scale datasets such as PubMed, MEDLINE, and S2ORC, and evaluate models using recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, Work Saved over Sampling, and related metrics.
- Model types: Both classical information retrieval models (such as BM25 and term frequency-inverse document frequency) and modern transformer-based or large language model approaches are tested.
- Findings: Some studies (e.g., Kang and Xiong, 2024) report that keyword-based retrieval still outperforms large language model-based methods in certain settings, while others (e.g., Skarlinski et al., 2024) demonstrate large language models matching or exceeding human performance in specific retrieval and synthesis tasks.
Benchmarks for Literature Summarization and Synthesis Tasks
This theme covers benchmarks designed to evaluate the ability of models to generate summaries, reviews, or structured tables from collections of scientific papers. Notable benchmarks include:
- SciReviewGen, ARXIV2TABLE, MSLR2022, ScholarQABench, and Manalyzer: Evaluation metrics include ROUGE, human preference ratings, semantic coverage, factual consistency, and hallucination rates.
- Findings: Some large language models and transformer-based models achieve promising results, but challenges remain in factuality, citation accuracy, and domain adaptation. Human evaluation is often used to supplement automated metrics, reflecting the complexity of assessing synthesis quality.
Benchmarks for Structured Literature Analysis Tasks
Structured analysis tasks include risk of bias assessment, data element extraction, compliance and traceability, and meta-evaluation. Key benchmarks:
- RoBBR, EvidenceBench, ECACT, SciArena-Eval, and internal datasets: These benchmarks use composite or task-specific metrics such as ECACT score, macro-F1, accuracy, and Cohen’s kappa.
- Findings: Automation is advancing, but human oversight remains essential, especially in high-stakes or regulatory contexts.
Evaluation Methodologies Across Benchmarks
| Benchmark Category | Primary Metrics | Evaluation Method | Performance Measurement Focus |
|---|---|---|---|
| Retrieval/Screening | Recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, Work Saved over Sampling, Average Precision | Automated scoring, comparison to human or baseline models | Relevance, workload reduction, ranking effectiveness |
| Summarization/Synthesis | ROUGE, human ratings, semantic coverage, hallucination rate | Automated metrics, human evaluation | Informativeness, factuality, coherence, citation accuracy |
| Structured Analysis | Macro-F1, accuracy, Cohen’s kappa, composite scores (ECACT) | Automated scoring, statistical analysis, human validation | Bias, compliance, extraction accuracy, traceability |
| Meta-evaluation | Elo ratings, accuracy, agreement | Human preference voting, model-based evaluator assessment | Alignment with human judgment, evaluator reliability |
Summary of Evaluation Approaches:
- Primary metrics: Accuracy was the only metric used in more than one category (2 out of 4 categories). All other metrics were unique to a single category, including recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, Work Saved over Sampling, Average Precision, ROUGE, human ratings, semantic coverage, hallucination rate, macro-F1, Cohen’s kappa, composite scores (ECACT), Elo ratings, and agreement.
- Evaluation methods: Automated evaluation methods were used in 3 out of 4 categories. Human-based evaluation methods were used in all 4 categories. Comparison-based methods, statistical analysis, and model-based evaluator assessment were each used in 1 out of 4 categories.
- Coverage: We did not find any categories without at least one automated and one human-based evaluation method. No information was missing for any category.
References
- Jianyou Wang, Weili Cao, Longtian Bao, Youze Zheng, Gil Pasternak, et al. (2024). Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark. arXiv.org
- Kun Zhu, Xiaocheng Feng, Xiachong Feng, Yin Wu, Bing Qin. (2023). Hierarchical Catalogue Generation for Literature Review: A Benchmark. Conference on Empirical Methods in Natural Language Processing
- Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, et al. (2024). LitSearch: A Retrieval Benchmark for Scientific Literature Search. Conference on Empirical Methods in Natural Language Processing
- David W. Brett, Anniek Myatt. (2025). Patience is all you need! An agentic system for performing scientific literature review. arXiv.org
- E. Kanoulas, Dan Li, L. Azzopardi, R. Spijker. (2018). CLEF 2017 Technologically Assisted Reviews in Empirical Medicine Overview. Conference and Labs of the Evaluation Forum
- Lucy Lu Wang, Jay DeYoung, Byron Wallace. (2022). Overview of MSLR2022: A Shared Task on Multi-document Summarization for Literature Reviews. SDP
- E. Kanoulas, Dan Li, Leif Azzopardi, René Spijker. (2019). CLEF 2019 Technology Assisted Reviews in Empirical Medicine Overview. Conference and Labs of the Evaluation Forum
- E. Kanoulas, Dan Li, Leif Azzopardi, René Spijker. (2018). CLEF 2018 Technologically Assisted Reviews in Empirical Medicine Overview. Conference and Labs of the Evaluation Forum
- Tetsu Kasanishi, Masaru Isonuma, Junichiro Mori, I. Sakata. (2023). SciReviewGen: A Large-scale Dataset for Automatic Literature Review Generation. Annual Meeting of the Association for Computational Linguistics
- Brian E. Howard, Jason R. Phillips, Kyle Miller, Arpit Tandon, D. Mav, et al. (2016). SWIFT-Review: a text-mining workbench for systematic review. Systematic Reviews
- Weiqi Wang, Jiefu Ou, Yangqiu Song, Benjamin Van Durme, Daniel Khashabi. (2025). Can LLMs Generate Tabular Summaries of Science Papers? Rethinking the Evaluation Protocol. arXiv.org
- Wojciech Kusa, G. Zuccon, Petr Knoth, A. Hanbury. (2023). Outcome-based Evaluation of Systematic Review Automation. International Conference on the Theory of Information Retrieval
- Wojciech Kusa, Óscar E. Mendoza, Matthias Samwald, Petr Knoth, Allan Hanbury. (2023). CSMeD: Bridging the Dataset Gap in Automated Citation Screening for Systematic Literature Reviews. Neural Information Processing Systems
- Harrisen Scells, G. Zuccon, B. Koopman, Anthony J Deacon, L. Azzopardi, et al. (2017). A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews. Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
- Leandra Budau, F. Ensan. (2024). Fully Automated Scholarly Search for Biomedical Systematic Literature Reviews. IEEE Access
- Yilun Zhao, Kaiyan Zhang, Tiansheng Hu, Sihong Wu, Ronan Le Bras, et al. (2025). SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks. arXiv.org
- Joshua Morriss, Tod Brindle, Jessica Bah Rösman, Daniel Reibsamen, Andreas Enz. (2024). The Literature Review Network: An Explainable Artificial Intelligence for Systematic Literature Reviews, Meta-analyses, and Method Development. arXiv.org
- Christian Jaumann, Andreas Wiedholz, Annemarie Friedrich. (2025). LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews. Annual Meeting of the Association for Computational Linguistics
- Mohammadreza Pourreza, F. Ensan. (2022). Towards semantic-driven boolean query formalization for biomedical systematic literature reviews. Int. J. Medical Informatics
- Maxime Gobin, Muriel Gosnat, Seindé Toure, Lina Faik, Joel Belafa, et al. (2025). From data extraction to analysis: a comparative study of ELISE capabilities in scientific literature. Frontiers in Artificial Intelligence
- Shubham Agarwal, Gaurav Sahu, Abhay Puri, I. Laradji, K. Dvijotham, et al. (2024). LitLLMs, LLMs for Literature Review: Are we there yet?. Trans. Mach. Learn. Res.
- Peter Brown, Ameya Sadguru Kulkarni, Osama Refai, Yaoqi Zhou, Abd Al-Bar Al-Farha. (2019). Large expert-curated database for benchmarking document similarity detection in biomedical literature search. Database J. Biol. Databases Curation
- Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess, Yuhui Zhang, et al. (2025). Can Large Language Models Match the Conclusions of Systematic Reviews?. arXiv.org
- Xuemei Tang, Xufeng Duan, Zhenguang G. Cai. (2024). Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
- Tanja Bekhuis, Eugene Tseytlin, K. Mitchell, Dina Demner-Fushman. (2014). Feature Engineering and a Proposed Decision-Support System for Systematic Reviewers of Medical Evidence. PLoS ONE
- Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela M. Hinks, et al. (2024). Language agents achieve superhuman synthesis of scientific knowledge. arXiv.org
- A. Lozano, Scott L. Fleming, Chia-Chun Chiang, Nigam H. Shah. (2023). Clinfo.ai: An Open-Source Retrieval-Augmented Large Language Model System for Answering Medical Questions using Scientific Literature. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
- Grace E. Lee, Aixin Sun. (2018). Seed-driven Document Ranking for Systematic Reviews in Evidence-Based Medicine. Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
- Alon Gorenshtein, Kamel Shihada, Moran Sorka, Dvir Aran, Shahar Shelly. (2025). LITERAS: Biomedical literature review and citation retrieval agents. Comput. Biol. Medicine
- Gaelen Adam, Melinda Davies, Jerusha George, Eduardo L Caputo, Ja Mai Htun, et al. (2025). Machine Learning Tools To (Semi-)Automate Evidence Synthesis: A Rapid Review and Evidence Map
- Aditya Nagori, Ricardo Accorsi Casonatto, Ayush Gautam, Abhinav Manikantha Sai Cheruvu, Rishikesan Kamaleswaran. (2025). Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
- G. Adam, Melinda Davies, Jerusha George, Eduardo L Caputo, Ja Mai Htun, et al. (2025). Machine Learning Tools To (Semi-) Automate Evidence Synthesis
- Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, et al. (2024). OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs. arXiv.org
- Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao. (2025). DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents. arXiv.org
- Hao Kang, Chenyan Xiong. (2024). ResearchArena: Benchmarking LLMs' Ability to Collect and Organize Information as Research Agents. arXiv.org
- Zifeng Wang, Qiao Jin, Jiacheng Lin, Junyi Gao, Jathurshan Pradeepkumar, et al. (2025). TrialPanorama: Database and Benchmark for Systematic Review and Design of Clinical Trials. arXiv.org
- Jianyou Wang, Weili Cao, Kaicheng Wang, Xiaoyue Wang, Ashish Dalvi, et al. (2025). EvidenceBench: A Benchmark for Extracting Evidence from Biomedical Papers. arXiv.org
- Wanghan Xu, Wenlong Zhang, Fenghua Ling, Ben Fei, Yusong Hu, et al. (2025). Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System. arXiv.org
- Jingcheng Du, Dong Wang, Bin Lin, Long He, Liang-chin Huang, et al. (2025). Use of deep learning-based NLP models for full-text data elements extraction for systematic literature review tasks. Scientific Reports
- Hao Kang, Chenyan Xiong. (2024). ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents
Report
Status: Gather sources 1000 sources found
Details: Screen sources 40 sources included
Details: Extract data 160 data points extracted
Details: Generate report Save PDF