Elicit: Benchmarks for Scientific Literature Review Performance (public)

Which benchmarks measure performance on scientific literature review tasks?

Benchmarks for scientific literature review tasks include datasets like LitSearch, SciReviewGen, SWIFT‐Review, and RoBBR that measure performance across retrieval, generation, screening, extraction, and meta-evaluation areas.

Abstract

Benchmarks for scientific literature review tasks address five primary areas: retrieval, generation, screening, extraction, and meta-evaluation. Seventeen studies target retrieval using datasets such as LitSearch, CLEF TAR, CSMeD, SIGIR2017-PICO-Collection, FASS‐BSLR, RELISH, and ResearchArena. These studies report performance using metrics such as recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, and Work Saved over Sampling. Eleven studies focus on generation tasks—including review, table, or abstract writing—with benchmarks like SciReviewGen, ARXIV2TABLE, MSLR2022, ScholarQABench, and Manalyzer; performance is measured by ROUGE scores, hallucination rates, semantic coverage, and human ratings. Nine studies on screening tasks (using tools such as SWIFT‐Review and LGAR) report measures such as precision at high recall levels and WSS@95, while five studies on extraction or structured analysis (using RoBBR, EvidenceBench, and ECACT) employ metrics such as macro‐F1, accuracy, and Cohen’s kappa. One study adopts a meta‐evaluation benchmark using Elo ratings and inter‐annotator agreement. Overall, papers combine automated retrieval and overlap metrics with human evaluation methods to capture the multifaceted nature of performance in scientific literature review tasks.

Methods

We analyzed 40 sources from an initial pool of 1000, using 5 screening criteria. Each paper was reviewed for 4 key aspects that mattered most to the research question.

Paper search

Using your research question “Which benchmarks measure performance on scientific literature review tasks?”, we searched across over 126 million academic papers from the Semantic Scholar corpus. We retrieved the 1000 papers most relevant to the query.

Screening

We screened in sources based on their abstracts that met these criteria:

Data extraction

We asked a large language model to extract each data column below from each paper. We gave the model the extraction instructions shown below for each column.

Results

Characteristics of Included Studies

Study Study Focus Benchmark/Dataset Name Literature Review Task Type Evaluation Approach Full text retrieved
Zhu et al., 2023 Hierarchical catalogue generation for literature reviews HiCaD Hierarchical catalogue generation CEDS (semantic/structural similarity), CQE (informativeness) Yes
Ajith et al., 2024 Literature search retrieval LitSearch Literature search/retrieval Recall at 5, recall at 20, normalized Discounted Cumulative Gain at 10 (nDCG@10) Yes
Kanoulas et al., 2019 Systematic review retrieval and ranking CLEF 2019 e-Health Technology-Assisted Review (TAR) Retrieval, ranking for systematic reviews No mention found in abstract No
Kasanishi et al., 2023 Automatic literature review generation SciReviewGen Literature review generation (summarization) Recall-Oriented Understudy for Gisting Evaluation (ROUGE), human evaluation (relevance, coherence, informativeness, factuality) Yes
Howard et al., 2016 Automated citation screening SWIFT-Review datasets Citation screening Work Saved over Sampling at 95% (WSS@95%), precision at 95% recall Yes
Wang et al., 2025a Literature review table generation ARXIV2TABLE Table generation for literature reviews Recall, precision, F1 for schema/cell/pairwise overlap Yes
Kusa et al., 2023a Outcome-based evaluation of systematic review automation CLEF TAR 2019 Systematic review automation (retrieval) Mean Average Precision (MAP), Recall at k%, WSS, Area Under the Curve (AUC), outcome difference Yes
Kusa et al., 2023b Automated citation screening CSMeD, CSMeD-FT Citation/full-text screening True Negative Rate at 95% (TNR@95%), normalized Precision at 95% (nP@95%), nDCG@10, MAP, macro-precision/recall/F1 Yes
Scells et al., 2017 Retrieval for systematic reviews SIGIR2017-PICO-Collection Retrieval, screening prioritization Precision-recall, F-beta, WSS, Average Precision (AP), nDCG, MAP Yes
Budau and Ensan, 2024 Automated study search for biomedical systematic literature reviews FASS-BSLR Study search (retrieval, Boolean query generation) Precision, Recall, NDCG, MAP, Recall at 1000 No

Distribution of Literature Review Automation Tasks

Distribution of Evaluation Approaches

Thematic Analysis

Benchmarks for Literature Retrieval and Screening Tasks

Benchmarks in this category focus on identifying, retrieving, and prioritizing relevant literature for systematic reviews or evidence synthesis. Key benchmarks include:

Benchmarks for Literature Summarization and Synthesis Tasks

This theme covers benchmarks designed to evaluate the ability of models to generate summaries, reviews, or structured tables from collections of scientific papers. Notable benchmarks include:

Benchmarks for Structured Literature Analysis Tasks

Structured analysis tasks include risk of bias assessment, data element extraction, compliance and traceability, and meta-evaluation. Key benchmarks:

Evaluation Methodologies Across Benchmarks

Benchmark Category Primary Metrics Evaluation Method Performance Measurement Focus
Retrieval/Screening Recall, precision, Mean Average Precision, normalized Discounted Cumulative Gain, Work Saved over Sampling, Average Precision Automated scoring, comparison to human or baseline models Relevance, workload reduction, ranking effectiveness
Summarization/Synthesis ROUGE, human ratings, semantic coverage, hallucination rate Automated metrics, human evaluation Informativeness, factuality, coherence, citation accuracy
Structured Analysis Macro-F1, accuracy, Cohen’s kappa, composite scores (ECACT) Automated scoring, statistical analysis, human validation Bias, compliance, extraction accuracy, traceability
Meta-evaluation Elo ratings, accuracy, agreement Human preference voting, model-based evaluator assessment Alignment with human judgment, evaluator reliability

Summary of Evaluation Approaches:

References

  1. Jianyou Wang, Weili Cao, Longtian Bao, Youze Zheng, Gil Pasternak, et al. (2024). Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark. arXiv.org
  2. Kun Zhu, Xiaocheng Feng, Xiachong Feng, Yin Wu, Bing Qin. (2023). Hierarchical Catalogue Generation for Literature Review: A Benchmark. Conference on Empirical Methods in Natural Language Processing
  3. Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, et al. (2024). LitSearch: A Retrieval Benchmark for Scientific Literature Search. Conference on Empirical Methods in Natural Language Processing
  4. David W. Brett, Anniek Myatt. (2025). Patience is all you need! An agentic system for performing scientific literature review. arXiv.org
  5. E. Kanoulas, Dan Li, L. Azzopardi, R. Spijker. (2018). CLEF 2017 Technologically Assisted Reviews in Empirical Medicine Overview. Conference and Labs of the Evaluation Forum
  6. Lucy Lu Wang, Jay DeYoung, Byron Wallace. (2022). Overview of MSLR2022: A Shared Task on Multi-document Summarization for Literature Reviews. SDP
  7. E. Kanoulas, Dan Li, Leif Azzopardi, René Spijker. (2019). CLEF 2019 Technology Assisted Reviews in Empirical Medicine Overview. Conference and Labs of the Evaluation Forum
  8. E. Kanoulas, Dan Li, Leif Azzopardi, René Spijker. (2018). CLEF 2018 Technologically Assisted Reviews in Empirical Medicine Overview. Conference and Labs of the Evaluation Forum
  9. Tetsu Kasanishi, Masaru Isonuma, Junichiro Mori, I. Sakata. (2023). SciReviewGen: A Large-scale Dataset for Automatic Literature Review Generation. Annual Meeting of the Association for Computational Linguistics
  10. Brian E. Howard, Jason R. Phillips, Kyle Miller, Arpit Tandon, D. Mav, et al. (2016). SWIFT-Review: a text-mining workbench for systematic review. Systematic Reviews
  11. Weiqi Wang, Jiefu Ou, Yangqiu Song, Benjamin Van Durme, Daniel Khashabi. (2025). Can LLMs Generate Tabular Summaries of Science Papers? Rethinking the Evaluation Protocol. arXiv.org
  12. Wojciech Kusa, G. Zuccon, Petr Knoth, A. Hanbury. (2023). Outcome-based Evaluation of Systematic Review Automation. International Conference on the Theory of Information Retrieval
  13. Wojciech Kusa, Óscar E. Mendoza, Matthias Samwald, Petr Knoth, Allan Hanbury. (2023). CSMeD: Bridging the Dataset Gap in Automated Citation Screening for Systematic Literature Reviews. Neural Information Processing Systems
  14. Harrisen Scells, G. Zuccon, B. Koopman, Anthony J Deacon, L. Azzopardi, et al. (2017). A Test Collection for Evaluating Retrieval of Studies for Inclusion in Systematic Reviews. Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
  15. Leandra Budau, F. Ensan. (2024). Fully Automated Scholarly Search for Biomedical Systematic Literature Reviews. IEEE Access
  16. Yilun Zhao, Kaiyan Zhang, Tiansheng Hu, Sihong Wu, Ronan Le Bras, et al. (2025). SciArena: An Open Evaluation Platform for Foundation Models in Scientific Literature Tasks. arXiv.org
  17. Joshua Morriss, Tod Brindle, Jessica Bah Rösman, Daniel Reibsamen, Andreas Enz. (2024). The Literature Review Network: An Explainable Artificial Intelligence for Systematic Literature Reviews, Meta-analyses, and Method Development. arXiv.org
  18. Christian Jaumann, Andreas Wiedholz, Annemarie Friedrich. (2025). LGAR: Zero-Shot LLM-Guided Neural Ranking for Abstract Screening in Systematic Literature Reviews. Annual Meeting of the Association for Computational Linguistics
  19. Mohammadreza Pourreza, F. Ensan. (2022). Towards semantic-driven boolean query formalization for biomedical systematic literature reviews. Int. J. Medical Informatics
  20. Maxime Gobin, Muriel Gosnat, Seindé Toure, Lina Faik, Joel Belafa, et al. (2025). From data extraction to analysis: a comparative study of ELISE capabilities in scientific literature. Frontiers in Artificial Intelligence
  21. Shubham Agarwal, Gaurav Sahu, Abhay Puri, I. Laradji, K. Dvijotham, et al. (2024). LitLLMs, LLMs for Literature Review: Are we there yet?. Trans. Mach. Learn. Res.
  22. Peter Brown, Ameya Sadguru Kulkarni, Osama Refai, Yaoqi Zhou, Abd Al-Bar Al-Farha. (2019). Large expert-curated database for benchmarking document similarity detection in biomedical literature search. Database J. Biol. Databases Curation
  23. Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess, Yuhui Zhang, et al. (2025). Can Large Language Models Match the Conclusions of Systematic Reviews?. arXiv.org
  24. Xuemei Tang, Xufeng Duan, Zhenguang G. Cai. (2024). Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
  25. Tanja Bekhuis, Eugene Tseytlin, K. Mitchell, Dina Demner-Fushman. (2014). Feature Engineering and a Proposed Decision-Support System for Systematic Reviewers of Medical Evidence. PLoS ONE
  26. Michael D. Skarlinski, Sam Cox, Jon M. Laurent, James D. Braza, Michaela M. Hinks, et al. (2024). Language agents achieve superhuman synthesis of scientific knowledge. arXiv.org
  27. A. Lozano, Scott L. Fleming, Chia-Chun Chiang, Nigam H. Shah. (2023). Clinfo.ai: An Open-Source Retrieval-Augmented Large Language Model System for Answering Medical Questions using Scientific Literature. Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing
  28. Grace E. Lee, Aixin Sun. (2018). Seed-driven Document Ranking for Systematic Reviews in Evidence-Based Medicine. Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
  29. Alon Gorenshtein, Kamel Shihada, Moran Sorka, Dvir Aran, Shahar Shelly. (2025). LITERAS: Biomedical literature review and citation retrieval agents. Comput. Biol. Medicine
  30. Gaelen Adam, Melinda Davies, Jerusha George, Eduardo L Caputo, Ja Mai Htun, et al. (2025). Machine Learning Tools To (Semi-)Automate Evidence Synthesis: A Rapid Review and Evidence Map
  31. Aditya Nagori, Ricardo Accorsi Casonatto, Ayush Gautam, Abhinav Manikantha Sai Cheruvu, Rishikesan Kamaleswaran. (2025). Open-Source Agentic Hybrid RAG Framework for Scientific Literature Review
  32. G. Adam, Melinda Davies, Jerusha George, Eduardo L Caputo, Ja Mai Htun, et al. (2025). Machine Learning Tools To (Semi-) Automate Evidence Synthesis
  33. Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, et al. (2024). OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs. arXiv.org
  34. Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao. (2025). DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents. arXiv.org
  35. Hao Kang, Chenyan Xiong. (2024). ResearchArena: Benchmarking LLMs' Ability to Collect and Organize Information as Research Agents. arXiv.org
  36. Zifeng Wang, Qiao Jin, Jiacheng Lin, Junyi Gao, Jathurshan Pradeepkumar, et al. (2025). TrialPanorama: Database and Benchmark for Systematic Review and Design of Clinical Trials. arXiv.org
  37. Jianyou Wang, Weili Cao, Kaicheng Wang, Xiaoyue Wang, Ashish Dalvi, et al. (2025). EvidenceBench: A Benchmark for Extracting Evidence from Biomedical Papers. arXiv.org
  38. Wanghan Xu, Wenlong Zhang, Fenghua Ling, Ben Fei, Yusong Hu, et al. (2025). Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System. arXiv.org
  39. Jingcheng Du, Dong Wang, Bin Lin, Long He, Liang-chin Huang, et al. (2025). Use of deep learning-based NLP models for full-text data elements extraction for systematic literature review tasks. Scientific Reports
  40. Hao Kang, Chenyan Xiong. (2024). ResearchArena: Benchmarking Large Language Models' Ability to Collect and Organize Information as Research Agents

Report

Status: Gather sources 1000 sources found

Details: Screen sources 40 sources included

Details: Extract data 160 data points extracted

Details: Generate report Save PDF