How we evaluated Elicit Reports - Elicit

How we evaluated Elicit Reports

We recently announced Elicit Reports: fully-automated research overviews for actual researchers, inspired by systematic reviews.

Through external evaluation by researcher specialists, we find that Elicit Reports produce higher quality research overviews and save more time than the other “deep research” tools, including well-known ones like ChatGPT Deep Research, Perplexity Deep Research, and Google Gemini.

Methodology

We recruited 17 professional researchers with backgrounds in fields such as neuroscience, speech therapy, economics, organizational psychology, information studies, ecology, and engineering. The evaluators were PhD holders and were not active Elicit users at recruitment.

We asked these evaluators to compare Elicit to competitors in their area of expertise, asking research questions of their choosing, such as:

They scored and compared reports on both qualitative and quantitative dimensions:

Evaluators compared Elicit against these well-known “deep research” tools:

Statistical analysis

Most results reported below are averages.

Due to the high cost and overhead of obtaining evaluations from skilled professionals, we obtained modest sample sizes of around 25-30 evaluations for each competitor.

We used the Wilcoxon signed-ranks test to determine statistical significance. All p-values were calculated on direct paired comparisons of Elicit vs. a competitor tool using the same research questions, but we grouped evaluations together in the graphics below for ease of reading.

Results

In total, 17 evaluators evaluated 29 Elicit reports, and 120 reports across competitor tools:

Overall quality

We asked evaluators to give a general rating of each report on a scale of 0 to 10, where 0 was defined as “useless, low quality” and 10 as “very useful, high quality.”

Feedback on Elicit Reports

Four Elicit Reports received a perfect grade of 10/10, with feedback such as:

This report was excellent: It spoke to all of the research questions I was hoping to be answered. Particularly impressive was the detailed ‘thematic analysis’ section.

The information was completely on par and displayed in depth data on the various themes I was looking for to answer the question fully.

This is a thorough and systematic examination of the question. In terms of saved labour, the 40 row table comparing studies across characteristics is a huge piece of work. The report is accurate and useful with many interesting insights that I hadn't heard of before. It surfaced some very interesting contemporary scholarship and did an excellent job integrating the findings.

Time savings

We asked evaluators to estimate how many hours the tool would save them when writing a systematic review.

Estimates varied substantially, from 0 to 960 hours, equivalent to 6 months of full-time work at 40 hours/week. Thus the mean estimated time saved is skewed by high values and we report the median instead.

Rating the report’s main answer

We asked: “ How good is the main answer of the report? Look at whatever is the closest to a ‘main answer’” for three metrics: accuracy, usefulness, and how well it answers the research question.

Elicit Reports scored highest across all metrics, though differences are small and generally not statistically significant, especially against ChatGPT Deep Research.

Accuracy and support for top claims

Beyond the main answer, it’s important that each claim in a report is accurate.

Since it is impractical to manually evaluate every claim, we asked evaluators to identify and evaluate the accuracy and citation support of the top 5 most important claims in a report. For example, in a report about evidence-based practices (EBP) in healthcare, an evaluator specified the top 5 claims as:

  1. “Healthcare organizations with supportive leadership, openness to change, and adequate resources reported better EBP implementation.”
  2. “Studies frequently mentioned time limitations, resource constraints, and weak research culture as challenges.”
  3. “Regarding technology, studies found that user-friendly information systems and digital resource access supported EBP implementation.”
  4. “Common individual barriers included limited literature search skills, skepticism about EBP, and uncertainty about professional roles in implementation.”
  5. “Practitioners with confidence in their abilities and positive views toward EBP were more likely to implement it.”

The evaluators then graded the accuracy of each claim as follows:

We also asked them to evaluate citation support, giving one of the following scores:

The scores were then summed to yield a score out of 5 for each report.

We found that several tools performed statistically equivalently as Elicit.

Qualitative impressions

Evaluators shared qualitative feedback in addition to quantitative scores, highlighted below.

Elicit uses quality academic sources

Evaluators noted that Elicit relies on trustworthy academic sources, unlike competitors.

Elicit doesn't include news articles, interviews, blog posts, SEO content farms, and other non-scholarly content.

In Elicit, all information included in a report is derived from a credible, peer-reviewed original publication.

Elicit makes it easy to check citations

Though evaluators did not rate Elicit's citation support as clearly superior, they did appreciate that Elicit cites specific passages in citations.

Elicit's methods are more transparent and flexible

Evaluators liked that they could understand and modify how Elicit finds, screens, and extracts data.

Elicit generates better tables

Most other research tools do not summarize data in tables that are easy to digest.

Limitations

Besides small sample size, our analysis is subject to the following limitations:

Further evaluation work

Elicit Reports use Elicit Systematic Reviews under the hood, including separate paper search, screening, extraction, and report-writing steps. We evaluated these steps individually, and will share those results soon.