Evaluation Engineer - Elicit

Evaluation Engineer

Department

Engineering

Location

Oakland, CA (or remote within US timezones)

Type

FullTime

Description

Application

About Elicit

Elicit is an AI research platform that uses language models to help researchers figure out what's true and make better decisions, starting with common research tasks like literature review.

What we're aiming for:

  1. Elicit radically increases the amount of good reasoning in the world.
    • For experts, Elicit pushes the frontier forward.
    • For non-experts, Elicit makes good reasoning more affordable. People who don't have the tools, expertise, time, or mental energy to make well-reasoned decisions on their own can do so with Elicit.
  2. Elicit is a scalable ML system based on human-understandable task decompositions, with supervision of process, not outcomes. This expands our collective understanding of safe AGI architectures.

The mission of Elicit evals

At Elicit, we're after something different. We want to understand, and hill-climb toward, models that help us make better decisions.

This is harder than "what will users like better." Decision support is difficult to evaluate, and users' knee-jerk reactions don't always track with what actually helps them decide. Because it's hard, and because the sales pitch is more complicated, few are doing it well. If we get this right, we have a real shot at pushing AI toward better decision-making, both inside Elicit and beyond.

Why we're hiring for this role

We need someone to own the technical foundation of our auto-evaluation systems. Our evals are much slower than they need to be, and our interfaces aren't built for the range of people who rely on them: ML engineers iterating on models, product managers monitoring quality, and customers assessing how much to trust a result.

This role goes beyond building infrastructure. You'll work out what it actually means for Elicit to support decision-making in pharma, and encode that understanding into our evaluation systems.

What you'll own

The core auto-eval platform

You'll build a comprehensive system that runs fast, is easy to use, and supports quickly building new evals:

Ensuring evaluations are accurate and reliable

A month in the role

In a typical month, expect to spend:

What you bring to the role

Requirements

Will make you more competitive for the role

This is a diverse list of nice-to-haves. We expect the candidate we select to have some, but not all, of these. Other team members can fill in for skills you lack.

Location and travel

We have a great office in Oakland, CA, and we'd love to see you there if you're local. That said, we're just as happy for you to work remotely. We do get the whole team together for a quarterly retreat somewhere fun, because in-person time matters to us.

Benefits and perks

In addition to working on important problems as part of a productive and positive team, we also offer great benefits (with some variation based on location):

Compensation

For all roles at Elicit, we use a data-backed compensation framework to keep salaries market-competitive, equitable, and simple to understand. For this role, we target starting ranges of:

Apply now