Discovering Alignment Windfalls Reduces AI Risk - Elicit

Discovering Alignment Windfalls Reduces AI Risk

My argument, in short:

  1. Just as there are alignment taxes, there are alignment windfalls.
  2. AI companies optimise within their known landscape of alignment taxes & windfalls.
  3. We can change what AI companies do by:
    1. Shaping the landscape of taxes and windfalls
    2. Shaping their knowledge of that landscape
  4. By discovering and advocating for alignment windfalls, we reduce AI risk overall because it becomes easier for companies to adopt more alignable approaches.

Alignment taxes

An “alignment tax” refers to the reduced performance, increased expense, or elongated timeline required to develop and deploy an aligned system compared to a merely useful one.

More specifically, let’s say an alignment tax is an investment that a company expects to help with alignment of transformative AI that has a net negative impact on the company's bottom line over the next 3-12 months.

A few examples, from most to least concrete:

All of these require more investment than the less aligned baseline comparison, and companies will face hard decisions about which to pursue.

Alignment windfalls

On the other hand, there are some ideas and businesses where progress on AI safety is intrinsically linked to value creation.

More specifically, let’s say an alignment windfall is an investment that a company expects to help with alignment of transformative AI that also has a net positive impact on the company's bottom line over the next 3-12 months.

For example:

In practice, almost all ideas will have some costs and some benefits: finding ways to shape the economic environment so that they look more like windfalls is key to getting them implemented.

Companies as optimisers

Startup companies are among the best machines we've invented to create economic value through technological efficiency.

Two drivers behind why startups create such an outsized economic impact are:

  1. Lots of shots on goal. The vast majority of startups fail: perhaps 90% die completely and only 1.5% get to a solid outcome. As a sector, startups take a scattergun approach: each individual company is likely doomed, but the outsized upside for the lucky few means that many optimists are still willing to give it a go.
  2. Risk-taking behaviour. Startups thrive in legal and normative grey areas, where larger companies are constrained by their brand reputation, partnerships, or lack of appetite for regulatory risk.

This optimisation pressure will be especially strong for artificial intelligence, because the upside for organisations leading the AGI race is gigantic.

Shaping the landscape

The technical approaches which lead to taxes and windfalls lie on a landscape that can be shaped in a few ways:

Regulation can levy taxes on unaligned approaches:

Public awareness can cause windfalls for aligned approaches:

Recruiting top talent is easier for safety-oriented companies:

Companies greedily optimise within the known landscape

An important nuance with the above model is that companies don’t optimise within the true landscape: they optimise within the landscape they can access.

Here are a couple of reasons why the full landscape tends to be poorly known to AI startup founders:

In contrast, researchers in academia have much more latitude to explore completely unproven ideas lacking any clear path to practical application.

Shaping knowledge of the landscape

What does it look like to shape the broader knowledge of this landscape?

Factored cognition

An example of an alignment windfall

Let’s consider a more detailed example. Elicit has been exploring one part of the landscape of taxes & windfalls, with a focus on factored cognition.

Factored cognition as a windfall

Since the deep learning revolution, most progress on AI capability has been due to a combination of:

Normally, we do all three at the same time. We have basically thrown more and more raw material at the models, then poked them with RLHF until it seems sufficiently difficult to get them to be obviously dangerous. This is an inherently fragile scheme, and there are strong incentives to cut corners.

Factored cognition offers a different path. Instead of solving harder problems with bigger and bigger models, we decompose the problem into smaller, more tractable problems. Each of these smaller problems is solved independently and their solutions combined to produce a final result.

How we’ve been exploring factored cognition at Elicit

Elicit, our AI research assistant, is built using factored cognition: we decompose common research tasks into a sequence of steps, using gold standard processes like systematic reviews as a guide.

For Elicit, creating a valuable product is the same thing as building a truthful, transparent system. Trustworthiness is our value proposition.

Conclusion:

Let's find and promote alignment windfalls!

I'm a proponent of other approaches—such as regulation—to guide us towards safe AI, but in high stakes situations like this my mind turns to the Swiss cheese model used to reduce clinical accidents.

In my view, Elicit is the best example of an alignment windfall that we have today. To have maximum impact, we need to show that factored cognition is a powerful approach for building high-stakes ML systems.