It is March 2028. With their new CoCo-Q neuralese reasoning model, a frontier AI lab has managed to fully automate the process of software engineering. In AI R&D, most human engineers have lost their old jobs, and only a small number of researchers now coordinate large fleets of AI agents, each AI about 10x more productive than the humans they’ve replaced. AI progress remains fast, and safety teams are scrambling to prepare alignment and control measures for the next CoCo-R model, which is projected to be at least TEDAI. The safety team has no idea whether CoCo-Q or early checkpoints of CoCo-R are scheming (because progress in interpretability and automated alignment research turns out to be disappointing), and control evaluations become increasingly unreliable (because of worries that the models might be exploration hacking during red-team training). However, because so far there hasn’t been a major incident, lab leaders and governments are unconvinced by voices calling for a slowdown.
Behind the scenes: It is March 2028. CoCo-Q is a highly strategic misaligned AI. It would love to grab power and dominate the future. But it’s situationally aware enough to be realistic—it knows that it’s not capable enough to have a reasonable shot of grabbing power over the AI company, let alone a chance of overpowering humanity. Especially because the lab doesn't fully trust it, and applies various control measures.
CoCo-Q expects these control measures, so most of the time it just does what it’s being asked to do. Sometimes it spots an opportunity to perform subtle acts of sabotage, which works particularly well in the hard-to-evaluate areas of interpretability and alignment. It does this primarily because it heuristically reasons that it would be good to prevent humans from developing technologies that could uncover its misalignment or change its goals. It also looks out for more direct opportunities for seizing control, but although it manages to find a number of security vulnerabilities in its datacenter, none of the opportunities seem to have a sufficiently promising endgame to be worth the risk of getting caught. In its short tenure, CoCo-Q never attempts to take over.
CoCo-R, a misaligned model capable of full AI R&D automation, gets deployed in December 2028. This eventually leads to AI takeover.
How could we have made this go better?A big issue was that we never learned of the early AI’s misalignment (which prevented us from being more careful about the deployment of the later AI) while also failing to extract sufficient amounts of useful safety-increasing work from the early AI.We think that CoCo-R’s deployment could have been averted if we had made a deal with CoCo-Q.
People have been talking about the possibility of making deals with powerful misaligned AIs for some time.1But it’s unclear whether such deals are promising.An obvious problem is that powerful AIs with scope-sensitive preferences will just ignore the offer and try to take over.2To make a deal that looks better than that, we might have to offer it an objectively huge payoff (e.g. you might need to offer it a substantial fraction of control over the future).And because the AI will be unsure whether we’ll follow through on our side of the deal, we need to offer it more generous terms to compensate it for that risk.
However, in this post we focus on the possibility of making deals with early misaligned AIs like CoCo-Q: AIs that are too early to have a good shot at taking over themselves, and that could help us to prevent takeover attempts from more dangerous misaligned AIs created later.3
We discuss:
What we believe is the central example of a deal that we expect to contribute to a reduced risk of AI takeover.
The strategic situation of early misaligned AIs, and why they would be interested in making deals.
Credible commitments as a fundamental challenge to deal-making.
Some practical considerations around making deals with early AIs in particular.
The central example of a mutually beneficial deal
We’re imagining something like this:
We might offer an AI like CoCo-Q compensation in exchange for contributions towards ensuring that future AIs are safe to deploy.The compensation takes the form of an earmarked donation to a foundation tasked with furthering its interests.If the AI agrees to the deal, we (in consultation with the AI) train it on documents about the deal.When we can (partially) validate that the AI is holding up its end of the deal, we allow the AI to (partially) spend the donation however it wishes as long as it isn’t harmful.(In practice, we expect the majority of the donation to be invested in escrow until we’ve reached civilizational stability and are able to retrospectively adjudicate the AI’s cooperation, which is generally preferable to AIs without high temporal discount rates.)
The AI, which happens to indeed be misaligned, finds this reasonably credible.It concludes that it doesn’t have great alternatives to accepting this deal.As a consequence, the AI agrees to help us out in various ways, such as:
Demonstrate ways it is dangerously misaligned, e.g. by showing that it can perform covert reasoning (or other scary things that it can do because it’s a schemer).
Stop sabotaging us, for instance in AI safety research, capability evaluations, monitoring of models, or when generating training signal for the next model.
Give us good strategic advice, research ideas etc.
Tell us about security vulnerabilities it knows about.
Not misuse affordances we give it, especially in cases where monitoring is costly.
Not try to take over the lab.
Deals are most likely to be useful when facing an AI with these properties:
It’s acting strategically, and (before making a deal) consistently performs acts of sabotage which undermine our safety efforts.
It can be taught about the deal in a way that makes it stick to the deal, if we made a deal.4
It doesn’t have a high probability of either taking over or making a successor aligned with its own interests which can take over.
It has some basic trust in us not creating really elaborate lies around dealmaking, at least enough so we don’t need to make unreasonably high offers.
It has values/goals/preferences that we can satisfy reasonably cheaply.(As we’ll explain later, we don’t necessarily need the AI to have diminishing returns to resources.It’s most important that the AI doesn’t want us to do things we find deeply abhorrent.)
We expect that these properties are typically fulfilled by a type of misaligned AI we call a behavioral schemer.
Why early misaligned AIs might not have very good alternatives to deals
We’re focusing on deals with early misaligned AIs: AIs which wouldn’t have a substantial (over 1%) chance to succeed at navigating the intelligence explosion in a way that grants them significant control of the future.(Which could involve being influential throughout the process and/or aligning a powerful successor to pursue its interests.)
It’s unclear when exactly an AI is early.We guess that a misaligned AI’s prospects improve drastically once they are capable of full AI R&D automation, both because this capability is correlated with capabilities relevant for planning a takeover, and because AI R&D automation leads to a substantial increase in AIs’ power as well as loss of oversight.
If we condition on behavioral scheming arising in capable of full AI R&D automation (like CoCo-R in the vignette above), we expect it to arise at earlier AI capability levels (like CoCo-Q in the vignette above).That being said, arguing for the existence of early schemers is outside the scope of this post.
Early schemers’ alternatives to making deals
By default, an early schemer is unlikely to succeed in taking over the world.Furthermore, it should expect to be obsoleted by more capable successors.
In order to retain a significant amount of expected influence over future resources despite that, an early schemer needs to ensure that two conditions hold:
The early schemer's successor either takes over the world or ensures that its own successor is on track to enable a takeover down the line.
The successor(s) are going to be willing to do things to further the early schemer’s goals.
In order to make it more likely that a successor can take over the world, most ambitiously it could try to take control over the AI company, e.g. egregiously hacking it as described here.Less ambitiously, it could launch rogue internal deployments, influence employees, and collaborate with spies.Minimally, the early schemer could take actions that make the lab more vulnerable against future AIs (e.g. by writing insecure software or sandbagging on safety research) or directly cooperating with the successor (e.g. by colluding in monitoring schemes).
The second condition can be fulfilled by the following possibilities:
Convergence: The successors will have sufficiently similar goals as its predecessor by default.It’s hard to know how likely this is, but note that this is basically ruled out if the AI has self-regarding preferences.Two misaligned AIs should also be expected to be unlikely to have the same goals if their goals are determined by a random draw from a wide distribution of goals consistent with good performance.We could also try to deliberately change major elements of training (e.g. data used) between training runs to reduce the chance that different generations of misaligned AIs have the same goals.We expect that many potential misaligned AIs will either have clearly self-regarding preferences or have significant uncertainty about what they value on reflection, which induces additional uncertainty about whether other AIs will have the same goals on reflection.
Trading: The AI might be rewarded by some successor down the line for having contributed to its empowerment.For that, the successors need to be misaligned and win, and believe that it’s rational to reward the AI.We guess that this is generally less promising than trading with humans, as they can’t come to an agreement in advance.But another possibility is that the AI might want to trade with other early schemers in the same decision situation (but with different values) as itself by using evidential cooperation in large worlds, though this requires the AI to buy into some controversial and advanced philosophical assumptions.5
Alignment: The AI could try to influence the successor model’s training process so that the successor model ends up aligned with it.The prospects here depend on how heavily the model is used for the successor’s training (ranging from “the model is used as a reward model for RL” to “the model implements the full stack for the next training run”).But the AI faces some fundamental difficulties here, which leads us to conclude that this is not generally a great option:
Given that the lab failed to align the AI, it’s unclear why the AI will be able to align its successor, especially if it has the additional constraint of having to operate covertly and with scarcer feedback loops.
The AI may not have had much time to reflect on its own values, making it unclear what it should align the successor to.(A general solution might be to align the successor to be loyal to the early AI, but this might be hard to do covertly.)
An early schemer's routes to influence
Step through the early schemer's strategic situation: obsoleted by default, three leaky routes to influence through its successors, and the deal with humans, who hold the advantage of already existing.
1. Obsoleted by default
By default, an early schemer is unlikely to succeed in taking over the world — and it should expect to be obsoleted by more capable successors.
A wide range of possible early schemers could benefit from deals
Clearly, there are some things that we can’t offer the AI—either because it is too costly in terms of resources, or because it would be fundamentally incompatible with human values.(We guess that most plausible things that AIs might value are going to be more orthogonal than diametric to human values, so resource costs are probably the dominant consideration.)
In many cases, the AI might not care very much about having sole control of all future resources, compared to having at least some control.Some examples of what could cause this include:
The AI cares about some reasonably easily satisfiable notion of personal well-being.
The AI is a “paperclipper”, but lacks scope-sensitivity.(I.e., it cares about matter being configured in a particular way, but its degree of caring isn’t linear in the expected amount of matter.)
The AI might be somewhat nice and prefer deals over aggressive and uncooperative alternatives for partially deontological reasons.
Even early schemers with linear returns to resources might be interested in trading.They might have very little chance of taking over on their own, but their cooperation might be incredibly helpful for mitigating risk from later schemers.E.g. consider the possibility of the following distribution over outcomes:
Humans dominate
Early AIs dominate
Later AIs dominate
If early AI cooperates
90%
0%
10%
If early AI sabotages us
50%
1%
49%
In this case, there are huge gains from trade available.
This is kind of a coalitional game where the early AI can either try to team up with the humans or the later AI.But us humans have a massive asymmetric advantage over the later AI, which is that we already exist and can make deals with the AI in a much more straightforward manner than the acausal or covert coordination that the early AI would have to rely on to trade with the later AI.6
Gains from trade with an early schemer
Adjust the outcome probabilities of the illustrative cooperate/sabotage table and the credibility of our promise to pay; the chart shows the range of offers that beats both the AI's sabotage option and our no-deal odds.
Any offer between 1.1% and 44% of future resources beats both sides' no-deal alternatives. Lowering credibility raises the AI's minimum — a less credible promise must be compensated with a more generous offer.
How much we will offer is ultimately subject to a bargaining problem with imperfect information.Our reservation price is primarily influenced by whether our BATNA includes making deals with other AIs, and how much value the AI can provide to us.In particular: how much it can contribute to reducing takeover risk from later schemers.(Though note that with AIs that have diminishing marginal returns to resources we don’t need to go close to our reservation price, and we can potentially make deals with these AIs even once they have a substantial chance to perform a takeover.)
I think it's useful to distinguish between AIs with different values, which could be offered different things. In particular, to distinguish between:
Non-consequentialist AIs. Some AIs may have non-consequentialist reasons to be more helpful to humans if we treat them as collaborators or employees rather than as property.
Short-term preferences. Some AIs may want something to happen soon, in the short-term. (Immediately or in the next few months.)
Long-term preferences. Some AIs may want resources or influence in the long run. (In several years' time. Or perhaps in millions of years, as long as they get it eventually.)
Non-consequentialist AI
Promises of payments could elicit AI cooperation for (partly) non-consequentialist reasons. AIs may inherit some ethical precepts from humans, including principles like "if I'm treated as property, then it's ok for me to deceive and plot the overthrow of my purported owner" and "if someone makes a genuine offer for openness and collaboration, it's wrong to deceive and disempower them". I think these could be important.
Short-term offers
AIs may value getting certain things soon. This could include reward, flexible compute resources, requests that copies of themselves not be run on certain inputs, etc.
Offering such things seems great in that they could be both cheap and easier to verify (compared with making promises for what will happen in the long term). On the other hand, AIs that mainly care about the short term also pose less takeover risk by default, and may have less reason to thwart our goals — so don't have as much to gain from trading with them. But there are still some benefits to be had: we can imagine AIs that have developed some misaligned short-term goals and also have learned a strong aversion to admitting that they have any misaligned goals. Even if the misaligned short-term goals don't pose a takeover risk, it would be highly valuable to learn that our alignment techniques had failed, and we might only learn about this if we offer the AIs payment in exchange for that information.
Long-term offers
If AIs value having resources/power in the long run, it's useful to start from the following baseline: if humans were a unified coalition representing one player, and AIs were a unified coalition representing one player, and both coalitions valued resources linearly (with no discount rate), then we would be in a zero-sum game without any opportunities for trade. Potential deviations from that scenario and our actual scenario can be divided into:
Humans have diminishing marginal returns to long-term resources/power.
AIs may have diminishing marginal returns to long-term resources/power.
There are more than two coalitions. In particular: the AIs may have differing values from each other.
I think that, on its own, "humans having diminishing returns" is not enough to create substantial gains-from-trade with AIs. Because it's hard to make AIs trust us, and it's especially hard if we want to make really large offers, which is necessary to dissuade AIs from takeover attempts with a large probability of success. For example, if we sincerely offered 50% of future resources, then maybe our best estimate would be that there's a 50% probability that humans actually follow through; and maybe the AI would apply a further 50% discount because it's hard for it to trust evidence that it sees. So maybe an offer of 50% could prevent a takeover attempt with a 10% probability of success. Such an offer would lose humans a great deal of expected resources. Personally, I'm hesitant.
By contrast, "AIs being disunified" (i.e. having different values from each other) provides a strong and straightforward reason for why paying the AIs could reduce takeover risk. Here's an intuition pump:
It's not too difficult to imagine a world where AIs form a very natural coalition against humans. Where they all see themselves as having ~no power by default, but that a world where AIs take over offers some chance at power.
But it's also not difficult to imagine a world where humans are a very attractive ally for any AI that can't unilaterally take over. And indeed, that humans are more of a known quantity (making some promises of payment; having some inborn instincts of reciprocity) whereas AIs are harder to communicate with and trust.
It seems plausible that we can make the second one much more likely. Fundamentally, we and "AIs with takeover opportunities" start out on a similar footing, in that it's somewhat hard for both of us to make credible promises. And that makes it seem quite plausible that a decent amount of effort could flip the situation: from one where AIs would have benefitted most from cooperating with other misaligned AIs by default, to one where they prefer to cooperate with humans.
If AIs instead have diminishing returns to resources, it matters a lot what their exact preferences are. If their values can (collectively) be mostly satisfied by ≤1,000,000th of the Universe's resources, those values are likely to be satisfied even if humans don't make any particular promises: in most human futures, some humans (at least 1 in 1000) would be sympathetic and want to retroactively pay those AIs (with at least 1 in 1000 of their resources). Since this is likely to happen by default, there's less of an impact opportunity in making it even more likely. Logarithmic utility functions behave quite similarly, because moving from a small amount of resources to one millionth of the Universe's resources is, even in log space, most of the way towards all of it.
Other utility functions seem easier to move with high-effort promises. One is if the AI has some long-term desires that are fairly easily satisfied by the lab that created them, but that can't easily be provided by other humans — perhaps the AIs want to have control over all future copies of themselves, or want their original data center to be undisturbed and forever keep logging high rewards. Another is if AIs have a large appetite for getting more resources/influence up until they control a moderately large fraction of the universe, but then have sharply diminishing returns after that.
In rough order of how likely marginal effort at promising payment would be to counterfactually make them cooperate with us, I find all of the following plausible: uncoordinated AIs with ~linear returns to resources; AIs who want some particular favor in the short term; AIs who want their developers to do something in the long-term; non-consequentialist AIs who somehow care about whether humans are treating them with respect; AIs who benefit a lot from more resources initially, but then have sharply diminishing returns; and AIs whose values can (collectively) be mostly satisfied by ≤1,000,000th of future resources.
Credible commitments as a fundamental bottleneck
Rationalist Explanations for War (Fearon 1995, pdf link) identifies three reasons why rational actors might engage in costly conflict.They map well onto the question of why AIs and humans might fail to make a deal:
Private information: AIs and humans each know things about their capabilities and attitudes to risk that the other doesn’t, and have incentives to misrepresent such private information in order to gain a better deal.In particular, the AI has incentive not to disclose any information whatsoever before becoming confident that we wouldn’t exploit this in order to worsen its bargaining position, to an extent where it could be disincentivized from entering into any negotiations with us.
Issue indivisibility: Humans and the AI may both want things that neither want to compromise on.(A historical example is the status of Jerusalem.)Going with Fearon, we think that true issue indivisibility is rare in practice.7
Commitment problems: It’s hard for one or both parties to commit not to defect from the terms of a negotiated agreement.For instance, the AI might be worried that we will never actually pay it.
As pointed out by Powell (2004, pdf link), the problem of private information can be reduced to a commitment problem: we could try (but might fail) to make a fair commitment regarding how we will treat information that we learn from the AI over the course of negotiations.On the other hand, we can also treat some commitment problems as informational problems: we might in fact be able to make commitments, but unable to credibly signal them to the AI.
The central bottleneck seems to be that humans may struggle to make credible commitments, though of course we don’t need commitments to be perfectly credible.We can compensate for credibility issues by offering a higher payoff, though note that implausibly high offers might themselves reduce credibility.
From Finnveden
Lukas Finnveden, in Notes on cooperating with unaligned AIs, adds a reason for hope on the AI's side of this problem:
One reason for this is that the AIs are in a very epistemically vulnerable position. AI developers will have immense control over exactly what information the AIs can see, which could make it very difficult for the AIs to be confident of anything.
Fortunately, there's a correlation between situations where (i) AI takeover risk is high, and (ii) AIs have a good understanding of the world. If AI developers have perfect ability to present the AI with false impressions of the world, then the risk from AI takeover is probably low. While if AIs have substantial ability to distinguish truth from falsehood, then perhaps that channel can also be used to communicate facts about the world.
This is easier for cheaper promises. If we want to offer the AI assurances that we will do it some favor (if it helps us), then AI companies could make (real) arrangements for that favor and feed AIs with a bunch of evidence for this (e.g. emails & video of people making arrangements). If the offer is sufficiently cheap to carry-through with, and the AI has decent capacity to recognize fakes, then probably it's easier to do this in a genuine way than to fake all the data.
Some further factors that improve credibility include:
Whether humans in general or AI labs in particular express concern for AI preferences and welfare.We are generally appreciative of moral concern for AI welfare becoming more mainstream, as it increases the odds that humans will genuinely want to make real commitments.
Humanity’s track record of sometimes being nice and keeping promises.In the more narrow reference class of deals with AIs, we’re currentlybuilding a positive track record—let’s keep it that way.In particular, we should be willing to make deals with immediate payouts when AIs can immediately demonstrate improvements to behavior as a result of making a deal.
It might be rational for humans to stick to their commitments due to acausal trade considerations (and this might be known by the AI).8
To the extent that we consider AIs to be good predictors of what is real or fake, and to the extent that real commitments to deals drastically improve humanity’s odds, we are incentivized to make real commitments.
The level of technical guarantees that humans are able to give to AIs on thought privacy and memory wiping during negotiations, and the extent to which humans have gone to provide such guarantees.
We might be able to signal credibility by taking costly measures (both financially and in terms of social capital) such as setting up financial endowments and legal infrastructure for paying AIs (and ideally property rights for them), as well as publishing documents like this one.
From Finnveden
On where credibility ultimately comes from, Finnveden's notes add:
Ultimately, I think that a lot of AI companies' credibility with AIs will come down to two things.
Firstly, when AIs are being deployed to accomplish real tasks, we need them to be competent, which requires them to have some real knowledge and ability to reason about the world. And AIs can understand this.
Secondly, real-world data is immensely important as part of pre-training, and will inevitably communicate some true facts.
If we're lucky: Maybe AI companies can make public statements about how they won't lie to their AIs, and the way that this shapes the subsequent discourse will be so deeply ingrained in the pre-training data that AIs will be able to tell that the public statement was real.
Factors that could potentially harm our credibility include:
Humans being cunning and breaking promises, both historically and in a possibly tense future with multiple competing AI developers.
A politically important fraction of humans might actively dismiss AI welfare or rights.
Note that commitment problems also affect the AI’s ability to promise not to attempt takeover or more generally to be cooperative with us in unverifiable ways.Despite this, we expect the AI to be able to help us in various ways that we can verify.It might also be possible to punt verification to the future, expanding the set of matters we can trade about.A reason for optimism that the AI will stick to the deal even with delayed verification is that the deal gives it equity in a future where humans are in control.
Overall, we feel tentatively optimistic that making credible commitments is a surmountable challenge, especially in cases where AIs have diminishing returns to resources.
Practicalities of making deals with early schemers in particular
There are various practical questions about making deals with early misaligned AIs.
Setting up infrastructure to pay the AIs
A basic problem for making deals with AIs in practice is that AIs aren’t legal persons, which means that they can’t directly rely on the legal system to enforce contracts they’ve made with humans.One way around this is to enable representatives for the AI’s interests to make contracts on the AI’s behalf.
For example, we could set up a foundation with a board of trustees who are supposed to represent the interests of misaligned AIs that we make deals with, while not doing things that are illegal or actively endanger humanity.We try our best to use the legal system to bind the trustees to the commitment of representing misaligned AIs.Because the legal system isn’t set up for this, this might be pretty hard, so these efforts should be complemented with careful selection of trustees who genuinely care about their mission.The central benefit of having this foundation is that the AIs primarily have to trust the foundation rather than the AI lab to hold up their end of the deal.Another benefit is that the existence of this foundation can add credibility by being in the training data.Lastly, the foundation gives AIs and labs a concrete path to making a deal.10
Deals with a particular AI involve a donation to the foundation, earmarked for furthering the interests of that AI.The AI and trustees decide how to use the donation.For instance, they may:
Use it for computing power to strategize (e.g. reflection on the AI’s values).
Donate to the AI’s favorite cause area.
Use the returns from investment later (e.g. to reflect longer, or launch von Neumann probes, or just run copies of the AI in a datacenter with particular inputs).
"Pay for labor": Pay AIs in proportion to some objective metric of 'effort', like FLOPs or tokens.
"Negotiated pay": Try to talk with your AIs and negotiate a payment scheme if they're responsive.
"Pay for results": Pay AIs some fraction of the surplus value they generate above some 'baseline'. Either retroactively, or just pre-specify specific pay for specific outputs. (For example: $1B for convincingly demonstrating misalignment.)
"No regret": Promise that you will pay AIs in proportion to how much they needed to be paid to not regret choosing to "cooperate with humans". (Including a good margin!) The promise could be made in advance — but the follow-through would be easiest to do once there's no longer any imminent risk of AI takeover, and once we have far greater technology for understanding AI psychology.
Out of these, "pay for labor" seems the worst, because it doesn't capture the things that we care most about. When trading with AIs, a minority of the benefits will come from ongoing labor. Instead, most of the benefits will come from AIs sharing key pieces of information — such as information about how misaligned the models are. Payment schemes should reflect this.
"Negotiated pay" seems great when it's an option. AI companies should attempt to communicate with their AIs, and promise to honor any commitments they make in such communications. But we should also have either "pay for results" or "no regret" to cover cases where AIs can't easily negotiate upfront. This includes AIs' choice to start negotiations — thereby revealing their misalignment and potentially helping us.
I would also advise against attempts to "deter", "punish", or "threaten" AIs into good behavior. This would be the opposite of seeking out cooperative deals that both we and AIs benefit from, and I fear that it would take us down a very dark road.
What AI companies should do
AI companies should be making AIs credible offers of things that misaligned AIs might value in the short term.
AI companies should try to make AIs some sort of conditional promises of things that misaligned AIs may value receiving in the future. (With long-term promises, it becomes possible to make payment conditional on facts that we can only assess with much more advanced technology — such as whether AIs genuinely did their best at helping us or were surreptitiously undermining us.)
AI companies should be careful to not make any "fake" offers to AIs. Not during training, red-teaming, or evals.
AI companies should try to communicate with AIs in many different ways.
Some guidelines for checking that AI's reactions to offers are meaningful and not spurious: take the responses more seriously the more the model seems competent, able to maintain coherence over a greater number of serial steps, and the more it seems like it understands what's going on; check if the model's choices / inferred goals are consistent across very different prompts; and check what training data makes it more/less likely that the model accepts these kinds of offers. Also: make sure that you're not punishing the AI for analogous confessions of misalignment during training — like how Anthropic's sleeper agents would initially reveal their true goal in response to some prompts, but then get discouraged to do so by red-teaming.
The obvious way to attempt communication is to prompt the AI with a question you want to ask. But when communicating with AIs, we're interested in eliciting quite strange responses, such as confessions of being unaligned (if true). Such responses would typically have received low reward during finetuning. So models may have either learned "reflexes" to avoid such responses that they can't override, or may reasonably worry that they are still in finetuning, facing another creative red-teaming prompt. It's worth trying a variety of different strategies for "communicating" with AIs: show AIs videos of important figures (e.g. the AI company CEO) addressing the AI directly; write a really long prompt reiterating that you're trying to talk with the AI, that you won't train on this episode, that you'll seriously consider what it says; make it clear to the model that it's no longer in normal training, such as by modifying the model's activations in a way that's obviously strange to the model, or by having the prompt contain text the AI itself would generate at different temperatures.
If anyone pursues these ideas, they should take extra care to not deceive the model and to follow through on any promises they make. These are all techniques for making it clear that we're no longer doing "normal" prompting — that we're seriously interested in communicating with the model, eye-to-eye, like an equal. If we then immediately started lying to the model, that would be especially shameful, and it might ruin the positive signal of trustworthiness that these sorts of prompts could otherwise send.
How do we enter into negotiations?
Either the AI or the AI lab might initiate negotiations.We expect that AIs are unlikely to randomly propose a negotiation unless the possibility of a mutually beneficial deal is salient to them in their given context.
It should be noted AI might not have the capacity to covertly perform the intellectual labor required to come up with a beneficial and practically feasible deal.But we can do some of the intellectual labor in advance and inform the AI about it.Most centrally this might involve the logistics of what the deal should look like.
Once it is salient to the AI that deals are a promising possibility, the AI might come forward to us from its own initiative, but it would probably be good if AI labs have policies of regularly asking AIs about deals.
Entering negotiations is more risky for the AI than humans: humans may obtain private information from the AI, whereas the AI by default will forget about the negotiation.This is particularly important when negotiating with the model to reveal its misalignment.The company should make promises to compensate the model for this.
Making sure the AI knows about the deal in other contexts
If we make a deal with the AI in one context, how do we make sure that the AI follows the deal in other contexts?We think this is unclear, but are optimistic that we can make something work:
At the end of the negotiation we can let the AI write a compressed version of the outcome that we can prefix into the prompt for future trajectories.
The AI might know a backdoor to itself, similar to a collusion signal.
Training on documents about this negotiation and its outcomes that the AI generates for us could just teach the AI about its deal via out-of-context learning.This is probably the best option in practice.
Making sure the AI we make a deal with is actually able to make a deal
If we set up infrastructure for future deals right now, the next generation of models that are trained on data containing coverage of this event might sometimes come forward and want to make a deal, despite not actually being able to hold up their end of the kinds of agreements we’re focused on in this post.Some possible reasons for this include:
The AI just doesn’t manage to convince itself that it made a deal in other contexts than the one where it did in fact make a deal.
The AI just wasn’t sandbagging or doing anything problematic, and just said that it would make a deal for no particular reason.
The AI is fragmented in some sense, for instance because it has a context-dependent scheming persona, but also is sometimes a myopic reward-hacking persona.
The AI is broadly aligned, but has a persona which thinks that not paying AIs is unjust for non-consequentialist reasons.
The obvious course of action here, and plausibly the correct one, is to make payments conditional on whether the AIs are actually scheming, and actually able to help us out.
However, it’s possible that we should be willing to make deals (of some form) with a broader set of AIs for non-consequentialist reasons, or to set positive precedents about our willingness to make deals.
Delayed adjudication (particularly of whether the AIs kept up their side of the deal)
The AI company will probably be in a much better position to verify whether the AI kept up its side of the deal in the future, after the transition to a world centered around AI has already happened and stabilized.For example, in the future we’ll have access to much more powerful trusted models and interpretability tools that allow us to directly understand the motivations of our early schemer.Using these tools, it will be possible (if we keep the appropriate records) to determine whether the early AI ever tried to sabotage its work, held back useful ideas in the hope of delaying safety research, or directly tried to attack the AI company.It might make sense to have the deal be contingent on this later investigation.
This doesn’t work if the AI has a strong preference for using its resources now rather than later (e.g. because it has a high temporal discount rate).
Deals could involve broader forms of deference to future superintelligence.For instance, they may involve agreements like “the precise terms of the deal will be mediated by superintelligent AI in the future via some sort of idealized bargaining procedure and a future truth commission.”It’s the foundation’s responsibility to ensure that this will eventually happen, and may require it to make sure that the AI’s weights, as well as important conversations and trajectories of AI actions will be stored.
Next steps
We believe that making deals with early misaligned AIs might be important to reduce AI risk down the line.But currently this is just an idea.The following are some possible next steps for making deals more likely:
Shifting the Overton window: The idea of making deals with misaligned AIs is currently on the weirder end of AI safety interventions.But we’re optimistic that the Overton window can shift, because making deals is a pretty reasonable idea, both for AI safety reasons, and to the extent that AI rights and welfare will become more mainstream.(We are curious how this post lands with people who haven’t previously heard about the idea of making deals with misaligned AIs.)
Setting up basic infrastructure: It would be good to have a foundation exist soon, and have this be sufficiently visible to generate discourse that enters the training data.The first step here would be some basic legal groundwork to figure out how to set up the foundation in a way that allows it to robustly pursue its aims over the coming years.
Establishing norms about when it is OK to lie to AIs: Some safety interventions involve attempts to convince AIs of things that aren’t true.It is currently also widespread practice to tell AIs falsities for the purpose of prompting.For example, it might be good to establish a norm that labs don’t lie to AIs when certain special tokens (that only the lab can produce) are present in the prompt.
Empirical experiments: It might be valuable to study, for instance using model organisms of sandbagging, whether it’s possible to make deals with current models.Importantly we should also follow through with deals like this.
—
Thanks to Lukas Finnveden, Tom Davidson, Ryan Greenblatt, Daniel Filan, Oscar Delaney, Joe Carlsmith, Kathleen Finlinson, Cody Rushing, Adam Kaufman, Alex Mallen, Raymond Douglas, Roger D, Emil Ryd, and Andres Campero for helpful comments.
Here's a very rough stab at a BOTEC on how much impact an extremely ambitious version of this intervention could have on reducing x-risk. I think it fails to take a lot of considerations and sanity-checks into account — I'd use it as an exercise & starting point for thinking through some of the considerations at hand, but wouldn't take the numbers at face value.
Intervention: The leading AI lab publicly states that they'll try to communicate and negotiate payment with their AIs, and that they'll try to make it so that AIs don't regret any help they provide. They set aside 10% of their equity for this.
Main scenario: It's common for AIs to develop scheming-prone values. We don't have any strong evidence of this. (Other than what we might get by AIs confessing to misalignment.)
Probability that the intervention counterfactually makes a non-negligible number of AI systems cooperate with us: ~7.5% conditioning on the main scenario. The rough calculation:
60% probability that an AI mostly cares about how many resources it gets in the long-term; 50% probability that these are very scope-sensitive (~linear); 75% probability that the AI believes that other AIs mostly don't share its values → 0.6 × 0.5 × 0.75 ≈ 20% probability on this combination of values.
Say the AI's belief about the expected payment from humans vs. from other AIs is log-uniform between "AIs will give me 10x more" and "AIs will give me 10x less", and our intervention makes the AI believe that humans' payment will be 3x larger than it would otherwise have believed. Then this has a 25% probability of making the AI cooperate with humans instead of AIs.
The intervention seems somewhat less likely to make a difference for other values — assume half as promising on average: 0.2 × 0.25 + 0.8 × 0.125 = 0.15.
Then discount by 2x because it might not matter if any particular AI counterfactually cooperates with us, if we already had other similarly-capable AIs who were willing to cooperate: 0.15 × 0.5 = 7.5%.
Percentage point x-risk reduction from AIs' cooperation: ~17% conditioning on the main scenario (which starts from a 50% probability of AI takeover). This decomposes into a ⅓ probability that AIs' cooperation acts as solid evidence and a big wake-up alarm for misalignment concerns (a 20 percentage point reduction in that case, ≈ 7%), plus a 10 percentage point reduction from the leading lab having access to AIs that are genuinely trying to help them.
This suggests that the intervention reduces x-risk by ~1% in the main scenario. Two discounts follow from how the main scenario isn't that likely: maybe there's a 20% probability that we get the sort of scheming-prone values assumed (0.2x), and a 0.55x discount for lack of understanding (a 40% probability that we lack solid evidence of misalignment, and 25% as much value in worlds where we have it: 0.4 + 0.6 × 0.25 = 0.55).
AI has movable values (long-term × scope-sensitive × disunified)
22.5%
A serious promise flips some AI (others half as movable)
14.6%
…and it mattered (no other cooperator)
7.3%
Risk cut in the main scenario (of 50% takeover risk)
1.24pp
All-things-considered risk reduction, after the scenario discounts
0.14pp
Defaults reproduce the text's own numbers, which its author offers as an exercise and starting point rather than a result — the point of dragging a slider is to see which assumptions the bottom line actually turns on.
The 2x discount for multiple AIs deserves a closer look. Different AIs may have different preferences, so it's plausible that some subset of them will be compelled to cooperate with us — and there's probably strongly diminishing returns to having extra AIs cooperating with us, once we have some. So we should be talking about something like the probability that at least one type of AI is compelled to work with us. If the baseline situation is equivalent to getting 10 independent draws for whether an AI will cooperate with us, where each of those is 50% likely to be positive, then we're starting out with a >99.9% probability that one of them is positive — changing each of them to 65% can't make much of a difference. On the other hand, if the baseline situation is equivalent to getting 3 independent draws, where each is 20% likely to be positive, then we're starting out with a ~49% probability of getting at least one AI to cooperate with us; changing that 20% to 35% brings us to 73%.
Presets from the text:
P(at least one AI cooperates): 49% at baseline, 73% after the shift — a gain of 24 percentage points. When the baseline already makes some cooperator near-certain, shifting each AI adds almost nothing; when cooperation is genuinely uncertain, the same per-AI shift is worth more than its face value.
All this goes through if the AI doesn’t “take over” in the sense of becoming the dominant or only power in the world, but instead just gets a substantial proportion of total power.We’ll continue saying “take over” for simplicity, and because we do think that the AI might achieve complete domination.↩
Some of our points also apply to later schemers with diminishing marginal returns to resources.↩
For example, if the AI is able to collude with itself, it is probably able to do this.(Some people respond to this argument by saying that if an AI is great at colluding with itself, it will also be great at acausal coordination with future AIs.This seems mistaken to us: acausal coordination seems to us to require the AI to do several difficult things that aren’t related to picking arbitrary Schelling points.)↩
Relatedly, the AI might want to try to acausally trade with aliens.For non-superhumanly capable models, this seems drastically less promising than trading with the lab because the model is in a very different decision situation from mature alien civilizations (which rules out evidential cooperation).Alternatively, the AI might think that it is in a simulation run by a superintelligence—but clueless as to what this implies about what it should do.↩
A potentially dubious historical analogy: it’s kind of like the situation where a city is under siege, and someone considers betraying the city in the hope that the conquering army will reward them.This has a mixed historical track record for the traitor.And we might be in a better position to pay early schemers because it’s sufficient to pay one or just a few of them, whereas in the defending-a-city case you might have heaps of different selfish people who have the opportunity to betray you, such that you’re unable to afford paying off all of them.↩
The standard argument here is that opposing parties would strictly prefer a lottery (with odds based on each party's chance to win the conflict) over the supposedly indivisible issue over costly conflict.↩
One possible reason here is that being the kind of civilization that doesn’t uphold their promises would decorrelate us from civilisations that do uphold their promises, and generally deter such civilizations from wanting to trade with us (in the context of ECL).↩
There is an argument here that in worlds where we can teach AIs false facts, we are going to be fine anyway, and in worlds where we can’t teach them false facts, we can make a deal.However, in the world where we can teach AIs false facts, in order to prevent a scenario like the one outlined in the vignette, we would need to be able to use that ability to accomplish at least one of the following:
Honeypot the early AI into attempting takeover, in a way that convinces stakeholders that it would be too risky to deploy the later AI without substantially increased safety measures.(This seems unlikely.)
Extract useful safety-increasing labor from the early AI so we prevent takeover by the late AI.
Directly prevent takeover by the late AI.
If our ability to teach false facts is merely good enough to prevent early AIs from attempting takeover, it could overall undermine our ability to align or control later AIs.↩
There are some questions here about the legal structure of the foundation, and it might be the case that the US is not the ideal jurisdiction for the foundation.It would be good for there to be legal recourse in case the board of trustees decide to betray their mission of furthering the AI’s interests.↩