
Lessons from Buck Shlegeris
Buck Shlegeris is CEO of Redwood Research, where he works on AI control, an approach to mitigating catastrophic misalignment risk from AI systems. — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 1: AI Control Fundamentals
- On AI Control: Buck Shlegeris: "I would say AI control is an approach to mitigating catastrophic misalignment risk." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On an Imagined World: Buck Shlegeris: "a lot of people who worry about AI catastrophic risk have imagined a world where people are trying really hard to mitigate these risks, where companies are strongly regulated" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Catching an Escape Attempt: Buck Shlegeris: "Suppose the AI tries to escape, and you catch it on a Tuesday, and then you keep deploying it. I think it’s seriously plausible that you’ll feel safer on Wednesday than you felt on Monday." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Five Actions: Buck Shlegeris: "it’s just generally a lot easier to tell the difference between five benign actions and five malign actions than to tell the difference between one benign action and one malign action." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Going Back in Time: Buck Shlegeris: "there’s actually no difficulty going back in time here." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Auditing: Buck Shlegeris: "Auditing is this process where we look into actions that AI is taking. And the core quantitative issue we have here is that we can audit so infrequently, right?" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Acute Problems: Buck Shlegeris: "the acute problems occur very rarely, but also probably are very incriminating." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Defer to Trusted: Buck Shlegeris: "which is when the action suggested was too sketchy, you replace it with an action by the trusted model." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Collusion Signals: Buck Shlegeris: "The main strategy available is trying to destroy collusion signals." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 2: Scheming and Deceptive Alignment
- On Correlated Scheming: Buck Shlegeris: "in the AI situation, it’s very correlated: there are probably either none scheming against you or all scheming against you." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Sting Operations: Buck Shlegeris: "I think it’s actually really rough for them to collude, basically because we can constantly run sting operations on both human employees and the AIs." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Policy Versus Monitor: Buck Shlegeris: "it’s hard to run sting operations on your policy, but this is not true of your monitor." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On the Monitor's Position: Buck Shlegeris: "when you’re a monitor, when you see an attack, it is a big deal. You are in a much worse situation if you ever make the wrong call in either direction." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On a One-in-a-Million Prior: Buck Shlegeris: "you start out with like a one is to a million prior on any particular attack that you see that you are being honeypotted." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Passwords: Buck Shlegeris: "it’s just pretty easy to have passwords that are very unlikely to come up by chance." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On the Word Aligned: Buck Shlegeris: "I’ve been perhaps foolishly using the word “aligned” to mean the model isn’t aggressively conspiring against you." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Evidence of Misalignment: Buck Shlegeris: "if it just constantly claims that there was not a backdoor in clearly backdoored code, that’s evidence the model is misaligned." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Skeptics After a Catch: Buck Shlegeris: "people who are very sceptical of misalignment before observing this will probably be pretty sceptical afterwards." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Rigged Training: Buck Shlegeris: "if I were an AI company that wished to race ahead, and then my competitor informed me that their AI had tried to escape, I would at least claim that it was plausible that my competitor had rigged the training of their AI" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 3: The Safety-Usefulness Tradeoff
- On Expensive Safety Techniques: Buck Shlegeris: "there’s a bunch of different reasons why it’s bad for safety techniques to be expensive. One reason is you might have limited political capital." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Usefulness and Safety: Buck Shlegeris: "The obvious way to measure the performance of techniques is to talk about their usefulness and safety." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Affordable Auditing: Buck Shlegeris: "The basic problem with that class of techniques, as we were saying, is that quantitatively you just can’t afford to do it very much." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On a 0.1% False-Positive Rate: Buck Shlegeris: "Suppose you can only afford to audit one thing in 1,000. Then your classifier for sketchiness has to have a 0.1% false-positive rate when you’re auditing, right?" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Halving the Risk: Buck Shlegeris: "there was previously a 30% chance that your AI succeeded at taking over your data centre, and now there’s a 15% probability or whatever. And to me that sounds pretty good." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Replacing at 2%: Buck Shlegeris: "I think that even replace at 2% gets you probably substantial safety compared to nothing." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Three Outcomes: Buck Shlegeris: "whenever the AI tries to attack, one of three things happens: it succeeds and then game is over for us" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Long Trajectories: Buck Shlegeris: "imagine that it only had a 50% chance of continuing every step: if the trajectory is six steps long, it’s going to make it to the end one time in 32." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Resampling: Buck Shlegeris: "before escalating to the auditor, before escalating to this more expensive process, is you can resample actions in various places in the trajectory and look at how suspicious the resamples were." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 4: Developer Incentives and Deployment Dynamics
- On Luck: Buck Shlegeris: "The world looks plausibly safer, but a lot of that is luck." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Misalignment Risk Over Time: Buck Shlegeris: "I think that AI misalignment risk looks plausibly lower than it looked 10 years ago." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Cheap Interventions: Buck Shlegeris: "there are a bunch of relatively cheap interventions that they could potentially get implemented inside these AI companies that the company wouldn’t fight them too much on." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Ten Chosen People: Buck Shlegeris: "there are a bunch of interventions that could be done with 10 appropriately chosen people that seem like they might substantially reduce risk." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Implementations Over Novel Research: Buck Shlegeris: "a lot of people should not be trying to produce novel research; they should be trying to get implementations of stuff set." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Industry Bodies: Buck Shlegeris: "suppose some regulator or suppose an industry body says that everyone should do the following set of control techniques." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On AI Feeling Far Away: Buck Shlegeris: "AI felt further away, so it felt better to people to look at it from kind of the big-picture perspective." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Implementation: Buck Shlegeris: "I’m optimistic about them implementing either mediocre or pretty good versions of these techniques." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Agent Scaffolds: Buck Shlegeris: "I would like it if the agent scaffold people did do some basic replacing" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 5: Empirical Research and Methodology
- On Red Teams: Buck Shlegeris: "we had a red team. So a blue team suggests one of these strategies — like defer to trusted or whatever — and then we have a red team which tries to build a scaffold out of our untrusted model" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On a While Loop: Buck Shlegeris: "use a ‘while loop’ somewhere where a ‘for loop’ would have been more idiomatic as the signal that this is a real attack." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Safety Cases: Buck Shlegeris: "one of the great things about AI control is it’s actually conceivable to write safety cases using it." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Not Rigging Results: Buck Shlegeris: "they are better than other techniques we tried in this paper, and I don’t think we rigged the results that much." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Restricting Context: Buck Shlegeris: "One way we could make that less likely is if we restricted the context that the agent has" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Weaker Monitors: Buck Shlegeris: "A problem with this is that the trusted model is much weaker, so it might miss stuff." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Coordinating on a Code: Buck Shlegeris: "can the models coordinate on a code, and can the models send a code to each other given our anti-code-sending countermeasures?" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Catching a Few Attempts: Buck Shlegeris: "I think it’s reasonably plausible that even catching a few escape attempts will reduce the probability that future escape attempts succeed." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Resamples Looking Benign: Buck Shlegeris: "when you resample, it’s probably not going to look so suspicious anymore." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Untrusted Monitoring: Buck Shlegeris: "I have another favourite class of control mechanisms I’m talking about: untrusted monitoring." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 6: Career Advice and Transitioning to AI Safety
- On Where Safety People Work: Buck Shlegeris: "AI safety people should be more broadly distributed between AI companies, and overall fewer of them should be inside AI companies." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On AI Futurism: Buck Shlegeris: "We like hiring people who are really into AI futurism." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Knowing What You Want: Buck Shlegeris: "you should try to figure out what you really want to be doing and who you really want to be good at, and try to do that before hiring a bunch of people." — The inaugural Redwood Research podcast (transcript).
- On Hiring Bloggers: Buck Shlegeris: "it’s plausible that we should hire a bunch of people to just blog and write LessWrong posts about AI safety." — The inaugural Redwood Research podcast (transcript).
Part 7: Evaluations and Safety Cases
- On Safety Cases: Buck Shlegeris: "Maybe one way of thinking about control is a way of establishing a safety case." — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On Capability Evaluations: Buck Shlegeris: "We can do a capability evaluation where we try to make our untrusted model as good as possible at coming up with problematic code that humans have trouble spotting, and just empirically evaluate how good the humans we were planning on using actually are at this task." — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On the Control Methodology: Buck Shlegeris: "you should follow the same kind of control evaluation methodology to try protocols kind of like the ones we suggested in order to determine which ones seem safe with future systems." — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On Resolvable Questions: Buck Shlegeris: "this is an extremely resolvable empirical question." — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On Evaluating Interpretability: Buck Shlegeris: "we started looking for good ways of evaluating whether our interpretability work was good, or whether anyone’s interpretability work was good." — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On the Threat Model: Buck Shlegeris: "the threat model we’re worried about is the AI doing a better job the first time that it starts taking obviously bad actions than the red team was able to do" — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On Conservatism: Buck Shlegeris: "I have mostly been ignoring this question out of a general desire for conservatism" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
Part 8: Epistemics, Deference, and Seeking Criticism
- On Work That Does Not Apply: Buck Shlegeris: "almost all of the work which actually goes on, especially by people who are competent researchers or are employed or something, definitely does not apply directly to working on arbitrarily powerful AI" — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).
- On a Sad Dynamic: Buck Shlegeris: "there’s a really sad dynamic among AI safety people where they focus much less than I think is appropriate on really scary situations" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Sleeping on a Point: Buck Shlegeris: "People inside the AI safety community, in my opinion, really slept on this point for a long time" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Confidence About Alignment: Buck Shlegeris: "it’s definitely not clear that we will be very confident the AIs aren’t misaligned." — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Iffy Situations: Buck Shlegeris: "feeling like the situation is much iffier for the AIs than people might have immediately assumed" — 80,000 Hours #214: Buck Shlegeris on controlling AI that wants to take over (transcript).
- On Action and Inaction Risk: Buck Shlegeris: "obviously I’m facing both inaction risk, which is if I just don’t move fast enough, other people might do really dangerous things and bad things might happen in the world" — AXRP episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (transcript).