Research Scientist in Scalable Oversight
Conduct theory research on AI safety via Debate as part of Iliad's Scalable Oversight team, drawing on computational complexity, algorithmic information, graph theory or game theory.
About Iliad
Iliad is a nonprofit dedicated to advancing foundational AI alignment research. We run the Iliad Intensive, the Iliad Fellowship, and a range of conferences that bring together researchers working on the hardest problems in AI safety. We also incubate research organizations and an academic journal, and have an internal research team. We're a small and quickly growing team.
About the Role
As a researcher on the Scalable Oversight team you will conduct theory research on AI safety via Debate, one of the most prominent approaches to AI safety. You will also be involved in hiring future team members, supervising mentees and shaping the research direction. We are looking for people with a deep understanding of any of the following: Computational Complexity, Algorithmic Complexity, Graph Theory or Game Theory.
About the group:
We want to study two questions that are intricately related:
- Scalable oversight: How can a weaker judge give a reward signal for honest reasoning for a stronger agent?
- Adversarial sensemaking, i.e., how can a weak judge make robust inferences from potentially adversarially selected evidence and arguments.
We look at these questions in the framework of AI safety via Debate (Christiano et al 2018), as well as extensions of it. The success of debate crucially hinges on judge-oracle complexity, robustness against errors in the judge’s world-model, and attractiveness of the game-theoretic equilibria.
Computational Complexity: We want to build on AI safety protocols like doubly-efficient Debate (Brown-Cohen, 2023) and Prover-Estimator Debate (Brown-Cohen, 2025), and investigate whether obfuscated arguments are a fundamental limitation, or can be remedied. We ideally want to find the limits of debate in terms of oracle complexity for powerful, but computationally constrained provers.
Graph Theory: Debate protocols can be described as games on (belief) graphs, such as factor graphs or Bayesian models, see Experimental Results by Hecket et al.. We are interested in which graph properties are necessary and sufficient for an honest strategy to be strategically favoured, and which types of argument graphs we expect to model real debates well.
Algorithmic Complexity Theory: We want to build a foundation of adversarial sense-making. Solomonoff induction can serve as a basis for a judge in the Bayesian limit, and we want to investigate optimal induction on evidence and arguments provided by both adversarial and cooperative provers.
Game Theory: We investigate which kinds of equilibria emerge in debate games, including their stability, basins of attraction, and robustness to judge errors. Are there more constrained settings than full debate for which we can show that a truthful strategy always dominates?
These directions are starting points rather than a fixed agenda. We expect researchers to challenge the framing and develop new approaches.
What We're Looking For
- Established research track record in computational complexity, algorithmic information, graph or game theory
- Highly creative and rigorous thinking
- Willingness to make great use of research automation,
- Strong vision for what rigorous, high-impact alignment research looks like
- Ability to set and communicate a coherent research agenda across a team
- Comfort working in a small, fast-moving organization where priorities shift
- Comfort working with distributed teams and working remotely
Benefits
Our benefits include:
- Time Off: 25 days paid vacation annually, plus local holidays.
- Flexible Working Hours: We're a distributed organization with employees in a variety of locations and offer flexible working hours.
- Team Retreats: We have occasional in-person team retreats in California and England.