How do you translate feasibility judgments into identifying, executing, and communicating an impactful research project? Below, you will learn some important strategies for AI governance research from MIRI senior researcher Aaron Scher.
Aaron’s Research Tips, or How I Wish I Did Research
The author keeps adding to the document. This page re-syncs from it, so follow the link if you want to be certain you are reading today’s version.
Links:
How To Do Research Well Reading List [Public, Shareable]
Medium Level AI Governance Reading List [Public, Shareable]
Project Selection and Scoping
- The primary goal of research is to get accurate answers to important questions.
- Therefore, the purpose of project scoping is to identify an important question for which it is possible to get more accurate answers about, and identify the pathway to doing so.
- Possibility here is constrained by your abilities, your time commitment on the project, your access to information, money, what existing tools and understanding humanity has to bear on the problem, and occasionally physical constraints in the world. Usually it’s constrained by your abilities and time.
- Importance is tough to assess but should broadly be based on a reasonable aggregation of worldviews, weighted toward the ones you yourself believe. An important question might be one that could substantially change AI governance plans depending on how it resolves. It might answer a question that future people are likely to ask and for which having an accurate answer seems really useful for those future people to make good choices.
- By putting importance and possibility together, you can think about project selection as an expected value calculation. You can also think about this as Scale times Tractability: how good would a successful outcome be, and how likely would you be to obtain that successful outcome (summed across different outcomes of success). Sometimes people use neglectedness to help them with problem selection, but neglectedness is just a heuristic for assessing tractability (because if there’s lots of people working on a problem, it’s more likely that somebody has picked the low hanging fruit).
- Research questions should ideally ground out in action-relevance: producing an output that is likely to influence further action in a meaningful way. If nobody should or will take any actions differently depending on what you find in your research, don’t do the research. Some good research questions are not directly action-relevant, but it is still fairly obvious why they will be useful to answer (e.g., they resolve an uncertainty that is upstream of many other questions with action-relevance).
- Relatedly, your research should have a Theory of Change (ToC) and you should write it down. A ToC is a clear pathway by which the work you are doing will lead to a better world state according to your long-term goals.
- Some research creates knowledge that is completely new for humanity, while some research takes known ideas or knowledge but uses them to answer a question you or your team cares about and propagate that answer to the right people.
- For every project, you should have in your head the answer to both the high level: “Why does TGT [your team] care about this project? For instance, what is the uncertainty TGT is trying to reduce with this project, or if we already have the answer, what is the key idea we are trying to convey to our audience?”
- And the lower level “What evidence or reasoning process will lead to a good answer to this question, such that I will find that evidence and reasoning compelling?” You should have this both for the “ideal version” of the project and the “realistic version” of the project (i.e., only 35% as good)
- You should only work on projects that you really believe would be useful. Ideally both the ideal version and the realistic version of the project.
- If you have a project direction you believe in, that is the core motivation. You can use it to motivate work on the project.
- Whenever you're trying to make a decision, a thing that can help you make a better decision is to have a clear picture of what your counterfactual is. What are your other options? Don’t decide on a project direction when you only have one decent idea; generate some other decent ideas.
Thinking Process
- Much of AI governance research is making connections between ideas, for instance, realizing when something you read six months ago is a relevant place to dive into further in order to answer your current question.
- Therefore, you should aim to have a large and high-quality cache of “things you read six months ago”. I recommend this reading list, and the most relevant part of it is the list of AI governance research organizations whose work you should skim over.
- What evidence comes to bear on this question?
- What is the ideal evidence that I wish I had?
- What evidence is easy to acquire? E.g., literature reviews often shed some light, asking an LLM is easy and sometimes useful
- What are the key quantities that have bearing on this question? Do a Fermi estimate or back of the envelope calculation for them.
- Which questions should you answer and which ones can you not answer? Lots of AI governance research questions are not resolvable. They rely on extremely hard-to-predict aspects of the future. Therefore, a key part of AI governance research is identifying the tractable questions that do actually matter and that are ideally not drowned out by intractable questions.
- For example, in this international agreement, we discuss how ideally inference on existing models can continue, but this is not a certainty. We structure the agreement around the assumption that it is fine for inference on existing models to continue, but we don't actually know if that will hold in practice. That is because it depends on when the agreement comes into effect and how strong AI capabilities are at that point. Our research process was roughly to identify that this is a key question, to identify the key considerations that determine which way the world should go on that question, and then to continue our research while leaving that as an open question.
- Another key aspect is figuring out which parts of the giant intractable questions are actually somewhat tractable and then working on them.
- For example, this research is motivated by the overall question: How fast will algorithmic progress be during an intelligence explosion and during a potential halt on AI capabilities? Neither of those is directly answerable, but a proxy question is: How fast has algorithmic progress been historically? Then I operationalized that as catch-up algorithmic progress in particular, as measured naively through pre-training compute and benchmark scores over time.
- For example, this research is motivated by the overall question: If you impose compute constraints and monitoring during an AI halt, will people still be able to do algorithmic research? This is not directly measurable, but one proxy is how much compute has been used for algorithmic research historically. That itself is also not directly measurable very easily, but a proxy for it is looking at the papers that introduce key algorithmic innovations and looking at the experiments done in those papers and seeing how much compute is used on those experiments. You can see the paper for discussion of important ways this proxy differs from the quantity we really care about.
- For example, this research is motivated by the question: Will the limits and definitions in our international agreement be sufficient to stop distributed training? This is hard to directly answer because distributed training methods are likely to improve in the future. But one approach is to look at existing distributed training methods and create a simulator to see how effective they would be at circumventing the rules. Another way that we attempted to answer a very similar question previously was to just do a quick calculation based on existing distributed training methods, multiplying together the ones that looked like they would stack, to get a rough order-of-magnitude estimate for how these methods might change communication requirements.
- Embrace the most relevant of the 12 virtues of rationality: Curiosity, relinquishment, lightness, evenness, argument, empiricism, simplicity, humility, perfectionism, precision, scholarship, and the void.
- Aim to understand things deeply.
- Frequently be asking yourself questions like:
- What would a skeptical reviewer say about this?
- Is a member of my target audience going to understand this conceptual breakdown?
- Know your audience. From the start, your Theory of Change states who your audience is. Have them in mind all the time: when you consider which types of evidence to collect, when you consider how to structure your output, when writing individual sentences.
- Backchain frequently.
- How to solve a problem
- In various parts of my life it is often useful that I have some strategies for solving problems. Many problems can be solved via applying the right methods, applying strategic thinking to your problem. I am by no means an expert on this, but I seem to be considerably better at it than many other people, based on using the following strategies.
- Try to understand the problem. If there are other people involved, check with them to make sure you understand, including saying things like “let me say this back to you to make sure I understand what you mean”.
- Spend 5 minutes brainstorming potential solutions. They don’t all have to be perfect, they don’t even have to solve the whole problem. During this initial time, don’t get stuck on one idea, and try not to even spend time analyzing how good they are (though often it is very fast to dismiss some ideas).
- Ask yourself “how would a really agentic and strategic person solve this problem?” and then do that.
- Make a pro / con list for the top potential solutions that you are considering.
Non-thinking Process
- Every morning I write a to-do list like this one and write reflections at the end of the day about how the to-do items went.
- Use a second monitor when you can. Keep the notes you are taking on one screen and searches, ChatGPT, etc. on a second screen.
- Keep a list of Future Project Ideas that you have a low bar for adding to. When you think of a project that you might want to work on, jot it down there so you don’t forget but also so you don’t get totally interrupted thinking about it.
- When in the writing / communication phase of a project, keep a list of project to-dos, including more aspirational to-dos, which are lower priority and likely to not be done for the current project but might be a good fit for follow-up work.
- A skill I wish I was better at is breaking larger tasks down into components that could be completed by someone on Upwork (or an AI, as is now the case). This is one way to use money to speed you up.
- In order to contribute to “the conversation”, you need to be intimately familiar with the conversation.
- Read lots of papers. Supervised learning works for humans; if you want to write good papers, you need to know what good papers look like, how they are written, etc.; that requires reading lots of papers.
- Do the crappy version first. When you're considering doing a project, often there is one option to directly try to do the thing, and another option to try to do more research investment into a longer-term, better version of it (bells and whistles, if you will). I.e., do you run simple experiment A, or spend more time trying to make the experiment more realistic first. You should usually do the crappy/simple version first. Often doing the crappy version is how you figure out what the good version should look like: you realize which parts are hard, you make a bunch of mistakes which you can then fix.
- Don’t outsource important work or thinking. Many of us have the desire to take on many mentees or use AI to do lots of our work for us.
- Write well. Here are my writing tips, manny of which are me-specific: Aaron’s Writing Tips [Public, Shareable]

