Your vision will become clear only when you look into your heart.... Who looks outside, dreams. Who looks inside, awakens. Carl Jung
Sunday, August 9, 2026
AI Safety and Mathematecians
https://www.mathforaisafety.org/
AI Safety For Mathematicians
A very high-level resource
A starting point for mathematicians who want to engage with AI safety.
Purpose of this site
This is designed to be a very high-level resource for mathematicians to get involved with AI safety work. We believe that:
AI is poised to have a substantial impact on society in the coming years.
With that come many new risks that we need more effort to manage effectively.
There are many opportunities for mathematicians to contribute, including some in which they have a strong comparative advantage.
This is not meant to put anyone in a box. Mathematicians and non-mathematicians have varied skill sets and interests, and we are glad if this site helps a broader group of people. Its target audience is professional mathematicians looking to engage.
What is AI safety?
AI safety asks how we can make sure AI systems do not lead to negative outcomes for humans. AI alignment asks how we can make sure those systems act, or try to act, in ways compatible with human values.
These fields tend to be composed of many heuristic notions that have not yet been made precise. Much of the work lies in formalizing vague notions and finding appropriate theoretical frameworks.
How do we understand what an AI is doing?
How do we get AI systems to cooperate with us and with each other?
How do we tell if an AI is lying or scheming?
Research directions
We lay out some research directions and explain them at a very high level aimed at mathematicians. At the end, we provide resources for further exploration. This is meant to be an expanding list of directions, and the writeups are subject to change (ostensibly for the better!)
Developing Good Heuristics
→
Interpretability and Feature Manifolds
→
Open-Source Game Theory
→
(More Research Directions forthcoming!)
Maintained by Jacob Tsimerman — all errors in the writing are due solely to me!
Contributors: Andrew Critch, Lionel Levine, Yevgeny Liokumovich, Arul Shankar