The ICINAi community

Building the science of interpretable AI, together.

Our steering group brings together researchers across institutions, disciplines, and sectors to help shape ICINAi’s scientific direction and shared infrastructure.

Steering group

Researchers helping guide the consortium’s priorities, programs, and standards.

David Bau

David Bau

Northeastern University

David Bau is an assistant professor at Northeastern University and director of the National Deep Inference Fabric. His research studies emergent internal mechanisms in large neural networks across language and vision, including causal tracing, model editing, and shared interpretability infrastructure.

Mor Geva

Mor Geva

Tel Aviv University

Mor Geva is an assistant professor in the School of Computer Science and AI at Tel Aviv University. Her research opens up the inner workings of large language models to improve their transparency, control, and reasoning.

Yonatan Belinkov

Yonatan Belinkov

Technion – Israel Institute of Technology

Yonatan Belinkov is an associate professor at the Technion. He develops probing, causal-analysis, and evaluation methods for understanding what language models represent and how their internal computations produce behavior.

Ellie Pavlick

Ellie Pavlick

Brown University · Google DeepMind

Ellie Pavlick is an associate professor at Brown University and a research scientist at Google DeepMind. She studies conceptual representations, language, reasoning, and learning, often by comparing the behavior of humans and language models.

Christopher Potts

Christopher Potts

Stanford University

Christopher Potts is a professor of linguistics at Stanford University, with a courtesy appointment in computer science. His research connects compositionality, causal abstraction, interpretability, information retrieval, and foundation-model programming.

Roger Grosse

Roger Grosse

University of Toronto · Vector Institute · Anthropic

Roger Grosse is a professor at the University of Toronto, a Vector Institute faculty member, and a researcher at Anthropic. His work spans mechanistic interpretability, alignment science, neural-network optimization, and representation learning.

Aaron Mueller

Aaron Mueller

Boston University

Aaron Mueller is an assistant professor of computer science at Boston University. His lab studies how concepts are learned and represented in language models, developing tools for circuit discovery, interpretability, and model control.

Naomi Saphra

Naomi Saphra

Boston University

Naomi Saphra is an assistant professor in Computing & Data Sciences at Boston University. Her research examines language-model training dynamics, how structure emerges during learning, and what those processes reveal about generalization and interpretability.

Tamar Rott Shaham

Tamar Rott Shaham

Weizmann Institute of Science

Tamar Rott Shaham is an assistant professor of computer science and applied mathematics at the Weizmann Institute. She develops automated tools to discover and explain model operations, then uses those insights to control model behavior.

Ana Marasović

Ana Marasović

University of Utah

Ana Marasović is an assistant professor in the Kahlert School of Computing at the University of Utah. Her work in natural language processing and human-centered AI develops more faithful explanations and evaluations of language models.

Ivan Titov

Ivan Titov

University of Edinburgh · University of Amsterdam

Ivan Titov is a professor at the University of Edinburgh and the University of Amsterdam. He studies trustworthy, robust, interpretable, and controllable language models and leads major European training and research programs in natural language processing.

Tal Linzen

Tal Linzen

New York University · Google

Tal Linzen is an associate professor at New York University and a staff research scientist at Google. His work connects language-model evaluation and interpretability with linguistics and cognitive science, especially the study of syntactic structure.

Byron Wallace

Byron Wallace

Northeastern University

Byron Wallace is a professor at Northeastern University and director of its BS in Artificial Intelligence program. His research develops machine-learning and natural-language-processing methods for health informatics, with particular attention to faithful explanations, biomedical evidence synthesis, and human–AI systems.

Yanai Elazar

Yanai Elazar

Bar-Ilan University

Yanai Elazar is an assistant professor of computer science and AI at Bar-Ilan University. He builds tools for a science of generative models, including methods for interpretability, data attribution, and rigorous evaluation.

Mario Giulianelli

Mario Giulianelli

University College London · Parallax

Mario Giulianelli is an associate professor of computational linguistics at University College London and a cofounder of Parallax. He combines behavioral evidence with analysis of internal mechanisms to make AI evaluation more rigorous and explanatory.

Gabriele Sarti

Gabriele Sarti

Parallax

Gabriele Sarti is a cofounder and technical director at Parallax. He develops white-box methods for context attribution, user modeling, deception detection, and model monitoring, alongside open-source libraries for language-model interpretability.

Sarah Schwettmann

Sarah Schwettmann

Transluce

Sarah Schwettmann is cofounder and chief scientist at Transluce, a nonprofit research lab building public infrastructure for responsible AI deployment. Her research reverse-engineers neural networks and develops tools for transparent, auditable, and steerable AI systems.

Eric J. Michaud

Eric J. Michaud

Simplex · Astera Institute

Eric J. Michaud is a research scientist at Simplex and the Astera Institute. He pursues empirical and conceptual accounts of learned computation, including work on sparse feature circuits, representation geometry, and mechanistic program synthesis.

R. Thomas McCoy

R. Thomas McCoy

Yale University

R. Thomas McCoy is an assistant professor of linguistics at Yale University, with a secondary appointment in computer science. He connects AI and cognitive science, using language-model interpretation and evaluation to study both machine intelligence and human language.

Noah D. Goodman

Noah D. Goodman

Stanford University

Noah D. Goodman is a professor of psychology and computer science at Stanford University. His research uses probabilistic models to study language, cognition, and reasoning, including what language-model behavior reveals about internal representations and thought.

Zachary Lipton

Zachary Lipton

Carnegie Mellon University

Zachary Lipton is the Raj Reddy Associate Professor of Machine Learning at Carnegie Mellon University. His research examines the conceptual foundations of interpretability, robust and adaptive machine learning, natural language processing, and real-world decision systems.

Anna Ivanova

Anna Ivanova

Georgia Tech

Anna Ivanova is an assistant professor of psychological and brain sciences at Georgia Tech, where she leads the Language, Intelligence, and Thought Lab. She combines brain imaging, behavioral studies, and computational modeling to study language and thought in humans and AI.

Hinrich Schuetze

Hinrich Schuetze

LMU Munich

Hinrich Schuetze is professor and chair of computational linguistics at LMU Munich. His research spans statistical natural language processing, representation learning, lexical semantics, and the analysis of multilingual language models.

Amir Globerson

Amir Globerson

Tel Aviv University

Amir Globerson is a professor of computer science and AI at Tel Aviv University. He studies machine learning and natural language processing, including the theory of learned representations, attention, and interpretable reasoning.

Fazl Barez

Fazl Barez

University of Oxford

Fazl Barez is a senior researcher at the University of Oxford and leads the TSG Lab. His work connects interpretability, evaluation, model control, and machine unlearning with the technical foundations needed for effective AI governance.

Robert West

Robert West

EPFL

Robert West is an associate professor at EPFL, where he leads the Data Science and AI Lab. His work spans AI, natural language processing, and computational social science, including safe AI and structured human–AI collaboration.

Emre Yavuz

Emre Yavuz

Cambridge Boston Alignment Initiative

Emre Yavuz directs programs at the Cambridge Boston Alignment Initiative. He builds research communities and talent pathways for AI safety and biosecurity, connecting researchers through fellowships, training, and cross-institutional collaboration.

Adam Tauman Kalai

Adam Tauman Kalai

Institute for Responsible Superintelligence

Adam Tauman Kalai is a cofounder of the Institute for Responsible Superintelligence and a former research scientist at OpenAI. His work spans AI safety and ethics, algorithms, fairness, AI theory, game theory, and crowdsourcing.

Tommaso Tosato

Tommaso Tosato

Tara Research · Mila

Tommaso Tosato is a cofounder of Tara Research and a postdoctoral fellow at Mila. His research measures honesty in frontier agents and evaluates white-box techniques, including activation steering, for detecting and reducing deceptive behavior.

Erik Miehling

Erik Miehling

IBM Research

Erik Miehling is a research scientist in the human-centered AI group at IBM Research. His work focuses on generative-model steerability, multi-agent coordination and alignment, and the dynamics of human–AI interaction.

Diego Garcia-Olano

Diego Garcia-Olano

Meta

Diego Garcia-Olano is a senior research scientist at Meta. His work spans safety alignment, language-model interpretability, and the evaluation and mitigation of risks including hallucination, memorization, and sensitive-data exposure.