Our Mission

Human understanding should grow with machine capability.

The International Consortium for Interpretable AI will develop the science needed for humans to understand, and thus effectively harness, artificial intelligence.

By building evidence-based scientific tools that clarify the internal mechanisms of AI, the Consortium will help people control AI behavior, center it on human goals, and enable the safe and beneficial adoption of AI technologies.

The imperative

Artificial intelligence has inverted the relationship between people and computers. Through the history of computing, programming has been humanity’s way of shaping and understanding automation. But today’s neural networks are trained, not programmed. Their internal mechanisms emerge from fitting vast amounts of data, leaving us with the challenge of reverse-engineering the learned programs that underlie their behavior.

Because human understanding is no longer a prerequisite to automation, computer science faces a new imperative: we must ensure that as AI capabilities rise, human agency grows.

The urgency of interpretability is stark: AI systems are already being adopted across society. AI has begun to displace human work across every field touched by software, mathematics, and writing. But ceding human thinking to machine intelligence will make it difficult for people to maintain real responsibility.

Government, science, engineering, and medicine are all predicated on the assumption that complex decision-making intertwines accountability with insight. As AI continues to take on intellectual work, humanity must find ways to incisively understand how AI reaches its conclusions. We must dispel our ignorance and uphold our intentions, so we can take responsibility for results.

A scientific foundation

The science of AI interpretability is nascent. While many studies have observed the structure of AI computations through a variety of lenses, the field lacks consensus around shared definitions, methods, and standards of evidence.

ICINAi will build a solid scientific foundation for empirical, evidence-driven AI interpretability by convening researchers, developing infrastructure, and creating scientific commons. Its work will be guided by the urgent mission of increasing human insight and responsibility as AI becomes more autonomous, and will advance three specific scientific goals.

Three scientific goals

01

White-box AI auditing

Can we build a lie detector for AI?

In both routine and high-stakes use, there is a gap between what an AI system says and what it knows internally. Closing this gap is a fundamental challenge for AI safety.

The Consortium will develop rigorous methods for inspecting AI computations to detect deception, hidden goals, censorship, hallucination, reward hacking, and other serious failures missed by ordinary evaluation. It will impartially steward benchmarks and experimental protocols—including competitions such as Aletheia’s Quest and model-organism benchmarks—and work with industry and other AI-safety efforts to build consensus, infrastructure, and expertise.

02

Understanding intelligence

What is thinking?

With AI, we now have silicon brains whose every computation can be observed, altered, and repeated. We face a historic opportunity to investigate ancient questions about the nature of reasoning and thought.

The Consortium will study how computation gives rise to concepts, abstraction, belief, planning, learning, and agency by dissecting neural representations and tracing causal pathways. We will develop measurable definitions for cognitive phenomena and evidence-based experimental methods to understand the profound capabilities of large neural networks.

03

Empowering people

How can AI amplify human agency?

The most valuable aspect of AI is its knowledge beyond human knowledge, but the science of AI will have failed if it only serves to make AI systems superhuman. People cannot answer for decisions they do not understand.

To make people more capable, AI capability must be paired with human comprehension. The Consortium will develop methods for identifying and explaining AI discoveries at the edge of human knowledge, translating AI knowledge into human insight, measuring whether those insights improve human judgment, and enabling people to understand the implications of complex choices made with AI.

Our principles

The Consortium is driven by an obligation to ask the important questions even when unprofitable, to hold fast to the optimism that hard problems can be solved, and to celebrate the serious play that connects understanding with action.

For humanity to be resilient in the face of unexpected AI challenges, we must translate insight into the technical means to correct AI, and build a culture in which people expect to exercise that power. The practice of interrogating and changing AI systems must be practical.

AI will not become understandable to us simply because we created it. Bringing AI within reach of human understanding will require us to build a new science; by creating a consortium, we have the opportunity to spearhead that effort.