A research library for mechanistic interpretability and Theory of Mind in large language models
-
Updated
Aug 31, 2026 - Python
A research library for mechanistic interpretability and Theory of Mind in large language models
Realtime unit labeling and debugging demo: streaming NeuronCards with confidence, stable identity under drift, and active probing for AI interpretability.
To associate your repository with the mechanisticinterpretability topic, visit your repo's landing page and select "manage topics."