MURANO: Design, Run, and Reproduce Mechanistic Interpretability Experiments as Composable Pipelines
Murano is an open source framework for designing, running, and reproducing mechanistic interpretability studies of large language models, intended for researchers across disciplines and builds on existing interpretability and machine learning libraries.
Alireza Bayat Makou, Emirhan Böge, Phu Gia Hoang et al.
· 0 citations