None
EN
The Building Blocks of Interpretability
['Olah', 'Google Brain', 'Satyanarayan', 'Johnson', 'Google Cloud', 'Carter', 'Schubert', 'Ye', 'Mordvintsev', 'Google Research']
Distill
For example, besides the one-way attribution that we currently perform with the input and output layers, semantic dictionaries allow us to attribute to-and-from specific hidden layers. An interface to understand how dataset examples influence the output classification, as presented by Koh & Liang An interface showing how examples influence the channels of hidden layers. Beyond interfaces for analyzing model behavior, if we add model parameters as a substrate, the design space now allows us to consider interfaces for taking action on neural networks.