None
EN
A Discussion of 'Adversarial Examples Are Not Bugs, They Are Features': Discussion and Author Responses
['Engstrom', 'Ilyas', 'Madry', 'Santurkar', 'Tran', 'Tsipras']
Distill
Our goal is to say that since adversarial examples can arise from well-generalizing features, simply patching up the “bugs” in ML models will not get rid of adversarial vulnerability — we also need to make sure our models learn the right features. However, we want to stress (as the comment itself does) that robust feature leakage does not have an impact on our main thesis — the D ^ d e t \widehat{\mathcal{D}}_{det} D det dataset explicitly controls for robust feature leakage (and in fact, allows us to quantify the models’ preference for robust features vs non-robust features — see Appendix D.6 in the paper). Response Summary: A fine-grained look at adversarial examples that neatly complements our thesis (i.e.