所属机构
已发表
尚未发表。
DOI
暂无 DOI。
Chris Olah, Adam Jermyn
本文就可解释性研究中定性方面可能比我们其他领域习惯的更为核心这一观点,提出一些个人见解。同时,本文也旨在描述定性研究工作中研究品味的一些启发式原则。
早期的科学领域通常非常定性化,随着它们逐渐成熟,会变得更加定量化。例如,发现细胞是一个定性结果,随后(经过数十年)可以发展成定量工具,比如在癌症研究中计数白细胞。发现化学光谱线是一个定性结果,直到玻尔意识到“原子的蝴蝶翅膀”揭示了电子轨道的奥秘,它才真正变得定量化。
绝大多数研究人员接受的是成熟学科的培训,因为真正全新的科学领域很少见。这些成熟学科拥有既定的范式,以及成熟的定量衡量标准和方法。但可解释性并非一个成熟的领域。它没有既定的范式。甚至连最基本的抽象概念(比如用“特征”来思考模型是否合理?)都尚存争议。
如果我们把来自成熟学科的培训经验直接套用到这样一个早期、混乱、尚未确立的科学领域,可能会产生错误的直觉。尤其值得注意的是,我们应该预期需要更多地依赖定性结果来指导方向。
需要明确的是,这并非说我们不应该在适当的时候进行定量研究!事实上,这两者常常可以协同增效,定性研究有助于我们确信自己使用了正确的定量工具。(它们之间的界限有时也很模糊!)相反,本文的目的仅仅是论证,定性结果应被真正视为一等公民,并且是我们需要不断回归的试金石,以免迷失方向或自欺欺人。
汇总统计量及其风险
安斯库姆四重奏是一个著名的例子,展示了几个截然不同的数据集如何拥有相同的均值、标准差和相关性:
这反映出一个更普遍的教训:将丰富的高维数据压缩成单一数字的汇总统计,总会让你对大部分实际情况视而不见。因此,在这样做时必须非常谨慎。尤其重要的是,如果你要依赖某个测量指标,就必须非常确定自己了解并信任所测量的对象。
在成熟的学科领域,可能存在一些被充分理解的标准测量方法和量化数值——人们清楚它们的含义,也知道如何解读它们。但在可解释性领域,我们并没有这种优势。
我们在一些可解释性研究以及与之相关的机器学习论文中看到一种模式:先定义一个声称能对应某种感兴趣属性的指标,然后非常严格地测量这个指标。我们认为这是一种“货崇拜式科学”。它看起来可能非常严谨,有大量带标准差柱状图的折线图等等。但往往并非如此,因为关键弱点在于该指标是否真正、可靠地追踪了感兴趣的属性,而非评估该指标的严谨程度。我们自身工作中最近的一个例子是字典学习中关于 tanh 正则化的研究,起初汇总统计指标显示出很大潜力,直到后来对特征进行定性检查时才发现我们被误导了。
我们推测,在像可解释性这样的早期科学领域,严谨性往往涉及更多的定性工作,量化指标应在此基础上发展而来,并且最初应受到相当审慎的对待。
因缺乏受限假设空间
成熟领域能够如此依赖汇总统计的原因之一,在于它们拥有一个受限的假设空间。
想想看,成熟的学科领域通常如何设计实验:它们往往是在两个对立假说之间进行检验。例如,引力波实验可能会被设计用来检验广义相对论与切伦-西蒙斯引力理论孰是孰非。这种做法在假说空间狭窄且被充分理解的情况下是合理的,此时概率质量的主体确实集中在被检验的少数几个假说上。在这些情况下,通常相对容易想出一个单一的度量标准,该标准在两个假说之间应该存在差异!
但在前范式学科中,我们根本不知道该考虑哪些假说!一位19世纪的物理学家若试图解释太阳的能量产生机制,完全不可能想到核物理学。
因此,在前范式学科中,科学的目标首先是弄清楚我们应该考虑哪些假说!这意味着要以一种相当不同的方式开展工作。虽然汇总统计量在区分少量假说方面可能非常有效,但单个数字无法提供足够丰富的信号,来引导我们在浩瀚的可能性空间中定位方向。
界面与汇总统计量的诱惑
为什么汇总统计量如此流行?原因之一涉及科学研究中的界面生态系统。我们通常不会以这种方式思考问题,但科学研究实际上隐含着用于思考数据的界面。例如,方程式、线图和术语都是界面。请参阅下方的延伸阅读,特别是《思考不可思考之事的媒介》(关于科学思维中界面的精彩讨论)和《分离理论》(关于“纸质工具”界面的讨论)。
最常见且可复用的界面(例如线图)通常需要一个标量汇总统计量。这些简单的数据类型是研究中的一种通用语言,因为它们非常普遍,几乎可以跨科学领域复用,并拥有标准化的实践方法。但它们要求我们将结果简化为一个汇总统计量。
呈现更复杂(且规模更大)的数据需要定制化界面。事实上,在可解释性研究中处理如此规模的数据时,如果不借助定制化(且往往是交互式)的界面,几乎不可能在不将其简化为汇总统计量的前提下完成工作。这正是可解释性研究与数据可视化紧密交织的原因。正如早期化学依赖定制玻璃器皿——实际上,许多科学家都自己动手制作玻璃器皿!——可解释性研究同样依赖数据可视化。
汇总统计量与可辩护性
汇总统计量流行的另一个原因是其具有可辩护性。更广泛的科学界和科学文化在潜移默化地告诉我们:“如果你以几种常见格式之一呈现数据,那么即便在最坏的情况下,你仍然是在做科学。”
在成熟的学科领域,这种做法是合理的:人们已经就使用哪些恰当的汇总统计量达成了共识,了解它们的缺陷,并且学界知道如何解读它们。但在新兴领域,我们并不知道该使用哪些汇总统计量。我们很容易被那些含义并非如我们所想、或者掩盖了我们试图研究的核心复杂性的数字引入歧途。
因此,尽管成熟学科领域有充分理由高度依赖汇总统计量,并形成了鼓励使用它们的文化,但这种文化在前范式领域可能会产生误导。在可解释性研究中,我们需要对此保持警惕。
严谨性与结构的信号
我们如何判断定性结果是“真实的”?严谨的定性研究是什么样的?
我们没有“统计显著性”这类工具可以依赖,因此这类研究很容易显得不够严谨。然而,在显微镜下发现细胞无疑是严谨的!同样,发现恒星光谱、发现超导现象(“电阻是否为零?”)等等也是如此。科学史上一些最引人注目、改变世界的发现,正是以定性结果的形式出现的——它们不需要误差线或汇总统计量,因为它们本身就足够引人注目。
在某些已确立的研究领域中,可能存在已知的定性方法——例如在动物学中发现新物种——这类方法利用了我们真正理解所观察对象这一事实,从而做出严谨的定性观察。但这一点同样不适用于可解释性研究。
于是我们又回到了最初的问题——我们如何知道定性结果是真实的?我们推测,判断定性结果是否可信的最可靠方法之一,就是我们所说的"结构信号":
结构信号是指定性观察中出现的某种结构,这种结构不可能是测量伪影,也不可能来自其他来源,而必须反映研究对象本身的某种内在结构,即便我们并不理解它。
你可以将其视为统计显著性的一种非正式、"无监督"版本。统计显著性检验的是某个(最好是预先注册的)假设相对于零假设的显著性,而结构信号则观察到一个未曾预料的高维模式,并拒绝"这是噪声或伪影"的假设——通常是因为该结构如此引人注目且复杂,以至于明显超出了可接受的阈值好几个数量级。
结构信号的实例
对细胞的观察不可能是显微镜的伪影——它们太复杂了!也不可能是噪声——它们太有结构了!像镜头眩光这样的伪影可以产生畸变,但绝不会产生这样的图像:
胡克《显微图谱》中的细胞图像。图片来自威尔士国家图书馆。
同样,DeepDream 图像也过于复杂(且没有其他来源可以产生这种结构),因此它只能是网络内部某种结构的反映。即使你不知道这种结构是什么,它也不可能是随机噪声!
多张 DeepDream 图像,来自 Mordvintsev 等人最初的 Inceptionism 博客文章
另一个例子是 InceptionV1 中曲线检测器之间的权重。其模式过于复杂,不可能是噪声。我们可能对其解释存在争议,但显然其中存在某种结构!
曲线检测器之间权重的图像,来自最初的 circuits 系列文章。
请注意,所有这些都涉及结构极其复杂且细节丰富,远超任何偶然因素所能产生的范畴。它们如此引人注目,以至于即便是有选择性地挑选出来,也依然具有显著意义:胡克通过显微镜观察细胞,这显然揭示了真实存在的事物,即便他可能还观察了其他数百种毫无趣味且未作报告的对象。
之所以能够发现这一点,很大程度上是因为存在“低垂的果实”——一旦看到,便是显而易见的模式。在成熟的领域中,结构信号的明亮灯塔通常早已被找到,剩下的课题则需要小心翼翼地梳理剖析。但对于可解释性研究而言,仍有大量“低垂的果实”等待摘取。
附注:经典定量研究中的结构信号
在更偏定量化的研究工作中,同样会出现注意到这种显著结构的情况。例如,当你在对数-对数坐标图上观察时,发现神经网络具有如此清晰的缩放定律,这就是“结构信号”的一个实例——这种结构必定在向你传递某种信息!尤其值得注意的是,这类结构往往揭示了一种新的抽象概念。
关键在于,你需要确认这种结构并非源自其他来源。例如,某些显著性图方法可以生成引人注目、看似高度结构化的图像——但它们主要反映的是原始图像的结构,因此并不清楚它们是否真的向你展示了关于网络本身的显著结构。这并非是说显著性图一定没有展示任何东西,只是结构信号无法在这方面给予我们信心:我们需要其他基于原理的论证或实验。更普遍地说,许多有趣的现象都涉及多个相互作用的对象(例如应用于特定输入的神经网络),这些现象非常有趣,但在我们援引结构信号之前,需要针对每个来源评估其结构,并判断是否存在某种不那么有趣的解释。
定性研究的成功与失败
值得指出的是,在定性研究中,人们也极其容易自欺欺人。(这很可能也是科学家们常常对其持怀疑态度的另一个原因!)结构信号是我们所知的、发现真实现象的最简便方法,但它并非唯一的方法;而当它能与其他方法结合时,其说服力会更强。
优秀定性工作的一些标志包括:
- 结构信号。
- 对你所观察到的现象及其观察依据有原则性的理解。放大镜能让物体变大;而神经网络的权重则是其基本的计算基础。
我们通常希望定性结果至少能在上述某一方面表现出色(当然,越多越好)。反之,以下特征则会使我们对定性工作产生怀疑:
- 观察缺乏原则性。如果结果具有极其清晰的结构,并且不依赖于特定的解读,那么这一点是可以克服的。
- 结构不够引人注目(即缺乏结构信号),尤其是与选择性展示结果相结合时。除非我们对测量方法非常有信心,否则我们希望结构足够清晰,以至于不用担心看到的是错觉。
- 结构可能源自其他来源(例如,显著性图中的结构可能来自图像本身,而非模型)。
延伸阅读
《科学革命的结构》——尤其关注其中关于电学研究领域相互竞争的特定“学派”的讨论,以及它们如何聚焦于不同的现象和抽象概念;还包括早期原子物理学以及在云室中可视化粒子的内容。请注意其中的普遍模式:这通常始于相对定性的观察,如何选择关注点并将其具体化为抽象概念,这一过程非常核心,并为更定量的工作奠定了基础。一旦你识别出细胞,你就可以对它们进行计数。
《分离理论》——尤其关注第一章关于“纸质工具”的讨论。
《思考不可想象之事的媒介》——如果你觉得“界面”与研究之间存在关联这一概念很陌生,请观看这个演讲!如果你喜欢这个,还可以看看《在抽象阶梯上上下下》。
证明与反驳——研究的一个重要部分,就是在黑暗中摸索,寻找恰当的定义和抽象概念。拉卡托斯的这部戏剧精彩地展现了这一点。你可以将具体实例与定义之间的相互作用,类比为定性研究与汇总统计量的创建。我们需要通过具体实例来找到正确的定义/统计量。
货舱崇拜科学——理查德·费曼的一篇经典文章,探讨了科学中流于形式的严谨性。
脚注
- 请参见下方的“延伸阅读”,特别是《思考不可思考之事的媒介》(关于科学思维中界面的精彩讨论)和《分离理论》(关于“纸质工具”界面的讨论)。[↩]
Affiliations
Published
Not published yet.
DOI
No DOI yet.
Chris Olah, Adam Jermyn
This note offers some opinionated thoughts on why interpretability research may have qualitative aspects be more central than we're used to in other fields. It also aims to describe some heuristics for research taste in qualitative work.
Early scientific fields are often quite qualitative and become more quantitative as they mature. For example, discovering cells is a qualitative result, which can then mature (over many decades) into quantitative tools like counting white blood cells in cancer research. Discovering chemical spectral lines was a qualitative result, which only really became quantitative when Bohr realized that the "butterfly wings of atoms" gave insight into electron orbitals.
The vast majority of researchers are trained in mature disciplines, because genuinely new scientific fields are rare. These mature disciplines have established paradigms, with established quantitative measures and methods. But interpretability is not a mature field. It doesn't have an established paradigm. Even the most basic abstractions (does it make sense to think of a model in terms of "features"?) are up for debate.
There's a risk that our training from mature fields may give us the wrong instincts if we translate them into such an early, messy, unestablished science. In particular, we should expect to need to be guided a lot more by qualitative results.
To be clear, this isn't saying we should not do quantitative research when appropriate! And in fact, often these can be synergistic, with qualitative research helping us be confident we're using the right quantitative tools. (The line between them can also be blurry!) Rather, the goal of this note is simply to argue that qualitative results should genuinely be seen as first class citizens, and something we want to keep returning to as a touchstone to avoid becoming lost or fooling ourselves.
Summary Statistics and Their Dangers
Anscombe's Quartetis a famous example of how several radically different datasets can have the same mean, standard deviation, and correlation:
This reflects a more general lesson: summary statistics which boil rich high dimensional data into a single number will always blind you to most of what's going on. And so, you need to be very careful when you do so. And in particular, you need to be very careful that you know and trust what you're measuring if you're going to rely on it.
In established fields, there may be standard measurements and quantitative values that are very well understood – in what they mean, in how to think about them, and so on. But in interpretability we don't have that benefit.
A pattern we see in some interpretability and interpretability-adjacent ML papers is defining some metric which is claimed to correspond to some property of interest, and then very rigorously measuring this metric. We see this as a kind of Cargo-Cult Science. It can seem very rigorous with lots of line plots with standard deviation bars and such. But it often isn't, because the critical weakness is whether the metric actually, reliably tracks the property of interest, not the rigor with which the metric is evaluated. A recent example of this in our own work was our study of tanh-regularization in dictionary learning, where summary statistics initially indicated a lot of promise and it was only later qualitative inspection of features that revealed that we had been led astray.
We suspect that often, in early stage scientific fields like interpretability, rigor involves much more qualitative work, with quantitative metrics growing out of that and initially being treated quite skeptically.
For Want of Constrained Hypothesis Space
One of the reasons mature fields can rely so much on summary statistics is that they have a constrained hypothesis space.
Consider how mature fields often frame experiments as testing one hypothesis against another. For instance, a gravitational wave experiment might be set up to test General Relativity against Chern-Simons gravity. This makes sense when the hypothesis space is narrow and well-understood, where the bulk of the probability mass truly is on the few hypotheses being tested. In these cases it's often relatively straightforward to come up with a single measure that should be different between two hypotheses!
But in pre-paradigmatic fields we don’t know what hypotheses to consider! A physicist in the 1800s coming up with explanations for energy production in the Sun would have entirely missed nuclear physics.
So the goal of science in pre-paradigmatic fields is to first figure out what hypotheses we should be considering! And this means working in a rather different way. Whereas summary statistics can be very good for discriminating between a small number of hypotheses, individual numbers don’t provide a rich enough signal to orient us in a vast space of possibilities.
Interfaces and the Lure of Summary Statistics
Why are summary statistics so popular? One reason involves the ecosystem of interfaces for scientific research. We don't often think of things this way, but scientific research implicitly involves interfaces for thinking about data. For example equations, line plots, and terminology are all interfaces. See Additional Reading below, and especially Media for Thinking the Unthinkable(for compelling discussion of interfaces in scientific thinking) and Drawing Theories Apart(for discussion of "paper tool" interfaces).
The most common and reusable interfaces (eg. line plots) often require a scalar summary statistic. These simple datatypes are a kind of lingua franca of research because they're so common that they can be reused across almost scientific fields and have standardized practices. But they require us to reduce our results to a summary statistic.
Presenting more complex (and simply more massive) data requires custom interfaces. In fact, working with the scale of data we do in interpretability without reducing to summary statistics is almost impossible without custom and often interactive interfaces. This is why interpretability has been so intertwined with data visualization. Just as early chemistry depended on custom glassware – and indeed, many scientists did their own glasswork! – so too does interpretability depend on data visualization.
Summary Statistics and Defensibility
Another reason summary statistics are popular is that they are defensible.There is a broader scientific community and culture that implicitly tells us “If you present your data in one of several common formats, you are in the very worst case still doing science.”
In mature fields this makes sense: there has been convergence on the right summary statistics to use, their pitfalls are known, and the community knows how to interpret them. But in young fields we don’t know what summary statistics to use. And we can easily be led astray by numbers that don’t mean what we think, or that hide the core complexity we’re trying to study.
So while there are good reasons that mature fields lean heavily on summary statistics, and develop a culture that encourages them, that culture can be misleading in pre-paradigmatic fields, and in interpretability we need to be mindful of this.
Rigor and The Signal of Structure
How do we know if qualitative results are "real"? What does rigorous qualitative research look like?
We don't have tools like "statistical significance" to fall back on, and so it's easy for this research to seem non-rigorous. And yet the discovery of cells under a microscope was certainly rigorous! And likewise the discovery of stellar spectra, of superconductivity (“Is the resistance zero?”), and so on. Some of the most striking and world-changing discoveries in the history of science came in the form of qualitative results that didn’t need error bars or summary statistics because they were so striking.
In certain well established topics there might be known qualitative methods – for example noticing a new species in zoology – which leverage the fact that we really understand what we're observing to make a rigorous qualitative observation. But this, also, isn't applicable to interpretability.
So this brings us back to the original question – how can we know if qualitative results are real? We suspect that one of the most reliable ways to know that a qualitative result is trustworthy is what we'll call the signal of structure:
The signal of structure is any structure in one's qualitative observations which cannot be an artifact of measurement or have come from another source, but instead must reflect some kind of structure in the object of inquiry, even if we don't understand it.
You might think of this as the informal, "unsupervised" version of statistical significance. Whereas statistical significance tests a particular (hopefully pre-registered) hypothesis against a null hypothesis, the signal of structure observes an unpredicted high-dimensional pattern and rejects the hypothesis it was noise or an artifact, typically because the structure is so compelling and complex that it's clearly orders of magnitude past the bar.
Examples of the Signal of Structure
The observation of cells can't be an artifact of the microscope – they're too complex! And it can’t be a noise – they’re too structured! An artifact like a lens flare can produce distortions, but not ones like this:
A picture of cells from Hooke's Micrographia. Images from the National Library of Wales.
Likewise, DeepDreamis too complex (and has no other source structure could come from) to be anything other than a reflection of some structure inside the network. Even if you don't know what the structure is, it can't be random noise!
Multiple DeepDream Images, from Mordvintsev et al's original Inceptionism blog post
Another example of this is the weights between curve detectorsin InceptionV1. The pattern is too complex to be noise. We could argue over interpretation, but there's clearly something there!
A picture of weights between curve detectors, from the original circuits thread.
Note that all of these involve structure with overwhelming complexity and detail which is just far beyond anything that could happen by chance. They're so striking that they hold up as notable even if they're cherry picked: Hooke seeing cells through a microscope is just obviously showing something real, even if he looked at hundreds of other things that were uninteresting and didn't report on them.
A big part of the reason one can find this is because there's low hanging fruit – glaringly obvious patterns once one sees them. In mature fields, the bright beacons of the signal of structure will generally have already been found, and remaining lessons will involve trying to carefully tease things apart. But for interpretability, there's an incredible amount of low hanging fruit.
Aside: The Signal of Structure in Classical Quantitative Work
Noticing striking structure like this also happens in more quantitative work. For example, noticing that neural networks have such clean scaling laws when you look at a log-log plot is an example of the "signal of structure" – the structure must be telling you something! And in particular, such structure is often telling you about a new abstraction.
The critical thing is that you need to know the structure isn't coming from some other source. For example, certain saliency map methods can produce striking, seemingly highly structured images – but they're mostly reflecting the structure of the original image, so it's not clear they're really showing you striking structure about the network. This isn't to say that saliency maps are necessarily not showing something, just that the signal of structure can't give us confidence in this: we need some other principled argument or experiment. More generally, many interesting phenomena involve cases where there are multiple interacting objects (such as a neural network applied to a particular input), and these are very interesting, but the structure needs to be evaluated with respect to every source and whether some less interesting explanation is possible, before we can invoke the signal of structure.
The Success and Failure of Qualitative Research
It's worth noting that it's also exceedingly easy to fool oneself with qualitative research. (This is likely another reason why scientists are often skeptical of it!) The signal of structure is the easiest way we know of to find real phenomena, but it's not the only one, and when it can combine with others, it's even more compelling.
Some signs of good qualitative work include:
- The Signal of Structure.
- A principled understanding of what you're seeing and why it makes sense to look at.A magnifying glass makes things bigger; the weights of a neural network are the fundamental computational substrate of them.
We usually want to see qualitative results shine on at least one of these (more is better of course!). Conversely, the following traits make us suspicious of qualitative work:
- Looking at something unprincipled.This can be overcome if results have extremely clear structure, and don't depend on a specific interpretation.
- Structure which isn't extremely striking(that is, lacking the signal of structure), especially if combined with cherry picking. Unless we’re really confident in the measurement, we want there to be such clear structure that we’re not worried about seeing an illusion.
- Other sources the structure could have come from(eg. structure in saliency maps can come from the image rather than the model)
Additional Reading
The Structure of Scientific Revolutions– Especially the discussion of particular competing "schools" in research into electricity and how they focused on different phenomena and abstractions; and also early atomic physics and visualizing particles in cloud chambers. Note the general pattern of how this often starts with relatively qualitative observations, how the choice of what to focus attention on and reify into an abstraction is very central and enables more quantitative work. Once you recognize cells, you can count them.
Drawing Theories Apart– Especially chapter 1 discussion of "paper tools".
Media for Thinking the Unthinkable– If the idea of "interfaces" and research being linked is foreign, watch this talk! If you like this, consider also looking at Up and Down the Ladder of Abstraction.
Proofs and Refutations– An important part of research is stumbling around in the dark for the right definitions and abstractions. This play by Lakatos beautifully gets at this. You might see the interplay between concrete examples and definitions as analogous to qualitative investigation and the creation of summary statistics. We need to engage with examples to find the right definitions / statistics.
Cargo Cult Science– A classic essay by Richard Feynman about pro forma rigor in science.
Footnotes
- See Additional Reading below, and especially Media for Thinking the Unthinkable(for compelling discussion of interfaces in scientific thinking) and Drawing Theories Apart(for discussion of "paper tool" interfaces).[↩]