From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence

ML NNGTB SSLRO Generative Models
我们能否从数据中学到比其生成过程本身所包含的更多的信息?仅仅通过对现有数据进行确定性变换,是否就能构造出新的、有用的信息?在不考虑下游任务的情况下,我们能否评估数据中可学习的内容?对于这些问题,香农信息论和柯尔莫哥洛夫复杂性几乎无法提供答案,部分原因在于它们假设观察者具有无限的计算能力,因而未能聚焦于真正有用的信息内容。本文中,我们指出并举例说明了信息论中的三个看似矛盾的现象:(1)确定性变换无法增加信息;(2)信息与数据的顺序无关;(3)似然建模仅仅是分布匹配。为了阐明这些理论结果与现代实践之间的张力,并量化数据的价值,我们提出了“表知性”(epiplexity)这一概念,用以形式化地刻画计算资源受限的观察者能够从数据中学到的内容。表知性能够捕捉数据中的结构性信息,同时排除时间受限熵——即由伪随机数生成器和混沌动力系统等所体现的不可预测的随机成分。借助这些概念,我们展示了信息如何通过计算被创造出来,如何依赖于数据的排列顺序,以及似然建模如何生成比原始数据生成过程本身更为复杂的程序。我们还提出了估计表知性的实用方法,实验表明这些方法能够捕捉不同数据源之间的差异,与下游任务性能变化保持一致,并凸显出有助于提升分布外泛化能力的数据集干预措施。与模型选择的原则不同,表知性为数据选择提供了理论基础,指导我们应如何为学习系统选择、生成或转换数据。
Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to existing data? Can the learnable content in data be evaluated without considering a downstream task? On these questions, Shannon information and Kolmogorov complexity come up nearly empty-handed, in part because they assume observers with unlimited computational capacity and fail to target the useful information content. In this work, we identify and exemplify three seeming paradoxes in information theory: (1) information cannot be increased by deterministic transformations; (2) information is independent of the order of data; (3) likelihood modeling is merely distribution matching. To shed light on the tension between these results and modern practice, and to quantify the value of data, we introduce epiplexity, a formalization of information capturing what computationally bounded observers can learn from data. Epiplexity captures the structural content in data while excluding time-bounded entropy, the random unpredictable content exemplified by pseudorandom number generators and chaotic dynamical systems. With these concepts, we demonstrate how information can be created with computation, how it depends on the ordering of the data, and how likelihood modeling can produce more complex programs than present in the data generating process itself. We also present practical procedures to estimate epiplexity which we show capture differences across data sources, track with downstream performance, and highlight dataset interventions that improve out-of-distribution generalization. In contrast to principles of model selection, epiplexity provides a theoretical foundation for data selection, guiding how to select, generate, or transform data for learning systems.
许愿