Dimensionality reduction via the Johnson-Lindenstrauss Lemma: theoretical and empirical bounds on embedding dimension

Fedoruk John; Schmuland Byron<sup>*</sup>; Johnson Julia; Heo Giseon

doi:10.1007/s11227-018-2401-y

摘要

The Johnson-Lindenstrauss (JL) lemma has led to the development of tools for dealing with datasets in high dimensions. The lemma asserts that a set of high-dimensional points can be projected into lower dimensions, while approximately preserving the pairwise distance structure. Significant improvements of the JL lemma since its inception are summarized. Particular focus is placed on reproving Matouek's versions of the lemma (Random Struct Algorithms 33(2):142-156, 2008) first using subgaussian projection coefficients and then using sparse projection coefficients. The results of the lemma are illustrated using simulated data. The simulation suggests a projection that is more effective in terms of dimensionality reduction than is borne out by the theory. This more effective projection was applied to a very large natural, rather than simulated, dataset thus further strengthening empirical evidence of the existence of a better than the proven optimal lower bound on the embedding dimension. Additionally, we provide comparisons with other commonly used data reduction and simplification techniques.

出版日期2018-8

全文

访问全文

收藏分享被引(2) 浏览

更新时间：2021-03-15 23:18

Dimensionality reduction via the Johnson-Lindenstrauss Lemma: theoretical and empirical bounds on embedding dimension

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友