Papers
Topics
Authors
Recent
Search
2000 character limit reached

Convergence of Graph Laplacian with kNN Self-tuned Kernels

Published 3 Nov 2020 in math.ST, cs.LG, stat.ML, and stat.TH | (2011.01479v2)

Abstract: Kernelized Gram matrix WW constructed from data points xi<em>i=1<sup>N{x_i}<em>{i=1}<sup>N as W</em>ij=k0(xixj<sup>2</sup>σ<sup>2</sup>)W</em>{ij}= k_0( \frac{ | x_i - x_j |<sup>2}</sup> {\sigma<sup>2}</sup> ) is widely used in graph-based geometric data analysis and unsupervised learning. An important question is how to choose the kernel bandwidth σ\sigma, and a common practice called self-tuned kernel adaptively sets a σi\sigma_i at each point xix_i by the kk-nearest neighbor (kNN) distance. When xix_i's are sampled from a dd-dimensional manifold embedded in a possibly high-dimensional space, unlike with fixed-bandwidth kernels, theoretical results of graph Laplacian convergence with self-tuned kernels have been incomplete. This paper proves the convergence of graph Laplacian operator LNL_N to manifold (weighted-)Laplacian for a new family of kNN self-tuned kernels W<sup>(α)ij</sup>=k0(xixj<sup>2</sup>ϵρ^(xi)ρ^(xj))/ρ^(xi)<sup>α</sup>ρ^(xj)<sup>αW<sup>{(\alpha)}_{ij}</sup> = k_0( \frac{ | x_i - x_j |<sup>2}{</sup> \epsilon \hat{\rho}(x_i) \hat{\rho}(x_j)})/\hat{\rho}(x_i)<sup>\alpha</sup> \hat{\rho}(x_j)<sup>\alpha, where ρ^\hat{\rho} is the estimated bandwidth function {by kNN}, and the limiting operator is also parametrized by α\alpha. When α=1\alpha = 1, the limiting operator is the weighted manifold Laplacian Δp\Delta_p. Specifically, we prove the point-wise convergence of LNfL_N f and convergence of the graph Dirichlet form with rates. Our analysis is based on first establishing a C<sup>0C<sup>0 consistency for ρ^\hat{\rho} which bounds the relative estimation error ρ^ρˉ/ρˉ|\hat{\rho} - \bar{\rho}|/\bar{\rho} uniformly with high probability, where ρˉ=p<sup>1/d\bar{\rho} = p<sup>{-1/d}, and pp is the data density function. Our theoretical results reveal the advantage of self-tuned kernel over fixed-bandwidth kernel via smaller variance error in low-density regions. In the algorithm, no prior knowledge of dd or data density is needed. The theoretical results are supported by numerical experiments on simulated data and hand-written digit image data.

Authors (2)
Citations (19)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.