Published 18 Apr 2016 in cs.LG, cs.IT, math.IT, and stat.ML | (1604.05307v1)
Abstract: A function f:R<sup>d</sup>→R is referred to as a Sparse Additive Model (SPAM), if it is of the form f(x)=l∈S∑​ϕl​(xl​), where S⊂[d], ∣S∣≪d. Assuming ϕl​'s and S to be unknown, the problem of estimating f from its samples has been studied extensively. In this work, we consider a generalized SPAM, allowing for second order interaction terms. For some S<em>1⊂[d],S2​⊂(2[d]​), the function f is assumed to be of the form: f(x)=∑</em>p∈S<em>1ϕ</em>p(xp​)+(l,l<sup>′)</sup>∈S<em>2∑​ϕ</em>(l,l<sup>′)</sup>(xl​,xl<sup>′​). Assuming ϕp​,ϕ(l,l<sup>′)​, S<em>1 and, S2​ to be unknown, we provide a randomized algorithm that queries f and exactly recovers S1​,S2​. Consequently, this also enables us to estimate the underlying ϕp​,ϕ</em>(l,l<sup>′). We derive sample complexity bounds for our scheme and also extend our analysis to include the situation where the queries are corrupted with noise -- either stochastic, or arbitrary but bounded. Lastly, we provide simulation results on synthetic data, that validate our theoretical findings.