Papers
Topics
Authors
Recent
Search
2000 character limit reached

Self-training Converts Weak Learners to Strong Learners in Mixture Models

Published 25 Jun 2021 in cs.LG, math.OC, and stat.ML | (2106.13805v3)

Abstract: We consider a binary classification problem when the data comes from a mixture of two rotationally symmetric distributions satisfying concentration and anti-concentration properties enjoyed by log-concave distributions among others. We show that there exists a universal constant $C_{\mathrm{err}}&gt;0$ such that if a pseudolabeler β<em>pl\boldsymbol{\beta}<em>{\mathrm{pl}} can achieve classification error at most C</em>errC</em>{\mathrm{err}}, then for any $\varepsilon&gt;0$, an iterative self-training algorithm initialized at β<em>0:=β</em>pl\boldsymbol{\beta}<em>0 := \boldsymbol{\beta}</em>{\mathrm{pl}} using pseudolabels y^=sgn(⟨β<em>t,x⟩)\hat y = \mathrm{sgn}(\langle \boldsymbol{\beta}<em>t, \mathbf{x}\rangle) and using at most O~(d/ε<sup>2)\tilde O(d/\varepsilon<sup>2) unlabeled examples suffices to learn the Bayes-optimal classifier up to ε\varepsilon error, where dd is the ambient dimension. That is, self-training converts weak learners to strong learners using only unlabeled examples. We additionally show that by running gradient descent on the logistic loss one can obtain a pseudolabeler β</em>pl\boldsymbol{\beta}</em>{\mathrm{pl}} with classification error CerrC_{\mathrm{err}} using only O(d)O(d) labeled examples (i.e., independent of ε\varepsilon). Together our results imply that mixture models can be learned to within ε\varepsilon of the Bayes-optimal accuracy using at most O(d)O(d) labeled examples and O~(d/ε<sup>2)\tilde O(d/\varepsilon<sup>2) unlabeled examples by way of a semi-supervised self-training algorithm.

Citations (15)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 8 likes about this paper.