Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Limits of Assumption-free Tests for Algorithm Performance

Published 12 Feb 2024 in math.ST, cs.LG, stat.ML, and stat.TH | (2402.07388v3)

Abstract: Algorithm evaluation and comparison are fundamental questions in machine learning and statistics -- how well does an algorithm perform at a given modeling task, and which algorithm performs best? Many methods have been developed to assess algorithm performance, often based around cross-validation type strategies, retraining the algorithm of interest on different subsets of the data and assessing its performance on the held-out data points. Despite the broad use of such procedures, the theoretical properties of these methods are not yet fully understood. In this work, we explore some fundamental limits for answering these questions with limited amounts of data. In particular, we make a distinction between two questions: how good is an algorithm AA at the problem of learning from a training set of size nn, versus, how good is a particular fitted model produced by running AA on a particular training data set of size nn? Our main results prove that, for any test that treats the algorithm AA as a ``black box'' (i.e., we can only study the behavior of AA empirically), there is a fundamental limit on our ability to carry out inference on the performance of AA, unless the number of available data points NN is many times larger than the sample size nn of interest. (On the other hand, evaluating the performance of a particular fitted model is easy as long as a holdout data set is available -- that is, as long as N−nN-n is not too small.) We also ask whether an assumption of algorithmic stability might be sufficient to circumvent this hardness result. Surprisingly, we find that this is not the case: the same hardness result still holds for the problem of evaluating the performance of AA, aside from a high-stability regime where fitted models are essentially nonrandom. Finally, we also establish similar hardness results for the problem of comparing multiple algorithms.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 2 likes about this paper.