VC-dimension of 𝑘-fold unions: General case

AbstractThe Vapnik–Chervonenkis dimension provides a notion of complexity for systems of sets. If the VC dimension is small, then knowing this can drastically simplify fundamental computational tasks such as classification, range counting, and density estimation through the use of sampling bounds. We analyze set systems where the ground set X is a set of polygonal curves in $$\mathbb {R}^d$$ R d and the sets $$\mathcal {R}$$ R are metric balls defined by curve similarity metrics, such as the Fréchet distance and the Hausdorff distance, as well as their discrete counterparts. We derive upper and lower bounds on the VC dimension that imply useful sampling bounds in the setting that the number of curves is large, but the complexity of the individual curves is small. Our upper and lower bounds are either near-quadratic or near-linear in the complexity of the curves that define the ranges and they are logarithmic in the complexity of the curves that define the ground set.

Download Full-text

Why concept lattices are large: extremal theory for generators, concepts, and VC-dimension

International Journal of General Systems ◽

10.1080/03081079.2017.1354798 ◽

2017 ◽

Vol 46 (5) ◽

pp. 440-457 ◽

Cited By ~ 2

Author(s):

Alexandre Albano ◽

Bogdan Chornomaz

Keyword(s):

Vc Dimension ◽

Concept Lattices ◽

Extremal Theory

Download Full-text

Neural Nets with Superlinear VC-Dimension

ICANN ’94 ◽

10.1007/978-1-4471-2097-1_136 ◽

1994 ◽

pp. 581-584 ◽

Cited By ~ 1

Author(s):

Wolfgang Maass

Keyword(s):

Neural Nets ◽

Vc Dimension

Download Full-text

Some Properties of Infinite VC-Dimension Systems Alexey Chervonenkis

Statistical Learning and Data Science ◽

10.1201/b11429-9 ◽

2011 ◽

pp. 69-76

Keyword(s):

Vc Dimension

Download Full-text

On the Complexity of Learning a Class Ratio from Unlabeled Data

Journal of Artificial Intelligence Research ◽

10.1613/jair.1.12013 ◽

2020 ◽

Vol 69 ◽

Author(s):

Benjamin Fish ◽

Lev Reyzin

Keyword(s):

Computational Complexity ◽

Unlabeled Data ◽

Training Data ◽

Pac Learning ◽

Vc Dimension ◽

Standard Set

In the problem of learning a class ratio from unlabeled data, which we call CR learning, the training data is unlabeled, and only the ratios, or proportions, of examples receiving each label are given. The goal is to learn a hypothesis that predicts the proportions of labels on the distribution underlying the sample. This model of learning is applicable to a wide variety of settings, including predicting the number of votes for candidates in political elections from polls. In this paper, we formally define this class and resolve foundational questions regarding the computational complexity of CR learning and characterize its relationship to PAC learning. Among our results, we show, perhaps surprisingly, that for finite VC classes what can be efficiently CR learned is a strict subset of what can be learned efficiently in PAC, under standard complexity assumptions. We also show that there exist classes of functions whose CR learnability is independent of ZFC, the standard set theoretic axioms. This implies that CR learning cannot be easily characterized (like PAC by VC dimension).

Download Full-text

Reasoning about Measures of Unmeasurable Sets

Proceedings of the Seventeenth International Conference on Principles of Knowledge Representation and Reasoning ◽

10.24963/kr.2020/27 ◽

2020 ◽

Author(s):

Marco Console ◽

Matthias Hofer ◽

Leonid Libkin

Keyword(s):

Asymptotic Behavior ◽

Unit Sphere ◽

Ad Hoc ◽

Discrete Analog ◽

Uniform Measure ◽

Entire Space ◽

Vc Dimension ◽

Know How ◽

First Order ◽

Specific Subset

In a variety of reasoning tasks, one estimates the likelihood of events by means of volumes of sets they define. Such sets need to be measurable, which is usually achieved by putting bounds, sometimes ad hoc, on them. We address the question how unbounded or unmeasurable sets can be measured nonetheless. Intuitively, we want to know how likely a randomly chosen point is to be in a given set, even in the absence of a uniform distribution over the entire space. To address this, we follow a recently proposed approach of taking intersection of a set with balls of increasing radius, and defining the measure by means of the asymptotic behavior of the proportion of such balls taken by the set. We show that this approach works for every set definable in first-order logic with the usual arithmetic over the reals (addition, multiplication, exponentiation, etc.), and every uniform measure over the space, of which the usual Lebesgue measure (area, volume, etc.) is an example. In fact we establish a correspondence between the good asymptotic behavior and the finiteness of the VC dimension of definable families of sets. Towards computing the measure thus defined, we show how to avoid the asymptotics and characterize it via a specific subset of the unit sphere. Using definability of this set, and known techniques for sampling from the unit sphere, we give two algorithms for estimating our measure of unbounded unmeasurable sets, with deterministic and probabilistic guarantees, the latter being more efficient. Finally we show that a discrete analog of this measure exists and is similarly well-behaved.

Download Full-text