The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions

Meyer, Anna P.; Albarghouthi, Aws; D'Antoni, Loris

Computer Science > Machine Learning

arXiv:2304.10655 (cs)

[Submitted on 20 Apr 2023]

Title:The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions

Authors:Anna P. Meyer, Aws Albarghouthi, Loris D'Antoni

View PDF

Abstract:We introduce dataset multiplicity, a way to study how inaccuracies, uncertainty, and social bias in training datasets impact test-time predictions. The dataset multiplicity framework asks a counterfactual question of what the set of resultant models (and associated test-time predictions) would be if we could somehow access all hypothetical, unbiased versions of the dataset. We discuss how to use this framework to encapsulate various sources of uncertainty in datasets' factualness, including systemic social bias, data collection practices, and noisy labels or features. We show how to exactly analyze the impacts of dataset multiplicity for a specific model architecture and type of uncertainty: linear models with label errors. Our empirical analysis shows that real-world datasets, under reasonable assumptions, contain many test samples whose predictions are affected by dataset multiplicity. Furthermore, the choice of domain-specific dataset multiplicity definition determines what samples are affected, and whether different demographic groups are disparately impacted. Finally, we discuss implications of dataset multiplicity for machine learning practice and research, including considerations for when model outcomes should not be trusted.

Comments:	25 pages, 8 figures. Accepted at FAccT '23
Subjects:	Machine Learning (cs.LG); Computers and Society (cs.CY)
Cite as:	arXiv:2304.10655 [cs.LG]
	(or arXiv:2304.10655v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2304.10655

Submission history

From: Anna Meyer [view email]
[v1] Thu, 20 Apr 2023 21:31:15 UTC (2,023 KB)

Computer Science > Machine Learning

Title:The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators