research-article

Linear Regression from Strategic Data Sources

Authors:

Nicolas Gast,

Stratis Ioannidis,

Patrick Loiseau,

Benjamin RoussillonAuthors Info & Claims

ACM Transactions on Economics and Computation (TEAC), Volume 8, Issue 2

Article No.: 10, Pages 1 - 24

https://doi.org/10.1145/3391436

Published: 13 May 2020 Publication History

Get Access

Abstract

Linear regression is a fundamental building block of statistical data analysis. It amounts to estimating the parameters of a linear model that maps input features to corresponding outputs. In the classical setting where the precision of each data point is fixed, the famous Aitken/Gauss-Markov theorem in statistics states that generalized least squares (GLS) is a so-called “Best Linear Unbiased Estimator” (BLUE). In modern data science, however, one often faces strategic data sources; namely, individuals who incur a cost for providing high-precision data. For instance, this is the case for personal data, whose revelation may affect an individual’s privacy—which can be modeled as a cost—or in applications such as recommender systems, where producing an accurate estimate entails effort.

In this article, we study a setting in which features are public but individuals choose the precision of the outputs they reveal to an analyst. We assume that the analyst performs linear regression on this dataset, and individuals benefit from the outcome of this estimation. We model this scenario as a game where individuals minimize a cost composed of two components: (a) an (agent-specific) disclosure cost for providing high-precision data; and (b) a (global) estimation cost representing the inaccuracy in the linear model estimate. In this game, the linear model estimate is a public good that benefits all individuals. We establish that this game has a unique non-trivial Nash equilibrium. We study the efficiency of this equilibrium and we prove tight bounds on the price of stability for a large class of disclosure and estimation costs. Finally, we study the estimator accuracy achieved at equilibrium. We show that, in general, Aitken’s theorem does not hold under strategic data sources, though it does hold if individuals have identical disclosure costs (up to a multiplicative factor). When individuals have non-identical costs, we derive a bound on the improvement of the equilibrium estimation cost that can be achieved by deviating from GLS, under mild assumptions on the disclosure cost functions.

References

[1]

Jacob Abernethy, Yiling Chen, Chien-Ju Ho, and Bo Waggoner. 2015. Low-cost learning via active data procurement. In Proceedings of the 16th ACM Conference on Economics and Computation (EC’15). 619--636.

Abstract

References

Cited By

Index Terms

Recommendations

Linear Regression as a Non-cooperative Game

The Effect of Strategic Noise in Linear Regression

The effect of strategic noise in linear regression

Comments

Information

Published In

Publisher

Publication History

Permissions

Check for updates

Author Tags

Qualifiers

Funding Sources

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Cited By

Get Access

Login options

Full Access

View options

PDF

eReader

HTML Format

Figures

Other

Share

Share this Publication link

Share on social media

Affiliations