Showing 1–2 of 2 results for author: Francis, B J

Search v0.5.6 released 2020-02-24

arXiv:2205.05993 [pdf, other]

stat.ME

On integrating the number of synthetic data sets $m$ into the 'a priori' synthesis approach

Authors: James Edward Jackson, Robin Mitra, Brian Joseph Francis, Iain Dove

Abstract: Until recently, multiple synthetic data sets were always released to analysts, to allow valid inferences to be obtained. However, under certain conditions - including when saturated count models are used to synthesize categorical data - single imputation ($m=1$) is sufficient. Nevertheless, increasing $m$ causes utility to improve, but at the expense of higher risk, an example of the risk-utility… ▽ More Until recently, multiple synthetic data sets were always released to analysts, to allow valid inferences to be obtained. However, under certain conditions - including when saturated count models are used to synthesize categorical data - single imputation ($m=1$) is sufficient. Nevertheless, increasing $m$ causes utility to improve, but at the expense of higher risk, an example of the risk-utility trade-off. The question, therefore, is: which value of $m$ is optimal with respect to the risk-utility trade-off? Moreover, the paper considers two ways of analysing categorical data sets: as they have a contingency table representation, multiple categorical data sets can be averaged before being analysed, as opposed to the usual way of averaging post-analysis. This paper also introduces a pair of metrics, $τ_3(k,d)$ and $τ_4(k,d)$, that are suited for assessing disclosure risk in multiple categorical synthetic data sets. Finally, the synthesis methods are demonstrated empirically. △ Less

Submitted 12 May, 2022; originally announced May 2022.
arXiv:2107.08062 [pdf, other]

stat.ME

Using saturated count models for user-friendly synthesis of categorical data

Authors: James Edward Jackson, Robin Mitra, Brian Joseph Francis, Iain Dove

Abstract: Over the past three decades, synthetic data methods for statistical disclosure control have continually evolved, but mainly within the domain of survey data sets. There are certain characteristics of administrative databases, such as their size, which present challenges from a synthesis perspective and require special attention. This paper, through the fitting of saturated count models, presents a… ▽ More Over the past three decades, synthetic data methods for statistical disclosure control have continually evolved, but mainly within the domain of survey data sets. There are certain characteristics of administrative databases, such as their size, which present challenges from a synthesis perspective and require special attention. This paper, through the fitting of saturated count models, presents a synthesis method that is suitable for administrative databases that is tuned by two parameters. The method allows large categorical data sets to be synthesized quickly and allows risk and utility metrics to be satisfied a priori, that is, prior to synthetic data generation. The paper explores how the flexibility afforded by two-parameter count models (the negative binomial and Poisson-inverse Gaussian) can be utilised to protect respondents' - especially uniques' - privacy in synthetic data. Finally, an empirical example is carried out through the synthesis of a database which can be viewed as a good substitute to the English School Census. △ Less

Submitted 12 May, 2022; v1 submitted 16 July, 2021; originally announced July 2021.

Comments: 37 pages, 6 figures

Search v0.5.6 released 2020-02-24