Rounding based continuous data discretization for statistical disclosure control
2023 (English)In: Journal of Ambient Intelligence and Humanized Computing, ISSN 1868-5137, E-ISSN 1868-5145, Vol. 14, no 11, p. 15139-15157Article in journal (Refereed) Published
Abstract [en]
“Rounding” can be understood as a way to coarsen continuous data. That is, low level and infrequent values are replaced by high-level and more frequent representative values. This concept is explored as a method for data privacy with techniques like rounding, microaggregation, and generalisation. This concept is explored as a method for data privacy in statistical disclosure control literature with perturbative techniques like rounding, microaggregation and non-perturbative methods like generalisation. Even though “rounding” is well known as a numerical data protection method, it has not been studied in depth or evaluated empirically to the best of our knowledge. This work is motivated by three objectives, (1) to study the alternative methods of obtaining the rounding values to represent a given continuous variable, (2) to empirically evaluate rounding as a data protection technique based on information loss (IL) and disclosure risk (DR), and (3) to analyse the impact of data rounding on machine learning based models. Here, in order to obtain the rounding values we consider discretization methods introduced in the unsupervised machine learning literature along with microaggregation and re-sampling based approaches. The results indicate that microaggregation based techniques are preferred over unsupervised discretization methods due to their fair trade-off between IL and DR.
Place, publisher, year, edition, pages
Springer, 2023. Vol. 14, no 11, p. 15139-15157
Keywords [en]
Micro data protection, Rounding for micro data, Unsupervised discretization, Discrete event simulation, Economic and social effects, Machine learning, Numerical methods, Volume measurement, Data protection techniques, Discretization method, Numerical data protection methods, Perturbative techniques, Statistical disclosure Control, Unsupervised machine learning, Data privacy
National Category
Computer Sciences
Research subject
Skövde Artificial Intelligence Lab (SAIL)
Identifiers
URN: urn:nbn:se:his:diva-17858DOI: 10.1007/s12652-019-01489-7Scopus ID: 2-s2.0-85074009425OAI: oai:DiVA.org:his-17858DiVA, id: diva2:1368632
Part of project
Disclosure risk and transparency in big data privacy, Swedish Research Council
Funder
Swedish Research Council, 2016-03346
Note
CC BY 4.0
Published: 25 September 2019
Correspondence to Navoda Senavirathne.
This work is supported by Vetenskapsrådet project: “Disclosure risk and transparency in big data privacy” (VR 2016-03346, 2017-2020)
DRIAT
2019-11-072019-11-072024-02-13Bibliographically approved