Beyond multivariate microaggregation for large record anonymization

Microaggregation is one of the most commonly employed microdata protection methods. The basic idea of microaggregation is to anonymize data by aggregating original records into small groups of at least k elements and, therefore, preserving k -anonymity. Usually, in order to avoid information loss, w...

Full description

Bibliographic Details
Author: Nin Guerrero, Jordi|||0000-0002-9659-2762
Format: article
Publication Date:2014
Country:España
Institution:Universitat Politècnica de Catalunya (UPC)
Repository:UPCommons. Portal del coneixement obert de la UPC
Language:English
OAI Identifier:oai:upcommons.upc.edu:2117/23297
Online Access:https://hdl.handle.net/2117/23297
https://dx.doi.org/10.1007/978-3-319-04178-0_8
Access Level:Open access
Keyword:Data protection
Database security
Microaggregation
k-anonymity
Privacy in statistical databases
Protecció de dades
Bases de dades -- Seguretat
Àrees temàtiques de la UPC::Informàtica::Seguretat informàtica
id ES_2e3e7be38ca47127465aa3bc2ac5b9eb
oai_identifier_str oai:upcommons.upc.edu:2117/23297
network_acronym_str ES
network_name_str España
repository_id_str
spelling Beyond multivariate microaggregation for large record anonymizationNin Guerrero, Jordi|||0000-0002-9659-2762Data protectionDatabase securityMicroaggregationk-anonymityPrivacy in statistical databasesProtecció de dadesBases de dades -- SeguretatÀrees temàtiques de la UPC::Informàtica::Seguretat informàticaMicroaggregation is one of the most commonly employed microdata protection methods. The basic idea of microaggregation is to anonymize data by aggregating original records into small groups of at least k elements and, therefore, preserving k -anonymity. Usually, in order to avoid information loss, when records are large, i.e., the number of attributes of the data set is large, this data set is split into smaller blocks of attributes and microaggregation is applied to each block, successively and independently. This is called multivariate microaggregation. By using this technique, the information loss after collapsing several values to the centroid of their group is reduced. Unfortunately, with multivariate microaggregation, the k -anonymity property is lost when at least two attributes of different blocks are known by the intruder, which might be the usual case. In this work, we present a new microaggregation method called one dimension microaggregation ( Mic1D-k ). With Mic1D-k , the problem of k -anonymity loss is mitigated by mixing all the values in the original microdata file into a single non-attributed data set using a set of simple pre-processing steps and then, microaggregating all the mixed values together. Our experiments show that, using real data, our proposal obtains lower disclosure risk than previous approaches whereas the information loss is preserved.20142014-01-0120142014-06-25journal articlehttp://purl.org/coar/resource_type/c_6501AMhttp://purl.org/coar/version/c_ab4af688f83e57aainfo:eu-repo/semantics/articleapplication/pdfhttps://hdl.handle.net/2117/23297https://dx.doi.org/10.1007/978-3-319-04178-0_8reponame:UPCommons. Portal del coneixement obert de la UPCinstname:Universitat Politècnica de Catalunya (UPC)Inglésengopen accesshttp://purl.org/coar/access_right/c_abf2info:eu-repo/semantics/openAccessoai:upcommons.upc.edu:2117/232972026-05-27T15:37:01Z
dc.title.none.fl_str_mv Beyond multivariate microaggregation for large record anonymization
title Beyond multivariate microaggregation for large record anonymization
spellingShingle Beyond multivariate microaggregation for large record anonymization
Nin Guerrero, Jordi|||0000-0002-9659-2762
Data protection
Database security
Microaggregation
k-anonymity
Privacy in statistical databases
Protecció de dades
Bases de dades -- Seguretat
Àrees temàtiques de la UPC::Informàtica::Seguretat informàtica
title_short Beyond multivariate microaggregation for large record anonymization
title_full Beyond multivariate microaggregation for large record anonymization
title_fullStr Beyond multivariate microaggregation for large record anonymization
title_full_unstemmed Beyond multivariate microaggregation for large record anonymization
title_sort Beyond multivariate microaggregation for large record anonymization
dc.creator.none.fl_str_mv Nin Guerrero, Jordi|||0000-0002-9659-2762
author Nin Guerrero, Jordi|||0000-0002-9659-2762
author_facet Nin Guerrero, Jordi|||0000-0002-9659-2762
author_role author
dc.subject.none.fl_str_mv Data protection
Database security
Microaggregation
k-anonymity
Privacy in statistical databases
Protecció de dades
Bases de dades -- Seguretat
Àrees temàtiques de la UPC::Informàtica::Seguretat informàtica
topic Data protection
Database security
Microaggregation
k-anonymity
Privacy in statistical databases
Protecció de dades
Bases de dades -- Seguretat
Àrees temàtiques de la UPC::Informàtica::Seguretat informàtica
description Microaggregation is one of the most commonly employed microdata protection methods. The basic idea of microaggregation is to anonymize data by aggregating original records into small groups of at least k elements and, therefore, preserving k -anonymity. Usually, in order to avoid information loss, when records are large, i.e., the number of attributes of the data set is large, this data set is split into smaller blocks of attributes and microaggregation is applied to each block, successively and independently. This is called multivariate microaggregation. By using this technique, the information loss after collapsing several values to the centroid of their group is reduced. Unfortunately, with multivariate microaggregation, the k -anonymity property is lost when at least two attributes of different blocks are known by the intruder, which might be the usual case. In this work, we present a new microaggregation method called one dimension microaggregation ( Mic1D-k ). With Mic1D-k , the problem of k -anonymity loss is mitigated by mixing all the values in the original microdata file into a single non-attributed data set using a set of simple pre-processing steps and then, microaggregating all the mixed values together. Our experiments show that, using real data, our proposal obtains lower disclosure risk than previous approaches whereas the information loss is preserved.
publishDate 2014
dc.date.none.fl_str_mv 2014
2014-01-01
2014
2014-06-25
dc.type.none.fl_str_mv journal article
http://purl.org/coar/resource_type/c_6501
AM
http://purl.org/coar/version/c_ab4af688f83e57aa
dc.type.openaire.fl_str_mv info:eu-repo/semantics/article
format article
dc.identifier.none.fl_str_mv https://hdl.handle.net/2117/23297
https://dx.doi.org/10.1007/978-3-319-04178-0_8
url https://hdl.handle.net/2117/23297
https://dx.doi.org/10.1007/978-3-319-04178-0_8
dc.language.none.fl_str_mv Inglés
eng
language_invalid_str_mv Inglés
language eng
dc.rights.none.fl_str_mv open access
http://purl.org/coar/access_right/c_abf2
dc.rights.openaire.fl_str_mv info:eu-repo/semantics/openAccess
rights_invalid_str_mv open access
http://purl.org/coar/access_right/c_abf2
eu_rights_str_mv openAccess
dc.format.none.fl_str_mv application/pdf
dc.source.none.fl_str_mv reponame:UPCommons. Portal del coneixement obert de la UPC
instname:Universitat Politècnica de Catalunya (UPC)
instname_str Universitat Politècnica de Catalunya (UPC)
reponame_str UPCommons. Portal del coneixement obert de la UPC
collection UPCommons. Portal del coneixement obert de la UPC
repository.name.fl_str_mv
repository.mail.fl_str_mv
_version_ 1869405391175024640
score 15,301603