Synthetic data through combinatorial optimization of pairwise probabilities
The generation of synthetic data is a critical area of research in domains where real data are either not available in large quantities or cannot be directly used. Different techniques have been developed to produce high-quality, realistic synthetic datasets which retain the statistical properties o...
| Autores: | , , |
|---|---|
| Formato: | artículo |
| Estado: | Versión publicada |
| Fecha de publicación: | 2026 |
| País: | España |
| Recursos: | Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| Repositorio: | Recercat. Dipósit de la Recerca de Catalunya |
| OAI Identifier: | oai:dnet:recercat____::932433e207a31ed682424ed2ad9fef2a |
| Acesso em linha: | https://doi.org/10.1007/s41060-026-01063-3 https://hdl.handle.net/10459.1/470089 |
| Access Level: | acceso abierto |
| Palavra-chave: | Synthetic data Data mining Data exploitation |
| id |
ES_80f439db06cebfa8e9736f39ee1a1fb8 |
|---|---|
| oai_identifier_str |
oai:dnet:recercat____::932433e207a31ed682424ed2ad9fef2a |
| network_acronym_str |
ES |
| network_name_str |
España |
| repository_id_str |
|
| spelling |
Synthetic data through combinatorial optimization of pairwise probabilitiesSalvia Hornos, Josep M.Fernàndez Camon, CésarMateu Piñol, CarlesSynthetic dataData miningData exploitationThe generation of synthetic data is a critical area of research in domains where real data are either not available in large quantities or cannot be directly used. Different techniques have been developed to produce high-quality, realistic synthetic datasets which retain the statistical properties of the original data. State-of-the-art results focus on the use of neural networks to capture the latent space extracted from the data. While recent advances in the field of deep learning motivate the application in this new context, this paper proposes a naive baseline to generate synthetic data based on pairwise probabilities. We name the technique DISCO (Discrete Intersection Synthesizer through Combinatorial Optimization), a novel synthetic generator for tabular data. DISCO models data by optimizing the intersection of pairwise probabilities on each generated row in order to resemble the original dataset. Our approach preserves marginal (and pairwise) distributions and as a result, resembles the original data with high fidelity with a very simple approach. Evaluation on various synthetic and real-world datasets as well as regression and classification tasks prove DISCO’s ability to generate high-quality data that rivals state-of-the-art models in both statistical accuracy and machine learning efficacy.Open Access funding provided thanks to the CRUE-CSIC agreement with Springer Nature. Josep Maria Salvia Hornos reports funding from Generalitat de Catalunya AGAUR DI-2024-00020. Cèsar Fernández reports funding by the Spanish MCIN/AEI/10.130- 39/501100011033/, FEDER, UE in project IDs PID2022-138564OAI00 and PID2022-137971OB-I00. Carles Mateu reports that this work was partially funded by the Ministerio de Ciencia e Innovación - Agencia Estatal de Investigación (AEI) (PID2021-123511OB-C31- MCIN/AEI/10.13039/501100011033/ FEDER, UE and RED2022-134219-T).Springer2026info:eu-repo/semantics/articleinfo:eu-repo/semantics/publishedVersionhttps://doi.org/10.1007/s41060-026-01063-3https://hdl.handle.net/10459.1/470089https://hdl.handle.net/10459.1/470089reponame:Recercat. Dipósit de la Recerca de Catalunyainstname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya)Inglésinfo:eu-repo/grantAgreement/AEI//PID2022-138564OA-I00info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2022-137971OB-I00info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2021-123511OB-C31info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/RED2022-134219-TReproducció del document publicat a https://doi.org/10.1007/s41060-026-01063-3International Journal of Data Science and Analytics, 2026, vol. 22, 154cc-by (c) Josep Maria Salvia Hornos, Cèsar Fernández Camón, Carles Mateu Piñol, 2026Attribution 4.0 Internationalinfo:eu-repo/semantics/openAccesshttp://creativecommons.org/licenses/by/4.0/oai:dnet:recercat____::932433e207a31ed682424ed2ad9fef2a2026-05-29T05:05:01Z |
| dc.title.none.fl_str_mv |
Synthetic data through combinatorial optimization of pairwise probabilities |
| title |
Synthetic data through combinatorial optimization of pairwise probabilities |
| spellingShingle |
Synthetic data through combinatorial optimization of pairwise probabilities Salvia Hornos, Josep M. Synthetic data Data mining Data exploitation |
| title_short |
Synthetic data through combinatorial optimization of pairwise probabilities |
| title_full |
Synthetic data through combinatorial optimization of pairwise probabilities |
| title_fullStr |
Synthetic data through combinatorial optimization of pairwise probabilities |
| title_full_unstemmed |
Synthetic data through combinatorial optimization of pairwise probabilities |
| title_sort |
Synthetic data through combinatorial optimization of pairwise probabilities |
| dc.creator.none.fl_str_mv |
Salvia Hornos, Josep M. Fernàndez Camon, César Mateu Piñol, Carles |
| author |
Salvia Hornos, Josep M. |
| author_facet |
Salvia Hornos, Josep M. Fernàndez Camon, César Mateu Piñol, Carles |
| author_role |
author |
| author2 |
Fernàndez Camon, César Mateu Piñol, Carles |
| author2_role |
author author |
| dc.subject.none.fl_str_mv |
Synthetic data Data mining Data exploitation |
| topic |
Synthetic data Data mining Data exploitation |
| description |
The generation of synthetic data is a critical area of research in domains where real data are either not available in large quantities or cannot be directly used. Different techniques have been developed to produce high-quality, realistic synthetic datasets which retain the statistical properties of the original data. State-of-the-art results focus on the use of neural networks to capture the latent space extracted from the data. While recent advances in the field of deep learning motivate the application in this new context, this paper proposes a naive baseline to generate synthetic data based on pairwise probabilities. We name the technique DISCO (Discrete Intersection Synthesizer through Combinatorial Optimization), a novel synthetic generator for tabular data. DISCO models data by optimizing the intersection of pairwise probabilities on each generated row in order to resemble the original dataset. Our approach preserves marginal (and pairwise) distributions and as a result, resembles the original data with high fidelity with a very simple approach. Evaluation on various synthetic and real-world datasets as well as regression and classification tasks prove DISCO’s ability to generate high-quality data that rivals state-of-the-art models in both statistical accuracy and machine learning efficacy. |
| publishDate |
2026 |
| dc.date.none.fl_str_mv |
2026 |
| dc.type.none.fl_str_mv |
info:eu-repo/semantics/article info:eu-repo/semantics/publishedVersion |
| format |
article |
| status_str |
publishedVersion |
| dc.identifier.none.fl_str_mv |
https://doi.org/10.1007/s41060-026-01063-3 https://hdl.handle.net/10459.1/470089 https://hdl.handle.net/10459.1/470089 |
| url |
https://doi.org/10.1007/s41060-026-01063-3 https://hdl.handle.net/10459.1/470089 |
| dc.language.none.fl_str_mv |
Inglés |
| language_invalid_str_mv |
Inglés |
| dc.relation.none.fl_str_mv |
info:eu-repo/grantAgreement/AEI//PID2022-138564OA-I00 info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2022-137971OB-I00 info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2021-123511OB-C31 info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/RED2022-134219-T Reproducció del document publicat a https://doi.org/10.1007/s41060-026-01063-3 International Journal of Data Science and Analytics, 2026, vol. 22, 154 |
| dc.rights.none.fl_str_mv |
cc-by (c) Josep Maria Salvia Hornos, Cèsar Fernández Camón, Carles Mateu Piñol, 2026 Attribution 4.0 International info:eu-repo/semantics/openAccess http://creativecommons.org/licenses/by/4.0/ |
| rights_invalid_str_mv |
cc-by (c) Josep Maria Salvia Hornos, Cèsar Fernández Camón, Carles Mateu Piñol, 2026 Attribution 4.0 International http://creativecommons.org/licenses/by/4.0/ |
| eu_rights_str_mv |
openAccess |
| dc.publisher.none.fl_str_mv |
Springer |
| publisher.none.fl_str_mv |
Springer |
| dc.source.none.fl_str_mv |
reponame:Recercat. Dipósit de la Recerca de Catalunya instname:Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| instname_str |
Varias* (Consorci de Biblioteques Universitáries de Catalunya, Centre de Serveis Científics i Acadèmics de Catalunya) |
| reponame_str |
Recercat. Dipósit de la Recerca de Catalunya |
| collection |
Recercat. Dipósit de la Recerca de Catalunya |
| repository.name.fl_str_mv |
|
| repository.mail.fl_str_mv |
|
| _version_ |
1869411938168995840 |
| score |
15,812455 |