Automatic error localisation for categorical, continuous and integer data

Data collected by statistical offices generally contain errors, which have to be corrected before reliable data can be published. This correction process is referred to as statistical data editing. At statistical offices, certain rules, so-called edits, are often used during the editing process to d...

Descripción completa

Detalles Bibliográficos
Autor: Waal, Ton de
Tipo de recurso: artículo
Fecha de publicación:2005
País:España
Institución:Universitat Politècnica de Catalunya (UPC)
Repositorio:UPCommons. Portal del coneixement obert de la UPC
Idioma:inglés
OAI Identifier:oai:upcommons.upc.edu:2099/3757
Acceso en línea:https://hdl.handle.net/2099/3757
Access Level:acceso abierto
Palabra clave:Mathematical logic
Statistics
Artificial intelligence
Mathematical programming
Lògica matemàtica
Estadística
Intel·ligència artificial
Programació (Matemàtica)
Classificació AMS::03 Mathematical logic and foundations::03B General logic
Classificació AMS::62 Statistics
Classificació AMS::68 Computer science::68T Artificial intelligence
Classificació AMS::90 Operations research, mathematical programming::90C Mathematical programming
Descripción
Sumario:Data collected by statistical offices generally contain errors, which have to be corrected before reliable data can be published. This correction process is referred to as statistical data editing. At statistical offices, certain rules, so-called edits, are often used during the editing process to determine whether a record is consistent or not. Inconsistent records are considered to contain errors, while consistent records are considered error-free. In this article we focus on automatic error localisation based on the Fellegi-Holt paradigm, which says that the data should be made to satisfy all edits by changing the fewest possible number of fields. Adoption of this paradigm leads to a mathematical optimisation problem. We propose an algorithm for solving this optimisation problem for a mix of categorical, continuous and integer-valued data. We also propose a heuristic procedure based on the exact algorithm. For five realistic data sets involving only integer-valued variables we evaluate the performance of this heuristic procedure.