Automatic Detection of Intestinal Content to Evaluate Visibility in Capsule Endoscopy

In capsule endoscopy (CE), preparation of the small bowel before the procedure is believed to increase visibility of the mucosa for analysis. However, there is no consensus on the best method of preparation, while comparison is difficult due to the absence of an objective automated evaluation method...

ver descrição completa

Detalhes bibliográficos
Autores: Noorda, Reinier, Nevárez, Andrea, Colomer, Adrián, Naranjo, Valery, Pons Beltrán, Vicente
Formato: capítulo de livro
Fecha de publicación:2020
País:España
Recursos:Universitat Politècnica de València (UPV)
Repositorio:RiuNet. Repositorio Institucional de la Universitat Politécnica de Valéncia
Idioma:inglés
OAI Identifier:oai:riunet.upv.es:10251/136059
Acesso em linha:https://riunet.upv.es/handle/10251/136059
Access Level:acceso abierto
Palavra-chave:Image processing
Machine learning
Support vector machines
Local binary patterns
Capsule endoscopy
Small bowel preparation
Convolutional neural networks
Descrição
Resumo:In capsule endoscopy (CE), preparation of the small bowel before the procedure is believed to increase visibility of the mucosa for analysis. However, there is no consensus on the best method of preparation, while comparison is difficult due to the absence of an objective automated evaluation method. The method presented here aims to fill this gap by automatically detecting regions in frames of CE videos where the mucosa is covered by bile, bubbles and remainders of food. We implemented two different machine learning techniques for supervised classification of patches: one based on hand-crafted feature extraction and Support Vector Machine classification and the other based on fine-tuning different convolutional neural network (CNN) architectures, concretely VGG-16 and VGG-19. Using a data set of approximately 40,000 image patches obtained from 35 different patients, our best model achieved an average detection accuracy of 95.15% on our test patches, which is similar to significantly more complex detection methods used for similar purposes. We then estimate the probabilities at a pixel level by interpolating the patch probabilities and extract statistics from these, both on per-frame and per-video basis, intended for comparison of different videos.