Logical Operators for Multimodal Fusion in Temporal Video Scene Segmentation

Early fusion techniques in content analysis aim to enhance efficacy by generating compact data models that retain semantic clues from multimodal data. Initial attempts used fusion operators at low-level feature space, which compromised data representativeness. This led to the development of complex...

Full description

Bibliographic Details
Authors: Barbosa, Letícia B., Goularte, Rudinei
Format: article
Status:Published version
Publication Date:2025
Country:Brasil
Institution:Sociedade Brasileira de Computação (SBC)
Repository:Revista Eletrônica de Iniciação Científica
Language:English
OAI Identifier:oai:journals-sol.sbc.org.br:article/5535
Online Access:https://journals-sol.sbc.org.br/index.php/reic/article/view/5535
Access Level:Open access
Keyword:Multimodal Fusion
Fusion Operators
Video Scene Segmentation
Video Analysis
Description
Summary:Early fusion techniques in content analysis aim to enhance efficacy by generating compact data models that retain semantic clues from multimodal data. Initial attempts used fusion operators at low-level feature space, which compromised data representativeness. This led to the development of complex operations inseparable from multimodal semantic clues processing. Previous studies showed that simple arithmetic-based operators could be as effective as complex operations when applied at the mid-level feature space, highlighting an unexplored opportunity to assess the efficacy of logical operators. This paper investigates the application of logical fusion operators (And, Or, Xor) at the mid-level feature space for Temporal Video Scene Segmentation. Comparative analysis demonstrates that Or and Xor logical operators are viable alternatives in the specific Temporal Video Scene Segmentation content analysis tasks.