Logical Operators for Multimodal Fusion in Temporal Video Scene Segmentation
Early fusion techniques in content analysis aim to enhance efficacy by generating compact data models that retain semantic clues from multimodal data. Initial attempts used fusion operators at low-level feature space, which compromised data representativeness. This led to the development of complex...
| Authors: | , |
|---|---|
| Format: | article |
| Status: | Published version |
| Publication Date: | 2025 |
| Country: | Brasil |
| Institution: | Sociedade Brasileira de Computação (SBC) |
| Repository: | Revista Eletrônica de Iniciação Científica |
| Language: | English |
| OAI Identifier: | oai:journals-sol.sbc.org.br:article/5535 |
| Online Access: | https://journals-sol.sbc.org.br/index.php/reic/article/view/5535 |
| Access Level: | Open access |
| Keyword: | Multimodal Fusion Fusion Operators Video Scene Segmentation Video Analysis |
| Summary: | Early fusion techniques in content analysis aim to enhance efficacy by generating compact data models that retain semantic clues from multimodal data. Initial attempts used fusion operators at low-level feature space, which compromised data representativeness. This led to the development of complex operations inseparable from multimodal semantic clues processing. Previous studies showed that simple arithmetic-based operators could be as effective as complex operations when applied at the mid-level feature space, highlighting an unexplored opportunity to assess the efficacy of logical operators. This paper investigates the application of logical fusion operators (And, Or, Xor) at the mid-level feature space for Temporal Video Scene Segmentation. Comparative analysis demonstrates that Or and Xor logical operators are viable alternatives in the specific Temporal Video Scene Segmentation content analysis tasks. |
|---|