Shehzad, Faheem (2026) Efficient Dual-Stream Architectures for Visual Question Answering. [Tesi di dottorato]

[thumbnail of Faheem_unina_thesis.pdf] Documento PDF
Faheem_unina_thesis.pdf
Visibile a [TBR] Amministratori dell'archivio

Download (8MB) | Richiedi una copia
Tipologia del documento: Tesi di dottorato
Lingua: English
Titolo: Efficient Dual-Stream Architectures for Visual Question Answering
Autori:
Autore
Email
Shehzad, Faheem
faheem.shehzad@unina.it
Data: 10 Febbraio 2026
Numero di pagine: 101
Istituzione: Università degli Studi di Napoli Federico II
Dipartimento: Biologia
Dottorato: Intelligenza artificiale Area Agrifood e ambiente
Ciclo di dottorato: 38
Coordinatore del Corso di dottorato:
nome
email
Loreto, Francesco
francesco.loreto@unina.it
Tutor:
nome
email
Esposito, Massimo
[non definito]
Data: 10 Febbraio 2026
Numero di pagine: 101
Parole chiave: Visual Question Answering, Artificial Intelligence, Medical Imaging
Settori scientifico-disciplinari del MIUR: Area 09 - Ingegneria industriale e dell'informazione > ING-INF/05 - Sistemi di elaborazione delle informazioni
Depositato il: 25 Feb 2026 13:46
Ultima modifica: 12 Ago 2026 05:37
URI: https://www.fedoa.unina.it/id/eprint/16290

Abstract

Visual Question Answering (VQA) requires the integration of visual and textual information to answer natural language questions about images. While powerful, state-of-the-art VQA systems typically rely on large, computationally expensive pre-trained multimodal models, rendering them unsuitable for domains constrained by data privacy, limited local computational resources, or the absence of large-scale annotated datasets. This doctoral thesis primarily addresses this gap by designing and systematically evaluating a novel, lightweight VQA architecture. The core objective is to develop a simple yet effective framework that strategically combines multimodal information from images and text without depending on massive pre-trained vision-language models and without prohibitive computational or memory overhead. The architecture is deliberately designed around efficient, independent visual and textual encoders, coupled with streamlined cross-modal fusion mechanisms, to achieve an optimal balance between performance and efficiency. To rigorously assess the practical applicability and effectiveness of this proposed architecture, we focus on two challenging, real-world domains: healthcare and agriculture. These domains are characterized by stringent data privacy regulations (often necessitating on-device/local processing), a critical need for domain-specific visual reasoning, and a severe scarcity of public, annotated VQA datasets. The evaluation across these domains serves as a stringent testbed for the architecture’s efficiency, adaptability, and performance. As a foundational enabler for this evaluation, a secondary contribution of this thesis is the creation of novel, application-focused VQA datasets for brain tumor MRI, gastrointestinal endoscopy, hematology microscopy, and rice leaf disease. While these datasets provide essential benchmarks, the primary emphasis remains on their role in facilitating the targeted design and comprehensive evaluation of domain-adapted VQA systems. Our methodological design and evaluation demonstrate that the proposed dual-stream transformer framework, employing distilled encoders and careful fusion strategy selection, achieves competitive accuracy against significantly larger baselines. Crucially, it does so with markedly reduced training times and GPU memory footprints. Extensive ablation studies and category-wise analyses provide insights into the trade-offs between architectural choices, model efficiency, and reasoning robustness. In summary, this thesis advances domain-specific VQA by shifting the focus from mere dataset collection or application of monolithic models to the principled design and rigorous evaluation of efficient, deployable multimodal architectures. The contributions offer a blueprint for developing VQA systems that are effective, resource-conscious, and adaptable to privacy-sensitive, data-scarce environments, with direct implications for clinical decision support and precision agriculture.

Downloads

Downloads per month over past year

Actions (login required)

Modifica documento Modifica documento