Shehzad, Faheem (2026) Efficient Dual-Stream Architectures for Visual Question Answering. [Tesi di dottorato]
|
Documento PDF
Faheem_unina_thesis.pdf Visibile a [TBR] Amministratori dell'archivio Download (8MB) | Richiedi una copia |
| Tipologia del documento: | Tesi di dottorato |
|---|---|
| Lingua: | English |
| Titolo: | Efficient Dual-Stream Architectures for Visual Question Answering |
| Autori: | Autore Email Shehzad, Faheem faheem.shehzad@unina.it |
| Data: | 10 Febbraio 2026 |
| Numero di pagine: | 101 |
| Istituzione: | Università degli Studi di Napoli Federico II |
| Dipartimento: | Biologia |
| Dottorato: | Intelligenza artificiale Area Agrifood e ambiente |
| Ciclo di dottorato: | 38 |
| Coordinatore del Corso di dottorato: | nome email Loreto, Francesco francesco.loreto@unina.it |
| Tutor: | nome email Esposito, Massimo [non definito] |
| Data: | 10 Febbraio 2026 |
| Numero di pagine: | 101 |
| Parole chiave: | Visual Question Answering, Artificial Intelligence, Medical Imaging |
| Settori scientifico-disciplinari del MIUR: | Area 09 - Ingegneria industriale e dell'informazione > ING-INF/05 - Sistemi di elaborazione delle informazioni |
| Depositato il: | 25 Feb 2026 13:46 |
| Ultima modifica: | 12 Ago 2026 05:37 |
| URI: | https://www.fedoa.unina.it/id/eprint/16290 |
Abstract
Visual Question Answering (VQA) requires the integration of visual and textual information to answer natural language questions about images. While powerful, state-of-the-art VQA systems typically rely on large, computationally expensive pre-trained multimodal models, rendering them unsuitable for domains constrained by data privacy, limited local computational resources, or the absence of large-scale annotated datasets. This doctoral thesis primarily addresses this gap by designing and systematically evaluating a novel, lightweight VQA architecture. The core objective is to develop a simple yet effective framework that strategically combines multimodal information from images and text without depending on massive pre-trained vision-language models and without prohibitive computational or memory overhead. The architecture is deliberately designed around efficient, independent visual and textual encoders, coupled with streamlined cross-modal fusion mechanisms, to achieve an optimal balance between performance and efficiency. To rigorously assess the practical applicability and effectiveness of this proposed architecture, we focus on two challenging, real-world domains: healthcare and agriculture. These domains are characterized by stringent data privacy regulations (often necessitating on-device/local processing), a critical need for domain-specific visual reasoning, and a severe scarcity of public, annotated VQA datasets. The evaluation across these domains serves as a stringent testbed for the architecture’s efficiency, adaptability, and performance. As a foundational enabler for this evaluation, a secondary contribution of this thesis is the creation of novel, application-focused VQA datasets for brain tumor MRI, gastrointestinal endoscopy, hematology microscopy, and rice leaf disease. While these datasets provide essential benchmarks, the primary emphasis remains on their role in facilitating the targeted design and comprehensive evaluation of domain-adapted VQA systems. Our methodological design and evaluation demonstrate that the proposed dual-stream transformer framework, employing distilled encoders and careful fusion strategy selection, achieves competitive accuracy against significantly larger baselines. Crucially, it does so with markedly reduced training times and GPU memory footprints. Extensive ablation studies and category-wise analyses provide insights into the trade-offs between architectural choices, model efficiency, and reasoning robustness. In summary, this thesis advances domain-specific VQA by shifting the focus from mere dataset collection or application of monolithic models to the principled design and rigorous evaluation of efficient, deployable multimodal architectures. The contributions offer a blueprint for developing VQA systems that are effective, resource-conscious, and adaptable to privacy-sensitive, data-scarce environments, with direct implications for clinical decision support and precision agriculture.
Downloads
Downloads per month over past year
Actions (login required)
![]() |
Modifica documento |


