Vento, Bruno Advanced AI Models for Environmental Monitoring and Disaster Prevention in Agriculture. [Tesi di dottorato]

[thumbnail of TESI_DOTTORATO_BRUNO_VENTO.pdf] Documento PDF
TESI_DOTTORATO_BRUNO_VENTO.pdf
Visibile a [TBR] Amministratori dell'archivio

Download (63MB) | Richiedi una copia
Tipologia del documento: Tesi di dottorato
Lingua: English
Titolo: Advanced AI Models for Environmental Monitoring and Disaster Prevention in Agriculture
Autori:
Autore
Email
Vento, Bruno
bruno.vento@unina.it
Numero di pagine: 258
Istituzione: Università degli Studi di Napoli Federico II
Dottorato: Intelligenza artificiale Area Agrifood e ambiente
Ciclo di dottorato: 38
Coordinatore del Corso di dottorato:
nome
email
Loreto, Francesco
francesco.loreto@unina.it
Tutor:
nome
email
Sansone, Carlo
[non definito]
Greco, Antonio
[non definito]
Numero di pagine: 258
Parole chiave: AI, Pedestrian Attribute Recognition, Weather and Ground Conditions Recognition, Federated Event-Driven Infrastructure
Settori scientifico-disciplinari del MIUR: Area 09 - Ingegneria industriale e dell'informazione > ING-INF/05 - Sistemi di elaborazione delle informazioni
Informazioni aggiuntive: 38 ciclo
Depositato il: 29 Dic 2025 15:45
Ultima modifica: 12 Ago 2026 05:39
URI: https://www.fedoa.unina.it/id/eprint/17081

Abstract

The thesis addresses the theme of intelligent perception in rural and agricultural contexts, positioned at the intersection between computer vision, distributed intelligence, and technological sustainability. In Agriculture 4.0, traditional sensor networks evolve into a broader and more adaptive perception paradigm: smart cameras, enhanced by deep learning and multimodal models, operate as intelligent sensors capable of detecting complex phenomena while reducing the costs and rigidity typical of sensor-specific architectures. The objective of the thesis is to develop an integrated and reliable automatic perception system capable of functioning in uncontrolled scenarios, operating in real time on resource-limited edge devices, and unifying within a single infrastructure the domains of fire detection, joint weather–ground condition recognition, and pedestrian attribute recognition. The first research axis concerns the automatic detection of fires in both natural and human-made environments. The analysis begins with a review of existing methods, highlighting limited generalizability due to scenario heterogeneity and the conventional separation between “wildfire” and “urban” contexts. To address this issue, a new taxonomy is proposed based on two measurable parameters: apparent distance (Short Range vs. Long Range) and scene activity level (Low Activity vs. High Activity). This classification defines four operational categories and supports a unified evaluation framework for consistent method comparison. Building on this foundation, the ONFIRE contest was developed as a public benchmark for the integrated evaluation of fire detection systems. It includes over three hundred annotated videos with frame-level metadata describing ignition timing, smoke presence, false look-alikes, and illumination and motion conditions. A hidden test set and a composite metric, the Fire Detection Score (FDS), are introduced to jointly account for precision, recall, latency, and computational cost. Based on these experimental resources, FOCUS was designed as a scenario-adaptive framework for real-time fire detection and classification. The system employs a modified YOLOv8s network optimized for explicit discrimination between smoke and flames, with dual detection heads and a balanced training scheme including hard negatives such as artificial lights, clouds, and reflections. The architecture integrates a temporal module that aggregates visual evidence across frames to mitigate false alarms due to flicker or dynamic noise, and an optional semantic component based on the BLIP-2 vision–language model for contextual disambiguation. Three operational configurations, Basic, Standard, and Advanced, cover different computational profiles, from edge-level inference to semantically enhanced processing on centralized infrastructures. Experimental evaluations on ONFIRE contest and independent datasets show that FOCUS achieves higher precision, recall, and cross-scenario robustness than baseline models, with mean accuracy values close to 98%. The second research axis addresses the joint analysis of weather and ground conditions, relevant for contextual perception in agricultural monitoring. The research proceeds along two main directions: data corpus creation and model design. To compensate for the lack of integrated datasets, the thesis introduces a public image archive annotated simultaneously for weather and ground states. The taxonomy distinguishes between cloudy, foggy, rainy, snowy, and sunny skies, and dry, wet, or flooded surfaces, allowing supervised learning even with missing labels. At the architectural level, the first proposed system is a single-task model designed for on-camera inference on commercial smart cameras. Based on MobileNetV2, it processes low-resolution inputs and applies temporal majority voting for output stabilization. This implementation demonstrates real-time operation feasibility on embedded hardware. A second, more advanced model performs multi-task classification of weather and ground conditions using shared low-level features and specialized high-level representations. The ConvNeXt backbone is adapted with Convolutional Block Attention Module (CBAM) attention modules to focus processing on relevant spatial regions, such as the sky for weather estimation and the ground for surface state recognition. Independent classification heads are trained with a Masked Asymmetric Loss, which excludes missing labels from loss computation, and with dynamic gradient balancing via GradNorm to maintain proportional learning. Experimental results indicate that the multi-task model provides higher accuracy and computational efficiency than single-task versions. It maintains throughput around 50 fps with low memory usage and stable inter-class performance. The third research axis concerns Pedestrian Attribute Recognition (PAR) as a component for safety and monitoring in rural environments, where semantic understanding of detected individuals, such as gender, clothing color, or accessory presence, is required. The study begins with an analysis of current literature, identifying limitations related to incomplete annotations, multi-task instability, and efficiency–semantic depth trade-offs. A reference dataset, MIVIA PAR KD, has been incrementally built in the two editions of the competition including first partially annotated samples and, then, complete labels derived through knowledge distillation. On this dataset, two complementary architectures are developed. SPARTAN is a multi-task CNN based on ConvNeXt-Base, optimized for pedestrian image geometry. It employs five parallel branches, each dedicated to a specific attribute, incorporating a Hybrid Attention Module (HAM) that combines channel and spatial attention. The training procedure handles incomplete labels and uses dynamic task balancing via GradNorm. It achieves high mean accuracy with real-time inference exceeding 110 fps and consistent results across attributes. PARVELOUS represents a hybrid model integrating efficiency in vision–language model semantic representation. It uses a SigLIP2 encoder, optimized through the NaFlex strategy to preserve the native aspect ratio of pedestrian crops. The five attribute branches are selectively fine-tuned to maintain semantic generalization without increasing model size. PARVELOUS attains performance comparable to larger transformer models while remaining compatible with edge hardware. All modules are integrated within FARM-VISION, a federated and event-driven infrastructure designed for distributed perception in Agriculture 4.0. The platform combines local and centralized intelligence: smart cameras perform lightweight inference and transmit metadata or cropped images upon event detection; the central node, based on BLIP-2, conducts semantic validation; a local event handler ensures communication continuity. Field experiments in rural settings report a reduction in bandwidth usage, a decrease in power consumption, and a reduction in server count, with consistent detection performance across modules. The results of the thesis in terms of accuracy, efficiency, and system scalability provide a reference foundation for the development of next-generation agricultural monitoring and surveillance infrastructures, where artificial intelligence, edge computing, and distributed collaboration support environmental resilience and sustainable innovation.

Downloads

Downloads per month over past year

Actions (login required)

Modifica documento Modifica documento