Corvi, Riccardo (2025) From Images to Videos: Analysis and Detection of AI-Generated Media. [Tesi di dottorato]
|
Documento PDF
Corvi_Riccardo_Thesis_ITEE_38cycle.pdf Download (40MB) |
| Tipologia del documento: | Tesi di dottorato |
|---|---|
| Lingua: | English |
| Titolo: | From Images to Videos: Analysis and Detection of AI-Generated Media |
| Autori: | Autore Email Corvi, Riccardo riccardo.corvi@unina.it |
| Data: | 10 Dicembre 2025 |
| Numero di pagine: | 142 |
| Istituzione: | Università degli Studi di Napoli Federico II |
| Dipartimento: | Ingegneria Elettrica e delle Tecnologie dell'Informazione |
| Dottorato: | Information technology and electrical engineering |
| Ciclo di dottorato: | 38 |
| Coordinatore del Corso di dottorato: | nome email Russo, Stefano sterusso@unina.it |
| Tutor: | nome email Verdoliva, Luisa [non definito] |
| Data: | 10 Dicembre 2025 |
| Numero di pagine: | 142 |
| Parole chiave: | Image Forensics, Deepfakes, Synthetic Image Detection, Synthetic Video Detection |
| Settori scientifico-disciplinari del MIUR: | Area 09 - Ingegneria industriale e dell'informazione > ING-INF/03 - Telecomunicazioni |
| Informazioni aggiuntive: | IL CICLO DI EFFETTIVA APPARTENENZA È IL CICLO 38 (XXXVIII) |
| Depositato il: | 10 Dic 2025 22:53 |
| Ultima modifica: | 02 Set 2026 08:04 |
| URI: | https://www.fedoa.unina.it/id/eprint/15961 |
Abstract
Synthetic media generation has improved enormously in the previous years and is now available to everyday users. Powered by large language models, these new generators allow users to create incredibly detailed content starting from a simple text description. These are incredible resources for all types of creative works. However, they provide extraordinarily powerful tools for malicious actors to spread disinformation. As such, Multimedia Forensics is more relevant than ever with the need for general and robust detectors to distinguish real from generated content. This thesis aims first to conduct a systematic study of a large number of generators to discover the most relevant characteristics for forensics applications. This initial analysis has shown that all generators leave artifacts in the generated content and that there are some major discrepancies in the mid-high frequencies between real and synthetic images. Then, the power of large pre-trained Vision-Language Model has been exploited to design a lightweight synthetic image detector based on CLIP. Such models have also been used to develop a robust strategy for the task of fully synthetic video detection. To this end, a novel forensic-oriented data augmentation strategy based on the wavelet decomposition is introduced, in which specific frequency-related bands are replaced to drive the model to exploit more relevant forensic cues. In fact, a well-designed forensic classifier should focus on identifying intrinsic low-level artifacts introduced by a generative architecture rather than relying on high-level semantic flaws that characterize a specific model. This approach improves the generalizability of the detectors, without the need for complex algorithms. Despite its simplicity, the method achieves a significant accuracy improvement over SoTA detectors and obtains excellent results even on very recent generative models.
Downloads
Downloads per month over past year
Actions (login required)
![]() |
Modifica documento |


