Corvi, Riccardo (2025) From Images to Videos: Analysis and Detection of AI-Generated Media. [Tesi di dottorato]

[thumbnail of Corvi_Riccardo_Thesis_ITEE_38cycle.pdf] Documento PDF
Corvi_Riccardo_Thesis_ITEE_38cycle.pdf

Download (40MB)
Tipologia del documento: Tesi di dottorato
Lingua: English
Titolo: From Images to Videos: Analysis and Detection of AI-Generated Media
Autori:
Autore
Email
Corvi, Riccardo
riccardo.corvi@unina.it
Data: 10 Dicembre 2025
Numero di pagine: 142
Istituzione: Università degli Studi di Napoli Federico II
Dipartimento: Ingegneria Elettrica e delle Tecnologie dell'Informazione
Dottorato: Information technology and electrical engineering
Ciclo di dottorato: 38
Coordinatore del Corso di dottorato:
nome
email
Russo, Stefano
sterusso@unina.it
Tutor:
nome
email
Verdoliva, Luisa
[non definito]
Data: 10 Dicembre 2025
Numero di pagine: 142
Parole chiave: Image Forensics, Deepfakes, Synthetic Image Detection, Synthetic Video Detection
Settori scientifico-disciplinari del MIUR: Area 09 - Ingegneria industriale e dell'informazione > ING-INF/03 - Telecomunicazioni
Informazioni aggiuntive: IL CICLO DI EFFETTIVA APPARTENENZA È IL CICLO 38 (XXXVIII)
Depositato il: 10 Dic 2025 22:53
Ultima modifica: 02 Set 2026 08:04
URI: https://www.fedoa.unina.it/id/eprint/15961

Abstract

Synthetic media generation has improved enormously in the previous years and is now available to everyday users. Powered by large language models, these new generators allow users to create incredibly detailed content starting from a simple text description. These are incredible resources for all types of creative works. However, they provide extraordinarily powerful tools for malicious actors to spread disinformation. As such, Multimedia Forensics is more relevant than ever with the need for general and robust detectors to distinguish real from generated content. This thesis aims first to conduct a systematic study of a large number of generators to discover the most relevant characteristics for forensics applications. This initial analysis has shown that all generators leave artifacts in the generated content and that there are some major discrepancies in the mid-high frequencies between real and synthetic images. Then, the power of large pre-trained Vision-Language Model has been exploited to design a lightweight synthetic image detector based on CLIP. Such models have also been used to develop a robust strategy for the task of fully synthetic video detection. To this end, a novel forensic-oriented data augmentation strategy based on the wavelet decomposition is introduced, in which specific frequency-related bands are replaced to drive the model to exploit more relevant forensic cues. In fact, a well-designed forensic classifier should focus on identifying intrinsic low-level artifacts introduced by a generative architecture rather than relying on high-level semantic flaws that characterize a specific model. This approach improves the generalizability of the detectors, without the need for complex algorithms. Despite its simplicity, the method achieves a significant accuracy improvement over SoTA detectors and obtains excellent results even on very recent generative models.

Downloads

Downloads per month over past year

Actions (login required)

Modifica documento Modifica documento