Pianese, Alessandro (2026) On the robust and generalizable detection of synthetic video and audio tracks. [Tesi di dottorato]
|
Documento PDF
Pianese_Alessandro_ITEE_Thesis_38cycle.pdf Visibile a [TBR] Amministratori dell'archivio Download (4MB) | Richiedi una copia |
| Tipologia del documento: | Tesi di dottorato |
|---|---|
| Lingua: | English |
| Titolo: | On the robust and generalizable detection of synthetic video and audio tracks |
| Autori: | Autore Email Pianese, Alessandro alepianese@gmail.com |
| Data: | 10 Febbraio 2026 |
| Numero di pagine: | 138 |
| Istituzione: | Università degli Studi di Napoli Federico II |
| Dipartimento: | Ingegneria Elettrica e delle Tecnologie dell'Informazione |
| Dottorato: | Information technology and electrical engineering |
| Ciclo di dottorato: | 38 |
| Coordinatore del Corso di dottorato: | nome email Russo, Stefano stefano.russo@unina.it |
| Tutor: | nome email Poggi, Giovanni [non definito] |
| Data: | 10 Febbraio 2026 |
| Numero di pagine: | 138 |
| Parole chiave: | Synthetic speech; detection; generalization; pre-trained models |
| Settori scientifico-disciplinari del MIUR: | Area 09 - Ingegneria industriale e dell'informazione > ING-INF/03 - Telecomunicazioni |
| Informazioni aggiuntive: | 38° Ciclo |
| Depositato il: | 11 Feb 2026 21:35 |
| Ultima modifica: | 12 Ago 2026 05:37 |
| URI: | https://www.fedoa.unina.it/id/eprint/16202 |
Abstract
Our time witnessed the explosive growth of artificial intelligence. New tools based on this technology have replaced their older counterparts, driving rapid development in existing fields of application and creating entirely new ones. This is certainly true in the field of natural language processing with powerful speech synthesis and vocoding models that have transformed the landscape of the field. These technologies enable very interesting uses, such as automatic translation of educational audio or live translation during online conferences. However, we are also witnessing a growing use of these technologies for fraudulent ends. Increasingly, synthetic audio and deepfakes are being used to spread political and personal disinformation or to commit fraud. The fight against these phenomena is Multimedia Forensics, a rapidly expanding field that also relies predominantly by now on AI-based tools. The most popular approach to detect synthetic or manipulated media is to train a model in a supervised manner on a subset of real and fake data. However, such an approach behaves poorly with newer technologies and it is too sensitive to signal degradations. To address this issues, we chose to explore innovative solutions that ensure better generalization and greater robustness to compression and background noise. We firstly introduced a deepfake detector that leverages contrastive learning to produce representative embeddings for both audio-visual tracks. We then modified the audio portion by introducing training-free solutions that, thanks to a pre-trained model, prevent bias and improve robustness. Finally, with the aim of supporting better interpretability of the results, we proposed a new encoding-decoding architecture to ensure that the models effectively focus on the individual's unique speech pattern.
Downloads
Downloads per month over past year
Actions (login required)
![]() |
Modifica documento |


