Pianese, Alessandro (2026) On the robust and generalizable detection of synthetic video and audio tracks. [Tesi di dottorato]

[thumbnail of Pianese_Alessandro_ITEE_Thesis_38cycle.pdf] Documento PDF
Pianese_Alessandro_ITEE_Thesis_38cycle.pdf
Visibile a [TBR] Amministratori dell'archivio

Download (4MB) | Richiedi una copia
Tipologia del documento: Tesi di dottorato
Lingua: English
Titolo: On the robust and generalizable detection of synthetic video and audio tracks
Autori:
Autore
Email
Pianese, Alessandro
alepianese@gmail.com
Data: 10 Febbraio 2026
Numero di pagine: 138
Istituzione: Università degli Studi di Napoli Federico II
Dipartimento: Ingegneria Elettrica e delle Tecnologie dell'Informazione
Dottorato: Information technology and electrical engineering
Ciclo di dottorato: 38
Coordinatore del Corso di dottorato:
nome
email
Russo, Stefano
stefano.russo@unina.it
Tutor:
nome
email
Poggi, Giovanni
[non definito]
Data: 10 Febbraio 2026
Numero di pagine: 138
Parole chiave: Synthetic speech; detection; generalization; pre-trained models
Settori scientifico-disciplinari del MIUR: Area 09 - Ingegneria industriale e dell'informazione > ING-INF/03 - Telecomunicazioni
Informazioni aggiuntive: 38° Ciclo
Depositato il: 11 Feb 2026 21:35
Ultima modifica: 12 Ago 2026 05:37
URI: https://www.fedoa.unina.it/id/eprint/16202

Abstract

Our time witnessed the explosive growth of artificial intelligence. New tools based on this technology have replaced their older counterparts, driving rapid development in existing fields of application and creating entirely new ones. This is certainly true in the field of natural language processing with powerful speech synthesis and vocoding models that have transformed the landscape of the field. These technologies enable very interesting uses, such as automatic translation of educational audio or live translation during online conferences. However, we are also witnessing a growing use of these technologies for fraudulent ends. Increasingly, synthetic audio and deepfakes are being used to spread political and personal disinformation or to commit fraud. The fight against these phenomena is Multimedia Forensics, a rapidly expanding field that also relies predominantly by now on AI-based tools. The most popular approach to detect synthetic or manipulated media is to train a model in a supervised manner on a subset of real and fake data. However, such an approach behaves poorly with newer technologies and it is too sensitive to signal degradations. To address this issues, we chose to explore innovative solutions that ensure better generalization and greater robustness to compression and background noise. We firstly introduced a deepfake detector that leverages contrastive learning to produce representative embeddings for both audio-visual tracks. We then modified the audio portion by introducing training-free solutions that, thanks to a pre-trained model, prevent bias and improve robustness. Finally, with the aim of supporting better interpretability of the results, we proposed a new encoding-decoding architecture to ensure that the models effectively focus on the individual's unique speech pattern.

Downloads

Downloads per month over past year

Actions (login required)

Modifica documento Modifica documento