Recent advances in generative AI have transformed the landscape of speech technologies. Modern voice cloning, text-to-speech, voice conversion, and neural audio generation systems can now produce highly natural and speaker-consistent speech from limited data. While these technologies open exciting opportunities for accessibility, entertainment, personalised interfaces, and assistive communication, they also introduce serious risks for digital identity systems, including spoofing of automatic speaker verification systems, voice phishing, misinformation, and erosion of trust in audio evidence. This tutorial will provide a comprehensive introduction to the security of voice biometric systems in the age of generative AI. It will first review the fundamentals of voice biometrics, including speaker verification pipelines, speaker embeddings, decision thresholds, and evaluation metrics. It will then examine how synthetic and manipulated speech can be used to attack biometric systems and deceive human listeners, with particular attention to logical access attacks, replay attacks, adversarial perturbations, and deepfake-based impersonation. The tutorial will also introduce the design and evaluation of spoofing countermeasures, including artefact-based detection, generalisation challenges, and lessons learned from international benchmarks such as ASVspoof. A new component of the tutorial will focus on audio watermarking and source tracing as complementary approaches to deepfake detection. Rather than only asking whether an audio sample is real or fake, watermarking aims to embed robust and imperceptible signatures into generated or processed audio so that authenticity can be verified, provenance can be traced, and misuse can be investigated. The tutorial will discuss the main principles of audio watermarking, including imperceptibility, robustness, detectability, payload capacity, security, and resistance to common signal processing transformations and adversarial removal. The tutorial is intended for researchers, students, and practitioners working in biometrics, speech processing, multimedia forensics, security, and trustworthy AI. It is designed to be accessible to participants without prior expertise in voice biometrics, deepfake detection, or audio watermarking, although a basic familiarity with machine learning and speech processing will be helpful. By the end of the tutorial, participants will understand the main threats posed by generative AI to voice biometric systems, the strengths and limitations of current countermeasures, and the emerging role of watermarking and source tracing in building trustworthy audio identity technologies.
Trustworthy voice biometrics in the age of generative AI
IJCB 2026, IEEE/IAPR International Joint Conference on Biometrics, 1-4 September 2026, Rome, Italy
Type:
Tutorial
City:
Rome
Date:
2026-09-01
Department:
Digital Security
Eurecom Ref:
8922
Copyright:
© 2026 IEEE. Personal use of this material is permitted. However, permission to reprint/republish this material for advertising or promotional purposes or for creating new collective works for resale or redistribution to servers or lists, or to reuse any copyrighted component of this work in other works must be obtained from the IEEE.
See also:
PERMALINK : https://www.eurecom.fr/publication/8922