Skip to main content
To KTH's start page

Synthesizing Speech and Gesture for Embodied Conversational Agents

Time: Mon 2026-10-12 10.00

Location: Kollegiesalen, Brinellvägen 8, Stockholm

Language: English

Subject area: Speech and Music Communication

Doctoral student: Siyang Wang , Tal, musik och hörsel

Opponent: Simon King, University of Edinburgh

Supervisor: Professor Joakim Gustafsson, Tal, musik och hörsel; Éva Székely, Tal, musik och hörsel; Simon Alexanderson,

Export to calendar

QC 20260922

Abstract

Embodied conversational agents (ECAs) require synthesis systems that produce natural, spontaneous speech together with appropriate co-speech gesture under real-world conversational conditions. This thesis addresses three intertwined research questions: spontaneous text-to-speech, integrated speech-gesture generation, and evaluation methodology. First, I ask how spontaneous speech can be synthesized; my research shows that disfluencies can be modeled explicitly through sampling-based insertion, that self-supervised speech representations outperform mel-spectrograms as acoustic model prediction targets — with intermediate layers of an ASR-finetuned model performing best — and that token-based speech language models gain spontaneity at the cost of robustness and speaker consistency. Second, I ask whether speech and co-speech gesture can be generated jointly within a single architecture rather than via a pipeline; I formulate this as integrated speech and gesture synthesis (ISG). I demonstrate that a Tacotron2-based ISG model matches a strong pipeline at a fraction of its parameter count and inference time. Third, I ask how such conversational synthesis should be evaluated; I develop a layered protocol that combines task-specific objective metrics, modality-decomposed listening tests, contextual mean-opinion-score evaluation, and live interactive evaluation with real users in an autonomous “20 Questions” dialogue system, and argue that no single layer is sufficient on its own. I close with lessons distilled from unsuccessful experiments on data preparation, on reinforcement-learning post-training for speech language models, and on the role of scale in conversational speech-text pretraining.

The unifying claim of this thesis is that spontaneous speech, co-speech gesture, and the evaluation that judges them are facets of one problem, and that progress on embodied conversational agents depends on both modeling and evaluating them together in the contexts in which conversation actually unfolds. 

Link to DiVA