TMH Publications (latest 50)
Below are the 50 latest publications from the Department of Speech, Music and Hearing.
TMH Publications
[1]
Thomé, C., Sturm, B., Pertoft, J., Jonason, N. (2026).
Applying Textual Inversion to Control and Personalize Text-to-Music Models.
I Machine Learning and Principles and Practice of Knowledge Discovery in Databases - International Workshops of ECML PKDD 2024, Revised Selected Papers. (s. 395-401). Springer Nature.
[2]
Sundberg, J. (2026).
Ugly and Beautiful Voices : A Rebuttal of Saruhan’s Critical Comments.
Journal of Voice.
[3]
Moëll, B. & Sand Aronsson, F. (2026).
High-accuracy prediction of mental health scores from English BERT embeddings trained on LLM-generated synthetic self-reports: a synthetic-only method development study.
Frontiers in Digital Health, 7.
[4]
Rugayan, J., Salvi, G., Svendsen, T. (2026).
Optimizing ASR Models with Semantic Information.
I TEXT, SPEECH, AND DIALOGUE, TSD 2025, PT I. (s. 25-35). Springer Nature.
[5]
Andin, J., Ellis, R., Ingo, E. & Nordqvist, P. (2026).
Effects of remote work and hearing loss status on well-being and communication in individuals with hearing loss before, during, and after the COVID-19 pandemic : a retrospective survey study.
International Journal of Audiology, 65(6), 703-713.
[6]
Axelsson, A., Vaddadi, B., Bogdan, C. M., Tobin, D. & Skantze, G. (2026).
Robots as Hosts in Autonomous Buses : A Field Trial.
ACM Transactions on Human-Robot Interaction, 15(1).
[7]
Irfan, B., Miniota, J., Thunberg, S., Lagerstedt, E., Kuoppamäki, S., Skantze, G. & Abelho Pereira, A. T. (2026).
Human-Robot Interaction Conversational User Enjoyment Scale (HRI CUES).
IEEE Transactions on Affective Computing, 17(2), 1438-1453.
[8]
Pandey, A., Edlund, J., Le Maguer, S. & Harte, N. (2026).
The use of variable length stimuli for assessing segmental distortion in TTS evaluation.
Computer speech & language (Print), 97.
[9]
Bokkahalli Satish, S. H., Henter, G. E., Székely, É. (2026).
When Voice Matters : Evidence of Gender Disparity in Positional Bias of SpeechLLMs.
I Speech and Computer - 27th International Conference, SPECOM 2025, Proceedings. (s. 25-38). Springer Nature.
[10]
Amerotti, M., Benford, S., Sturm, B. L.T., Vear, C. (2026).
A Live Performance Rule System Informed by Irish Traditional Dance Music.
I Music and Sound Generation in the AI Era - 16th International Symposium, CMMR 2023, Revised Selected Papers. (s. 127-139). Springer Nature.
[11]
Vaddadi, B., Axelsson, A., Skantze, G. (2026).
The Role of Social Robots in Autonomous Public Transport.
I Transport Transitions: Advancing Sustainable and Inclusive Mobility: Proceedings of the 10th TRA Conference, 2024, Dublin, Ireland - Volume 1: Safe and Equitable Transport. (s. 711-716). Springer Nature.
[12]
Ekström, A. G., Tennie, C., Moran, S. & Everett, C. (2026).
The Phoneme as a Cognitive Tool.
Topics in Cognitive Science, 18(1), 43-61.
[13]
Cai, H., Ternström, S., Chaffanjon, P. & Henrich Bernardoni, N. (2026).
Effects on Voice Quality of Thyroidectomy : A Qualitative and Quantitative Study Using Voice Maps.
Journal of Voice, 40(4), 1249.e21-1249.e42.
[14]
Mihajlik, P., Székely, É., Barta, P., Kádár, M. S., Dobsinszki, G., Tóth, L. (2025).
Improved Dysarthric Speech to Text Conversion via TTS Personalization.
I 2025 33rd European Signal Processing Conference, EUSIPCO 2025 - Proceedings. (s. 521-525). Institute of Electrical and Electronics Engineers (IEEE).
[15]
Balsells-Rodas, C., Sumba, X., Narendra, T., Tu, R., Schweikert, G., Kjellström, H., Li, Y. (2025).
Causal Discovery from Conditionally Stationary Time Series.
I Proceedings of Machine Learning Research - International Conference on Machine Learning, ICML 2025. (s. 2715-2741). ML Research Press.
[16]
Olsson, E. J., Madison, G. & Ekström, A. G. (2025).
Is Google liberal on immigration? Attitude bias, politicisation and filter bubbles in search engine result pages.
Heliyon, 11(3).
[17]
Al-Jaff, M., Marchetti, G. L., Welle, M. C., Lundell, J., Gustafsson, M., Henter, G. E., Azizpour, H., Kragic, D. (2025).
A Non-Adversarial Approach to Idempotent Generative Modelling.
I Proceedings ECAI 2025 - 28th European Conference on Artificial Intelligence. (s. 1993-2000). IOS Press.
[18]
Huang, R. S., Holzapfel, A. & Sturm, B. L. T. (2025).
FROM PHILOSOPHY TO PRACTICE : A Culturally Informed Ethics of Music AI in Asia.
I Artificial Intelligence and Music Ecosystem: Second Edition ( (2 uppl.) s. 171-187). Taylor and Francis.
[19]
Huang, R. S., Holzapfel, A. & Sturm, B. L. T. (2025).
GLOBAL ETHICS : Reflection 2.0 (2024).
I Artificial Intelligence and Music Ecosystem: Second Edition ( (2 uppl.) s. 168-170). Informa UK Limited.
[20]
Gurstad-Nilsson, H., Kanhov, E., Bryngelsson, P., Niklasson, M. & Degerman, P. (2025).
Going forward by moving backwards : a perpetual dialectic movement.
I Patrick Hopkinson; Mats Niklasson (Red.), Discovery of International Digital Collaborative Autoethnographical Psychobiography: Knowing You Knowing Me (s. 63-92). Emerald.
[21]
Borg, A., Schiött, J., Ivegren, W., Gentline, C., Huss, V., Hugelius, A. ... Parodis, I. (2025).
AI-Enhanced Social Robotic Versus Computer-Based Virtual Patients for Clinical Reasoning Training in Medical Education : Observational Crossover Cohort Study.
Journal of Medical Internet Research, 27.
[22]
Borg, A., Jobs, B., Huss, V., Gentline, C., Espinosa, F., Ruiz, M. ... Parodis, I. (2025).
A qualitative comparison of clinical reasoning training : LLM-powered social robotic versus computer-based virtual patients for undergraduate medical education in rheumatology.
Scandinavian Journal of Rheumatology, 54(Suppl. 132), 302-302.
[23]
Willemsen, B., Skantze, G. (2025).
Detecting Referring Expressions in Visually Grounded Dialogue with Autoregressive Language Models.
Presenterad vid XLLM @ ACL 2025, The 1st Joint Workshop on Large Language Models and Structure Modeling, Vienna, Austria, Aug 1st, 2025. Association for Computational Linguistics (ACL).
[24]
Stinkeste, C., Dreber, A., Olofsson, J. & Skantze, G. (2025).
Comparing the audience effect of anthropomorphic robots and humans in economic games.
Computers in Human Behavior: Artificial Humans, 6.
[26]
Francis, J., Gustafsson, J., Székely, É. (2025).
From Static to Dynamic : Enhancing AAC with Generative Imagery and Zero-Shot TTS.
I Interspeech 2025. (s. 4960-4962). International Speech Communication Association.
[27]
Bokkahalli Satish, S. H., Henter, G. E., Székely, É. (2025).
Hear Me Out : Interactive evaluation and bias discovery platform for speech-to-speech conversational AI.
I Interspeech 2025. (s. 2151-2152). International Speech Communication Association.
[28]
Park, M., Ontakhrai, S., Kittimathaveenan, K., Alfredsson, J., Ternström, S. (2025).
How to make closed-back headphones transparent for avocalist’s own direct sound.
Presenterad vid AES 159th Convention 2025 October 23–25, Long Beach, CA, USA. (s. 8). Audio Engineering Society, Inc.
[29]
Tånnander, C., House, D., Beskow, J., Edlund, J. (2025).
Intrasentential English in Swedish TTS : perceived English-accentedness.
I Interspeech 2025. (s. 1638-1642). International Speech Communication Association.
[30]
Malisz, Z., Foremski, J., Kul, M. (2025).
Contextual predictability effects on acoustic distinctiveness in read Polish speech.
I Interspeech 2025. (s. 335-339). International Speech Communication Association.
[31]
Thulinsson, F., Söderlund, N., Rafiei, S., Schenkman, B., Djupsjöbacka, A., Andrén, B., Brunnström, K. (2025).
Impact of Camera height and Field-of-View on distance judgement and gap selection in digital rear-view mirrors in vehicles.
I IS and T International Symposium on Electronic Imaging Science and Technology. Society for Imaging Science & Technology.
[32]
Ekström, A. G., Gärdenfors, P., Snyder, W. D., Friedrichs, D., McCarthy, R. C., Tsapos, M. ... Moran, S. (2025).
Correlates of Vocal Tract Evolution in Late Pliocene and Pleistocene Hominins.
Human Nature, 36(1), 22-69.
[33]
Moëll, B. (2025).
Evaluation of Artificial Intelligence in the Medical Domain : Speech, Language and Applications
(Doktorsavhandling , KTH Royal Institute of Technology, Stockholm, TRITA-EECS-AVL 2025:83). Hämtad från https://urn.kb.se/resolve?urn=urn:nbn:se:kth:diva-371738.
[34]
Moell, B. & Sand Aronsson, F. (2025).
Automatic Evaluation of the Pataka Test Using Machine Learning and Audio Signal Processing.
Acta Logopaedica, 2.
[35]
Leite, I., Ahlberg, W., Pereira, A., Sestini, A., Gisslen, L., Tollmar, K. (2025).
A Call for Deeper Collaboration Between Robotics and Game Development.
I Proceedings of the IEEE 2025 Conference on Games, CoG 2025. Institute of Electrical and Electronics Engineers (IEEE).
[36]
Jacka, R., Peña, P. R., Leonard, S. J., Székely, É., Cowan, B. R. (2025).
Impact Of Disfluent Speech Agent On Partner Models And Perspectve Taking.
I CUI 2025 - Proceedings of the 2025 ACM Conference on Conversational User Interfaces. Association for Computing Machinery (ACM).
[37]
Moëll, B. & Sand Aronsson, F. (2025).
Journaling with large language models : a novel UX paradigm for AI-driven personal health management.
Frontiers in Artificial Intelligence, 8.
[38]
Grouwels, J., Jonason, N., Sturm, B. (2025).
Exploring the Expressive Space of an Articulatory Vocal Modal using Quality-Diversity Optimization with Multimodal Embeddings.
I GECCO 2025 - Proceedings of the 2025 Genetic and Evolutionary Computation Conference. (s. 1362-1370). Association for Computing Machinery (ACM).
[39]
Cavalcanti, J. C., Skantze, G. (2025).
"Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking.
I Proceedings of the Interspeech 2025. Rotterdam, The Netherlands: International Speech Communication Association.
[40]
Moëll, B. & Sand Aronsson, F. (2025).
Harm Reduction Strategies for Thoughtful Use of Large Language Models in the Medical Domain : Perspectives for Patients and Clinicians.
Journal of Medical Internet Research, 27.
[41]
Best, P., Araya-Salas, M., Ekström, A. G., Freitas, B., Jensen, F. H., Kershenbaum, A. ... Marxer, R. (2025).
Bioacoustic fundamental frequency estimation : a cross-species dataset and deep learning baseline.
Bioacoustics, 34(4), 419-446.
[42]
Torubarova, E. (2025).
Brain-Focused Multimodal Approach for Studying Conversational Engagement in HRI.
I HRI 2025 - Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. (s. 1894-1896). Institute of Electrical and Electronics Engineers (IEEE).
[43]
Irfan, B., Churamani, N., Zhao, M., Ayub, A., Rossi, S. (2025).
Lifelong Learning and Personalization in Long-Term Human-Robot Interaction (LEAP-HRI) : Overcoming Inequalities with Adaptation.
I HRI 2025 - Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. (s. 1970-1972). Institute of Electrical and Electronics Engineers (IEEE).
[44]
Skantze, G., Irfan, B. (2025).
Applying General Turn-Taking Models to Conversational Human-Robot Interaction.
I HRI 2025 - Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. (s. 859-868). Institute of Electrical and Electronics Engineers (IEEE).
[45]
Irfan, B., Skantze, G. (2025).
Between You and Me: Ethics of Self-Disclosure in Human-Robot Interaction.
I HRI 2025 - Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. (s. 1357-1362). Institute of Electrical and Electronics Engineers (IEEE).
[46]
Kamelabad, A. M., Inoue, E., Skantze, G. (2025).
Comparing Monolingual and Bilingual Social Robots as Conversational Practice Companions in Language Learning.
I Proceedings of the 2025 ACM/IEEE International Conference on Human-Robot Interaction. (s. 829-838).
[47]
Cai, H. & Ternström, S. (2025).
A WaveNet-based model for predicting the electroglottographic signal from the acoustic voice signal.
Journal of the Acoustical Society of America, 157(4), 3033-3044.
[48]
Marcinek, L., Beskow, J., Gustafsson, J. (2025).
A Dual-Control Dialogue Framework for Human-Robot Interaction Data Collection : Integrating Human Emotional and Contextual Awareness with Conversational AI.
I Social Robotics - 16th International Conference, ICSR + AI 2024, Proceedings. (s. 290-297). Springer Nature.
[49]
Irfan, B., Kuoppamäki, S., Hosseini, A. & Skantze, G. (2025).
Between reality and delusion : challenges of applying large language models to companion robots for open-domain dialogues with older adults.
Autonomous Robots, 49(1).
[50]
Włodarczak, M., Ludusan, B., Sundberg, J. & Heldner, M. (2025).
Classification of voice quality using neck-surface acceleration : Comparison with glottal flow and radiated sound.
Journal of Voice, 39(1), 10-24.