Can MLLMs Distinguish Human Laughter?

The 18th International Conference on Social Robotics (ICSR + Art 2026) is currently taking place in London from July 1–4, bringing together researchers, academics, and industry professionals from around the world to explore the latest developments in social robotics. The conference serves as an international platform for exchanging ideas on how intelligent systems can better understand, interact with, and support people in everyday life. On the second day of the conference, Sahan Hatemo, a student at the FHNW School of Computer Science, presented the paper “Reading Between the Laughs: A Human-Referenced Audio Evaluation of MLLMs for Social Robotics”, co-authored with Dr. Katharina Kühne (University of Potsdam) and Prof. Dr. Oliver Bendel (FHNW School of Business). The study investigates whether today’s leading multimodal large language models (MLLMs) can distinguish authentic from non-authentic laughter using audio signals alone. As laughter is an important social cue, the ability to recognize its authenticity could significantly improve how robots and AI systems communicate with people in social settings. The researchers found notable differences in how the evaluated AI models interpreted laughter. OpenAI models showed a clear tendency to classify most laughter as genuine, while Gemini models were generally more skeptical in their assessments. Despite these contrasting biases, several models performed significantly better than chance, with Gemini 2.5 Pro achieving the strongest overall performance. A closer analysis also revealed qualitative differences in the models’ decision-making. Less capable models appeared to rely on superficial acoustic features, such as pitch, and were more likely to classify higher-pitched laughter as less authentic. In contrast, the best-performing model seemed to focus on more sophisticated aspects of voice quality, indicating a deeper understanding of the characteristics that distinguish genuine from non-authentic laughter. The findings demonstrate the growing potential of multimodal AI for social robotics. As robots increasingly become part of everyday environments, the ability to accurately interpret subtle social signals such as laughter could play a crucial role in fostering trust, improving communication, and strengthening human-robot relationships. Further information is available at icsr2026.uk.

Authentic and Non-Authentic Laughter

The paper “Reading Between the Laughs: A Human-Referenced Audio Evaluation of MLLMs for Social Robotics” by Sahan Hatemo, Katharina Kühne, and Oliver Bendel has been accepted at ICSR + Art 2026. In this work, the researchers investigated whether today’s leading AI models can distinguish authentic from non-authentic laughter based solely on audio signals. The results revealed striking differences in model behavior: OpenAI systems showed a strong tendency to interpret most laughter as genuine, while Gemini models were generally more skeptical. Despite these contrasting biases, several models performed significantly better than chance, with Gemini 2.5 Pro achieving the strongest overall results. Their analysis also demonstrated that less capable models often relied on superficial cues such as pitch, disproportionately labeling higher-pitched laughter as less authentic, whereas the top-performing model appeared to focus on more sophisticated voice quality features, suggesting a deeper understanding of laughter authenticity. These findings highlight the growing potential of multimodal large language models in social robotics, where accurately interpreting subtle social signals like laughter could play an important role in trust, communication, and relationship building between humans and robots. The 18th International Conference on Social Robotics will take place in London, UK, from 1-4 July 2026. ICSR is the leading international forum that brings together researchers, academics, and industry professionals from across disciplines to advance the field of social robotics.

About Authentic Laughter

From November 2025 to February 2026, Sahan Hatemo of the FHNW School of Computer Science, Dr. Katharina Kühne of the University of Potsdam, and Prof. Dr. Oliver Bendel of the FHNW School of Business are conducting a research study. As part of this project, they are launching a sub-study that includes a short computer-based task and a brief questionnaire. Participants are asked to listen to a series of laughter samples and evaluate whether each one sounds authentic or not. The task involves 50 samples in total and typically takes about ten minutes to complete. Participation is possible via PC, laptop, or smartphone. Before starting, participants should ensure that their device’s sound is turned on and that they are in a quiet, distraction-free environment. The computer-based task and the brief questionnaire can be accessed at research.sc/participant/login/dynamic/3BE7321C-B5FD-4C4B-AF29-9A435EC39944.