Benchmarking Foundation Models for Human-Robot Interaction
This project started fall 2025 as a collaboration between myself, and researchers at Google DeepMind and Lund University.
Project Overview
Social robotics and foundation model (FM) research are developing in parallel, with limited conversation between them. In the FM and benchmarking discourse, robotics features prominently, yet the systems discussed are typically humanoids built for mechanical manipulation rather than social or affective interaction. In human-robot interaction (HRI), however, FMs have become the default architecture for social robots, yet researchers have no formal basis for choosing among models, as existing benchmarks do not measure socio-affective embodied competence. The first two years of my PhD have been devoted to co-design and hands-on HRI studies with vulnerable populations, studying directly how FM-driven social robots succeed and fail in socially and ethically relevant ways. This has now evolved into a project on what ought to be evaluated (evaluation desiderata), which we will continue to operationalise into the first sociotechnical and affective benchmark for FMs in social robots.
Related Publications
- Nichols, E., Markelius, A., & Gunes, H. (2026). How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots. Accepted at the FoRMA workshop at 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). [DOI]
- Markelius, A., Brady, D., Tanqueray, L. & Gunes, H. (2026). Desiderata for Foundation Models in Social Robots: Capturing Embodied and Social Aspects for Benchmarking. In 2026 35th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)
