Benchmarking Desiderata

Benchmarking Foundation Models for Human-Robot Interaction

This project started fall 2025 as a collaboration between myself, and researchers at Google DeepMind and Lund University.

Project Overview

Social robotics and foundation model (FM) research are developing in parallel, with limited conversation between them. In the FM and benchmarking discourse, robotics features prominently, yet the systems discussed are typically humanoids built for mechanical manipulation rather than social or affective interaction. In human-robot interaction (HRI), however, FMs have become the default architecture for social robots, yet researchers have no formal basis for choosing among models, as existing benchmarks do not measure socio-affective embodied competence. The first two years of my PhD have been devoted to co-design and hands-on HRI studies with vulnerable populations, studying directly how FM-driven social robots succeed and fail in socially and ethically relevant ways. This has now evolved into a project on what ought to be evaluated (evaluation desiderata), which we will continue to operationalise into the first sociotechnical and affective benchmark for FMs in social robots.

Related Publications

image
From "Desiderata for Foundation Models in Social Robots: Capturing Embodied and Social Aspects for Benchmarking". Our proposed set of desiderata with attributes (what should be assessed) and example capabilities (how it may be assessed).
Cover Image: Polytope 1 by Richard A Carter. Richard A Carter / Polytope 1 / Licenced by CC-BY 4.0