Abstract
This thesis is about how meaning can be represented, identified, and interpreted in situated multimodal interactions. There have been impressive advances in dialogue systems in recent years, culminating in the development of large language models capable of responding to prompts with fluent, natural-looking text. However, when humans communicate with each other, they use multiple modalities beyond language; for example, gesture. Understanding the meaning conveyed through gesture and other communicative modalities is crucial for artificial intelligence systems to effectively interact with humans in embodied settings. Moreover, such systems must also be able to perform situated grounding, linking communicated information to the outside world.
In this thesis, we present a corpus of multimodal communication in a task-based setting, using Abstract Meaning Representation (AMR) to encode speech, gesture, and action. We perform experiments demonstrating how gestures in our corpus can be automatically detected, and how they may be automatically annotated with their meanings. Finally, we use the speech, gesture, and action annotations to show how meaning is constructed in multimodal interactions, and how objects and actions in the communicative channels are grounded to those in the environment.