Logo image
Uh, Actually: Grunt Acts as Grounding Contributions in Multimodal Dialogue
Dissertation   Open access

Uh, Actually: Grunt Acts as Grounding Contributions in Multimodal Dialogue

Richard Brutti
Doctor of Philosophy (PhD), Brandeis University
2026
DOI:
https://doi.org/10.48617/etd.1658

Abstract

common ground gesture grunts non-verbal communication Semantics
This dissertation argues that non-lexical communicative acts and actions are first-class contributions to the common ground in multimodal dialogue. There is more than oh going on. The acknowledging minimal vocalizations, conversational grunts like oh, mm-hm, and uh-huh, are among the primary signals through which grounding occurs in co-situated interaction. Despite their frequent appearance, these phenomena are systematically excluded from computational models of dialogue, which assume that the propositional content of an interaction is carried by its lexical content. This dissertation develops the formal and empirical steps to update that assumption. The dominant frameworks in dialogue modeling reflect this assumption in concrete ways. Dialogue state tracking systems represent information state as slot-value pairs populated by utterances. Meaning representation frameworks were designed for written sentences, and while their extensions can handle speech acts, they have no principled account of non-lexical modalities. Automatic speech recognition (ASR) systems explicitly filter out the vocalizations that are most communicatively relevant: OpenAI's Whisper removes hmm, mm, mhm, uh, and um as a matter of policy. BERT has been shown to use similar representations for sentences that have similar meaning with or without so-called disfluencies. Large language models are pre-trained predominantly on written text, and written language conventions actively suppress the non-lexical vocalizations that are ubiquitous in speech. The result is that the signals that most directly encode grounding in situated dialogue are invisible to the systems designed to model it. The first contribution is Gesture AMR (GAMR), an extension of Abstract Meaning Representation that captures the semantic content of gesture through a taxonomy of gesture act relations. GAMR is realized in an annotated corpus of multimodal speech and gesture. The theoretical result established by this work is foundational: non-lexical communicative acts have representable propositional content, and the formal machinery for capturing that content can be extended beyond gesture to other non-lexical modalities. The scaffolding for the grunt analysis is developed through collaborative work. The Weights Task Dataset, a multimodal corpus of collaborative problem-solving, provides the primary empirical resource for the dissertation and was designed specifically to elicit genuine multimodal communication in a co-situated, shared physical environment. The Common Ground Tracking (CGT) framework operationalizes multimodal common ground formally as a dynamic structure of questions under discussion, evidenced propositions, and accepted facts, updated by dialogue moves whose formal semantics are specified using Evidence-Based Dynamic Epistemic Logic (EB-DEL). The second major contribution is a systematic annotation and formal model of grunts as grounding contributions. Drawing on approximately three hours of collaborative task interaction, an annotation scheme is presented that captures grunt tokens, their conversational triggers, and their dialogue act functions. The central finding is that grunts respond to speech and observable events at nearly equal rates, demonstrating that non-verbal events function as conversational contributions on par with utterances. Grunts are formalized as operations within EB-DEL that promote evidenced propositions to accepted facts in the common ground. The third contribution is a prosodic analysis showing that token selection is the primary signal of grounding source, with prosody operating as a secondary channel. Population-level acoustic differences are largely absorbed by token identity, but where the token is held constant, in standalone oh in the Portal corpus, event-following productions are significantly higher and wider in pitch than speech-following ones. This is exploratory evidence that speakers can modulate acoustics to signal what is being grounded. The fourth contribution extends the analysis to the structural dimension of turn position. Using a four-way positional typology (standalone, turn-initial, turn-medial, turn-final) annotated across all three corpora, position is shown to align with the same functional grouping found in the token and acoustic analyses acknowledgment tokens such as mm-hm are almost always standalone complete turns, while such as oh are predominantly turn openers. In the Weights Task and Glider, position also tracks grounding source; speech-following grunts tend to stand alone as complete turns, while event-following grunts disproportionately open turns. In the faster-paced Portal, this positional contrast flattens somewhat. Across all three encoding channels, the speech--event distinction is found to be marked redundantly rather than carried by any single feature. The Portal Dialogue Corpus, a structurally distinct virtual collaborative setting, serves as a test of generalization. Grunts are used differently to ground speech versus observable events across all three corpora, while their surface realization is likely shaped by the task. Taken together, this dissertation demonstrates that a complete account of grounding in multimodal dialogue must bring these small sounds into the formal and evidential picture. The grunt work, it turns out, is doing quite a lot.
pdf
Brandeis_University_Dissertation_RB_20265.41 MBDownloadView
Open Access

Metrics

1 Record Views

Details

Logo image