Imagine directing a robot through a building in a conflict zone. Entering the structure would be too dangerous, and poor signal makes continuous video communication impossible. As a US Army teleoperator searching for mission-critical information, you piece together the fragments by asking the robot a series of spoken questions. Has anyone been here recently? What was this room used for? Do you see signs of foreign occupation?
With near-perfect recall and the capacity to navigate extreme environments, you might have assumed that an AI-assisted robot would be an ideal accomplice. But how can you ensure that it “wants” what you want – that it will identify the exact information you need?
Establishing common ground #
At the USC Institute for Creative Technologies (ICT), David Traum and Kallirroi Georgila work together to investigate the problem of how two fundamentally different kinds of intelligence – humans and artificial agents – can successfully communicate to achieve shared goals.
Traum, who directs ICT’s Natural Language Dialogue Group, is a research professor of computer science in the Thomas Lord Department of Computer Science (CS) within USC Viterbi School of Engineering and USC Stevens School of Computing and Artificial Intelligence. Drawing upon theories of linguistics, philosophy and psychology, he emphasises the centrality of “common ground” – the idea that a conversation is not simply an exchange of information but a dynamic construction of a shared mental model. That model accounts for what has already been said, what both (or multiple) participants know, what remains uncertain and what they are trying to achieve together.
Georgila, research associate professor of computer science at CS within USC Viterbi and USC Stevens, specializes in natural language dialogue, reinforcement learning, speech processing and user simulation; she develops computational methods that allow dialogue systems to adapt to a user’s goals, behavior and changing circumstances. ”By mapping out multiple human-system interaction scenarios in simulation, conversational AI agents can safely learn from failure before deployment, which is particularly important in high-stakes situations,” she explained.
Sight unseen #
Traum and Georgila’s recent collaboration with the US Army Research Laboratory (ARL), recognized with the Best Paper Award at MAGMaR, examines how robots can identify and communicate the observations needed to answer human questions during remote exploration.
To study that problem, the researchers recreated the environmental constraints of an interior space in a zone with limited signal – imagine an apartment in a partly bombed building, which might once have housed a sniper or served as enemy headquarters.
Participants guided a mobile robot through the environment without seeing a continuous video feed; instead, they relied on spoken dialogue, a map of the robot’s movement and occasional still images requested by the participant during the mission.
“Robots are not humans, and so you wouldn’t talk to them exactly the way you would talk to people,” said Traum. “The flip side is, how would you engineer a robot that’s going to be able to talk to people the way people want to talk to it?”
Rather than waiting for a fully autonomous system to exist, the team adopted a “Wizard of Oz” methodology, in which researchers temporarily performed capabilities the robot would eventually be expected to carry out itself (think of the famous curtain scene). This involved a researcher controlling the robot’s movement while another managed the dialogue with the participant, allowing the team to observe how people naturally asked questions, clarified misunderstandings and referred to places the robot had visited. Those interactions were collected as Video-SCOUT, a dataset of sixty 20-minute robot exploration videos accompanied by human dialogue.
The USC and ARL researchers developed the AI framework Non-Event Oriented Video Assessments (NOVA) to search the footage and identify the frames most relevant to a particular question. “We tested NOVA against different AI frameworks to compare how well it performed at retrieving the right frames,” said Georgila. “Combined with human judgement, we found that the systems performed better than either humans or automated models working independently.”
Better together #
If humans and intelligent systems achieve the strongest outcomes by working together, the next question is no longer how either should operate independently, but how they should communicate when each possesses different kinds of knowledge. In the case of the MAGMaR paper, the robot contributes perception and computational search; the human contributes interpretation and judgement. That finding reflects the broader scientific direction of Traum and Georgila’s research. Over more than two decades, they have explored dialogue across an expanding range of settings: between humans and robots; across speech, vision and spatial information; between multiple conversational partners and multiple conversational channels; and among teams of humans and intelligent agents working towards a shared objective. Although the applications differ, each asks the same question: how should communication evolve when knowledge is distributed across people, machines and modalities?
Published on August 20th, 2026
Last updated on August 20th, 2026