{"slug": "how-to-see-eye-to-eye-with-ai", "title": "How to See Eye to Eye With AI", "summary": "Researchers at the USC Institute for Creative Technologies, in collaboration with the US Army Research Laboratory, won the Best Paper Award at MAGMaR for a study on how robots can communicate observations to human teleoperators in conflict zones. The study, led by David Traum and Kallirroi Georgila, used a 'Wizard of Oz' methodology to simulate robot dialogue, finding that spoken interaction can help humans and robots establish common ground even without continuous video. The research aims to improve human-robot communication in high-stakes environments.", "body_md": "Imagine directing a robot through a building in a conflict zone. Entering the structure would be too dangerous, and poor signal makes continuous video communication impossible. As a US Army teleoperator searching for mission-critical information, you piece together the fragments by asking the robot a series of spoken questions. Has anyone been here recently? What was this room used for? Do you see signs of foreign occupation?\n\nWith near-perfect recall and the capacity to navigate extreme environments, you might have assumed that an AI-assisted robot would be an ideal accomplice. But how can you ensure that it “wants” what you want – that it will identify the exact information you need?\n\n## Establishing common ground\n\nAt the [USC Institute for Creative Technologies](https://ict.usc.edu/) (ICT), [David Traum](https://viterbi.usc.edu/directory/faculty/Traum/David) and [Kallirroi Georgila](https://viterbi.usc.edu/directory/faculty/Georgila/Kallirroi) work together to investigate the problem of how two fundamentally different kinds of intelligence – humans and artificial agents – can successfully communicate to achieve shared goals.\n\nTraum, who directs ICT’s [Natural Language Dialogue Group,](https://nld.ict.usc.edu/) is a research professor of computer science in the [Thomas Lord Department of Computer Science](https://www.cs.usc.edu/) (CS) within [USC Viterbi School of Engineering](https://viterbischool.usc.edu/) and [USC Stevens School of Computing and Artificial Intelligence](https://stevens-computing-ai.usc.edu/). Drawing upon theories of linguistics, philosophy and psychology, he emphasises the centrality of “common ground” – the idea that a conversation is not simply an exchange of information but a dynamic construction of a shared mental model. That model accounts for what has already been said, what both (or multiple) participants know, what remains uncertain and what they are trying to achieve together.\n\nGeorgila, research associate professor of computer science at CS within USC Viterbi and USC Stevens, specializes in natural language dialogue, reinforcement learning, speech processing and user simulation; she develops computational methods that allow dialogue systems to adapt to a user’s goals, behavior and changing circumstances. ”By mapping out multiple human-system interaction scenarios in simulation, conversational AI agents can safely learn from failure before deployment, which is particularly important in high-stakes situations,” she explained.\n\n## Sight unseen\n\nTraum and Georgila’s recent [collaboration](https://aclanthology.org/2026.magmar-main.8/) with the [US Army Research Laboratory](https://arl.devcom.army.mil/) (ARL), recognized with the Best Paper Award at [MAGMaR](https://nlp.jhu.edu/magmar/), examines how robots can identify and communicate the observations needed to answer human questions during remote exploration.\n\nTo study that problem, the researchers recreated the environmental constraints of an interior space in a zone with limited signal – imagine an apartment in a partly bombed building, which might once have housed a sniper or served as enemy headquarters.\n\nParticipants guided a mobile robot through the environment without seeing a continuous video feed; instead, they relied on spoken dialogue, a map of the robot’s movement and occasional still images requested by the participant during the mission.\n\n“Robots are not humans, and so you wouldn’t talk to them exactly the way you would talk to people,” said Traum. “The flip side is, how would you engineer a robot that’s going to be able to talk to people the way people want to talk to it?”\n\nRather than waiting for a fully autonomous system to exist, the team adopted a “Wizard of Oz” methodology, in which researchers temporarily performed capabilities the robot would eventually be expected to carry out itself (think of the famous [curtain scene](https://www.google.com/search?q=wizard+of+oz+man+behind+the+curtain&sca_esv=121635e38b0479b5&rlz=1C5GCEM_enUS1158US1158&sxsrf=APpeQnv6zBVDnRCo92IQ7Ajsx3Vl4ThTlQ%3A1785798439245&ei=Jx9xatzTDo_AkPIP29eYuQo&biw=1457&bih=840&oq=wizard+&gs_lp=Egxnd3Mtd2l6LXNlcnAiB3dpemFyZCAqAggBMgsQLhiABBiKBRiRAjIKEAAYgAQYigUYQzILEAAYgAQYigUYkQIyCxAAGIAEGIoFGJECMg0QLhiABBiKBRhDGLEDMgsQABiABBiKBRiRAjILEC4YgAQYigUYkQIyChAAGIAEGIoFGEMyChAAGIAEGIoFGEMyCBAAGIAEGLEDMhoQLhiABBiKBRiRAhiXBRjcBBjeBBjgBNgBAUj3FVAAWOgKcAB4AJABAJgBfaABmwWqAQM1LjK4AQHIAQD4AQGYAgegArwFwgIEECMYJ8ICDRAjGKIHGJ4GGPAFGCfCAgoQLhhDGIAEGIoFwgIKEC4YgAQYigUYQ8ICEBAAGIAEGIoFGEMYsQMYgwHCAhAQLhiABBiKBRhDGLEDGIMBwgIOEC4YgAQYigUYsQMYgwHCAhQQABiABBiKBRiRAhixAxiDARjwBsICCBAuGIAEGLEDwgILEC4YgAQYsQMYgwGYAwC6BgYIARABGBSSBwM0LjOgB_VxsgcDNC4zuAe8BcIHBTAuMi41yAcZgAgB&sclient=gws-wiz-serp#fpstate=ive&vld=cid:e1123133,vid:YWyCCJ6B2WE,st:0)). This involved a researcher controlling the robot’s movement while another managed the dialogue with the participant, allowing the team to observe how people naturally asked questions, clarified misunderstandings and referred to places the robot had visited. Those interactions were collected as [Video-SCOUT](https://uscedu-my.sharepoint.com/:w:/r/personal/bathurst_usc_edu/Documents/USC%20Articles/Articles/ICT/Traum_Georgila/Drafts/How_to_See_Eye_to_Eye_With_AI_revised_final.docx?d=we4eb3f69a666487ab577e52593cc2d40&csf=1&web=1&e=3Rjctb&nav=eyJoIjoiODQzMTUxMDYzIn0), a dataset of sixty 20-minute robot exploration videos accompanied by human dialogue.\n\nThe USC and ARL researchers developed the AI framework Non-Event Oriented Video Assessments (NOVA) to search the footage and identify the frames most relevant to a particular question. “We tested NOVA against different AI frameworks to compare how well it performed at retrieving the right frames,” said Georgila. “Combined with human judgement, we found that the systems performed better than either humans or automated models working independently.”\n\n## Better together\n\nIf humans and intelligent systems achieve the strongest outcomes by working together, the next question is no longer how either should operate independently, but how they should communicate when each possesses different kinds of knowledge. In the case of the MAGMaR paper, the robot contributes perception and computational search; the human contributes interpretation and judgement.\n\nThat finding reflects the broader scientific direction of Traum and Georgila’s research. Over more than two decades, they have explored dialogue across an expanding range of settings: between humans and robots; across speech, vision and spatial information; between multiple conversational partners and multiple conversational channels; and among teams of humans and intelligent agents working towards a shared objective. Although the applications differ, each asks the same question: how should communication evolve when knowledge is distributed across people, machines and modalities?\n\nPublished on August 20th, 2026\n\nLast updated on August 20th, 2026", "url": "https://wpnews.pro/news/how-to-see-eye-to-eye-with-ai", "canonical_source": "https://viterbischool.usc.edu/news/2026/08/how-to-see-eye-to-eye-with-ai/", "published_at": "2026-08-20 16:00:24+00:00", "updated_at": "2026-08-20 16:15:32.258338+00:00", "lang": "en", "topics": ["artificial-intelligence", "natural-language-processing", "robotics"], "entities": ["USC Institute for Creative Technologies", "US Army Research Laboratory", "David Traum", "Kallirroi Georgila", "USC Viterbi School of Engineering", "USC Stevens School of Computing and Artificial Intelligence", "MAGMaR"], "alternates": {"html": "https://wpnews.pro/news/how-to-see-eye-to-eye-with-ai", "markdown": "https://wpnews.pro/news/how-to-see-eye-to-eye-with-ai.md", "text": "https://wpnews.pro/news/how-to-see-eye-to-eye-with-ai.txt", "jsonld": "https://wpnews.pro/news/how-to-see-eye-to-eye-with-ai.jsonld"}}