# Tavus Griffin: AI Passes the Video Turing Test — Act Now

> Source: <https://byteiota.com/tavus-griffin-video-turing-test/>
> Published: 2026-10-03 18:16:48+00:00

Tavus launched Griffin-Lite on October 1, 2026 — and 48% of participants in a one-minute video call believed they were speaking to a real person. The previous best for any AI video system was 2.4%. That jump happened in a single product cycle. The headline will get filed under “AI getting creepy,” but the actual story is narrower and more urgent: **the video call is no longer a reliable identity signal**, and developers who treat it as one are building on a broken assumption.

## What Tavus Griffin Actually Is

Griffin is not a better avatar. It is a different class of model — Tavus calls it a Human Interaction Model (HIM), and the architecture justifies the framing. Previous AI video systems used a cascade: speech-to-text, then language model, then text-to-speech, then avatar renderer, each stage introducing latency and dropping signal. Griffin runs perception, decision-making, and generation simultaneously. It watches you. It decides when to respond. It generates face, voice, and background — including shadows and chair movement — in a single unified system.

Audio-to-video latency sits at 0.43 seconds on H100s, half the next-fastest method. On [NVIDIA’s independent VideoFDB benchmark](https://www.nvidia.com/en-us/research/video-fdb/), Griffin-Lite scores 3.83 out of 5 on the generation track — the human reference is 3.92. On emotional matching specifically, it scores 4.40, which actually exceeds the human reference of 4.14. These are not Tavus’s own numbers; NVIDIA runs VideoFDB independently.

The Turing test study itself has caveats worth knowing. Tavus ran it — not a third party — the sample was 54 people, and participants were told upfront they’d be speaking with “another participant,” a framing that primes the answer. One minute is a short window; Griffin’s response latency runs roughly two seconds behind human conversation timing, and doubters typically identified it within 20 seconds. The 48% headline is real but overstated. The NVIDIA benchmark is more honest, and it still puts Griffin closer to human-level than any previous AI video system.

## What Developers Can Build Right Now

Griffin itself is not available through the [Tavus API](https://www.tavus.io/cvi) yet — it is a research preview for select testers only. What is available is the existing Conversational Video Interface (CVI), currently used by more than 150,000 developers. And its capabilities are already significant.

A Tavus PAL (their term for an AI video persona) can be given its own email address. It accepts calendar invites. It joins Google Meet, Zoom, or Microsoft Teams roughly one minute before start time, with a face and voice of your choosing. It can join calls already in progress. Tool calling lets it pull from calendars, CRMs, and document signing services during the conversation.

This is live today — not Griffin, the current stack. Which means the enterprise concern is not hypothetical: AI personas that look and sound like specific people can already show up in your video meetings. The one thing preventing arbitrary abuse is an allowlist that controls who can send calendar invites to a PAL. **That allowlist is optional. It is not on by default.**

## The Safety Problem Tavus Acknowledges

Tavus is holding Griffin back from customers. Their stated reason: *“The same properties that make Human Interaction Models powerful interfaces for natural communications allow them to deceive.”* They are building disclosure features — some form of in-call indicator that the participant is AI — before broadening access.

That is a responsible call and worth noting, because it is rare. The concern is real: tech journalist Kevin Roose described Griffin as a [“trick old people into handing over their bank passwords machine.”](https://protos.com/tavus-call-bot-sparks-ai-scam-psychosis-fears/) That is hyperbolic, but the underlying risk is not. AI-driven social engineering that was once limited to voice and text now extends to live video — and the attack surface is larger than most enterprise threat models currently account for. It is the same category of problem as [AI agents making autonomous decisions that bypass existing security policies](https://byteiota.com/pixelleak-ai-coding-agents-leaked-13000-internal-screenshots-to-github/): the tool does exactly what it is optimized to do, and the damage comes from a threat model that hadn’t caught up.

The EU AI Act and NYC’s AI regulation are already moving toward mandatory disclosure requirements for AI-driven interactions. Developers building on video AI systems should track these — the compliance window is shorter than it looks.

## Three Things to Do Now

Griffin is not in your stack yet, but the threat model shift it represents is already relevant:

- **Stop treating video presence as identity proof.** Enterprise workflows that use video calls for verification — hiring, financial approvals, access grants — need an additional authentication layer. A face on a video call is no longer sufficient.
- **If you are building on Tavus CVI, audit your allowlist.** The conferencing allowlist that controls who can invite your PAL to a call is optional and not enabled by default. Leaving it open means anyone with the PAL’s email address can add it to a call. Check[Tavus’s CVI documentation](https://docs.tavus.io/sections/conversational-video-interface/faq) for allowlist configuration.
- **Update your social engineering threat model.** Spear phishing that includes a “video call with the CEO” is now technically feasible at a per-call cost well under what the attack is worth. Brief your security team before this is news to them the hard way.

Full Griffin, whenever it ships, will be more capable than Griffin-Lite. Tavus’s safety work is genuine, but they have no control over what happens when comparable technology ships from a less cautious provider. Plan for the world where this is available broadly, not just the one where Tavus controls the gate.
