Tavus Griffin AI Fooled 48% in Video Turing Test

San Francisco startup Tavus said Oct. 1 that its Tavus Griffin real-time video AI was mistaken for a human by 26 of 54 participants—about 48%—in one-minute live video calls, a company-run result it bills as a video Turing-test milestone, while keeping Griffin-Lite limited to a research preview over safety concerns.

Tavus Griffin’s 48% “human” result

According to Runtime Wire and THE DECODER, participants in the United States and Europe were told they would chat for a minute about what they looked forward to that year and only afterward asked whether their partner was real. The partner was a Griffin-Lite video persona. Under the same protocol, Tavus says its prior Phoenix-4.5 system fooled just 1 of 41 people (2.4%). The sample is small and company-run, so it is not a broad measure of everyday deception risk.

Full-duplex video, not just talking heads

Tavus describes Griffin as a “Human Interaction Model” that continuously processes incoming audio and video, decides when to speak or yield, and generates voice and picture together—playing Simon Says, coaching a Rubik’s Cube, or reacting to objects held up on camera. The firm says NVIDIA’s Video Full-Duplex Benchmark ranked Griffin-Lite first on generation (3.83/5 vs. 2.80 for the next system; human reference 3.92) and perception (3.73 vs. 3.44 baseline; human 4.20), with video-generation latency averaging 0.43 seconds on H100 GPUs. Those scores use a language-model judge and measure different things than the live perception study.

The announcement fits a crowded AI week that also includes frontier models such as Gemini 4 Argon and agent work like OpenAI Dots, plus hardware safety efforts around Nvidia’s open agent safety platform.

Why Tavus will not ship Griffin widely yet

CEO Hassaan Raza’s company argues computers should adapt to face-to-face human communication. Tavus raised a $40 million Series B in November 2025 (about $64 million total per THE DECODER) and already sells conversational video tools—but Griffin itself is not on those plans. The same realism that impressed testers can enable deception, so Tavus says disclosure features and further safety work must land before broader release. Potential uses listed include tutoring, practicing hard conversations, and camera-based tech support—if regulators and customers accept the risk tradeoff.

Sources: Runtime Wire; THE DECODER; Tavus Griffin.