Beyond chatbots: Tavus's Griffin model passes video Turing test with 48% human success rate
Unlike traditional conversational AI setups that rely on a stitched-together cascade of text transcription, large language model generation, and separate voice or video synthesis, Griffin operates as a single, full-duplex video-to-video pipeline.

- Oct 3, 2026,
- Updated Oct 3, 2026 10:06 AM IST
Machine interaction took a dramatic step toward human-like fluidity today as AI firm Tavus announced Griffin, an end-to-end Human Interaction Model (HIM) engineered to perceive and conduct real-time, face-to-face video conversations.
Unlike traditional conversational AI setups that rely on a stitched-together cascade of text transcription, large language model generation, and separate voice or video synthesis, Griffin operates as a single, full-duplex video-to-video pipeline.
The architecture allows the system to see, hear, interpret, speak, move, react, and respond simultaneously. Because its perception remains active even while speaking, any sudden visual change, pause, or verbal interruption instantly alters what the model says and does as the conversation unfolds.
"Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human," Tavus wrote in a announcement on X. "Previous systems have had a pass rate
To achieve seamless face-to-face dynamics, Griffin runs on a dual-engine structure: a Continuous Conversational Modeling engine that continuously assesses sub-second inputs to decide when to interject or back-channel with expressions like "mm-hm", and a unified Audio-Visual Generation engine that renders real-time speech, movement, and scene-level pixel adjustments.
The system goes beyond generating isolated facial movements. From a single reference image, Griffin renders entire dynamic scenes in real time — controlling everything from natural finger and arm gestures to shadows and background shifts.
It also factors in temporal awareness, intuitively understanding how long a silence has lasted to determine whether a user is taking a moment to gather their thoughts or waiting for the model to guide the next step.
Tavus has made an early research preview version, titled Griffin-Lite, available to a select group of developers, with plans for a broader rollout of the full model in the coming months.
For Unparalleled coverage of India's Businesses and Economy – Subscribe to Business Today Magazine
Machine interaction took a dramatic step toward human-like fluidity today as AI firm Tavus announced Griffin, an end-to-end Human Interaction Model (HIM) engineered to perceive and conduct real-time, face-to-face video conversations.
Unlike traditional conversational AI setups that rely on a stitched-together cascade of text transcription, large language model generation, and separate voice or video synthesis, Griffin operates as a single, full-duplex video-to-video pipeline.
The architecture allows the system to see, hear, interpret, speak, move, react, and respond simultaneously. Because its perception remains active even while speaking, any sudden visual change, pause, or verbal interruption instantly alters what the model says and does as the conversation unfolds.
"Introducing Griffin, the first model to pass the video Turing test. 48% of people who talked to it live thought it was a real human," Tavus wrote in a announcement on X. "Previous systems have had a pass rate
To achieve seamless face-to-face dynamics, Griffin runs on a dual-engine structure: a Continuous Conversational Modeling engine that continuously assesses sub-second inputs to decide when to interject or back-channel with expressions like "mm-hm", and a unified Audio-Visual Generation engine that renders real-time speech, movement, and scene-level pixel adjustments.
The system goes beyond generating isolated facial movements. From a single reference image, Griffin renders entire dynamic scenes in real time — controlling everything from natural finger and arm gestures to shadows and background shifts.
It also factors in temporal awareness, intuitively understanding how long a silence has lasted to determine whether a user is taking a moment to gather their thoughts or waiting for the model to guide the next step.
Tavus has made an early research preview version, titled Griffin-Lite, available to a select group of developers, with plans for a broader rollout of the full model in the coming months.
For Unparalleled coverage of India's Businesses and Economy – Subscribe to Business Today Magazine
