Conversational video products
Give an AI video agent fresh context about visible and audible cues during a live interaction.
Independent tool overview
Raven-1 is Tavus's perception layer for real-time AI conversations, combining audio, video, and timing signals into natural-language descriptions of user context and intent.
Visit the official Raven-1 site ↗.png&w=3840&q=75)
Overview
Raven-1 is not a standalone consumer app. It is the multimodal perception model built into Tavus conversations, where it processes vocal tone, prosody, facial expression, gaze, posture, gesture, visual context, and timing together. Its output is meant to help a downstream language model respond to how something was communicated—not only the transcript.
Instead of forcing each moment into a fixed emotion label, Raven-1 produces natural-language descriptions and updates them throughout a turn. Developers can also define perception events through an OpenAI-compatible tool schema, allowing an application to react when Raven detects a specified cue such as laughter, attention shifting, or rising frustration. These outputs remain machine inferences and should not be treated as objective readings of a person's inner state.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Give an AI video agent fresh context about visible and audible cues during a live interaction.
Change explanations or offer help when the system infers confusion, hesitation, or disengagement, with human review and appropriate safeguards.
Test whether multimodal cues can help an agent slow down, clarify, or escalate when an interaction appears to be deteriorating.
Call application functions when narrowly defined cues occur, while keeping consequential decisions outside the model.
Capabilities
Combines tone, prosody, facial expression, posture, gaze, and other context in one representation rather than analyzing each channel independently.
Tracks how inferred emotional and attentional context changes at sentence-level granularity throughout a conversational turn.
Produces nuanced descriptions that a downstream language model can consume directly instead of only numeric scores or fixed labels.
Continuously refreshes its interpretation as a conversation progresses so responses can use recent context.
Processes expressions, gaze, posture, gestures, surrounding context, and supported screen-sharing inputs.
Lets developers define specified cues and receive callbacks through an OpenAI-compatible schema.
Works alongside Sparrow-1 for conversational flow and Phoenix-4 for rendered facial behavior in Tavus's real-time video pipeline.
Process
Step 1
Choose a specific interaction where perceptual context may improve the experience and document what the model must never decide.
Step 2
Create a conversational experience through Tavus and confirm the required camera, audio, consent, and retention settings.
Step 3
Use Raven's built-in context or configure narrowly scoped tool events for the application behavior you need.
Step 4
Evaluate different cultures, accents, lighting, cameras, disabilities, speaking styles, and levels of background noise.
Step 5
Treat detections as uncertain signals, require confirmation for sensitive cases, and provide a clear path to a human.
Step 6
Measure false triggers, missed cues, user comfort, and downstream effects—not just perceived conversational realism.
Cost
Raven-1 is included as the perception and vision layer in Tavus developer plans rather than sold through a separate model subscription. Conversational minutes, concurrency, and overage rates vary by plan.
Free
Entry plan for testing Tavus's developer APIs and conversational video pipeline.
$59/month
Paid developer plan for individuals, startups, and early product use.
$397/month
Higher-volume plan for teams productionizing conversational experiences.
Custom
Volume plan with negotiated scaling, concurrency, support, and service terms.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Marketing
A competing real-time interactive avatar platform for building face-to-face AI experiences.
Explore LiveAvatar by HeyGen →Marketing
Offers AI avatars and conversational visual agents for customer-facing and training experiences.
Explore D-ID →Marketing
A more established avatar-video platform for teams prioritizing scripted training and business content over real-time perception.
Explore Synthesia →Questions
Raven-1 is Tavus's real-time multimodal perception model for combining audio, video, and timing signals into descriptions of user context, expression, and inferred intent.
No. Tavus says Raven-1 is available across Tavus conversations and through its perception layer in Tavus APIs.
Tavus lists signals such as vocal tone, prosody, facial expression, gaze, posture, gestures, visual context, and temporal changes during a conversation.
Its stated approach is to produce natural-language descriptions rather than reduce each moment to a single categorical emotion label.
There is no separate Raven-1 price. It is included in Tavus developer plans, which currently range from a free Basic tier to $59 Starter, $397 Growth, and custom Enterprise pricing.
Yes. Tavus documents OpenAI-compatible tool-call events for developer-defined cues, but applications should require confirmation before sensitive or irreversible actions.
No model can directly know a person's internal state. Raven-1 infers context from observable signals, and those inferences can be wrong or biased.
Bottom line
Raven-1 is an ambitious perception layer for developers building live video agents, especially where tone and visible context affect conversational quality. Its value depends less on impressive demos than on careful validation: teams need to measure false interpretations, obtain meaningful consent, and prevent uncertain emotional inferences from becoming consequential decisions.
Visit Raven-1 website ↗
Model Council - Perplexity's new tool for querying and synthesizing outputs from multiple models into a single answer

Phoenix-4 - Tavus' real-time human rendering model with emotional intelligence and active listening

Voxtral Transcribe 2 - A new speech-to-text family for transcription across 13 languages, including an open-weights Realtime model for live transcription.

Gemini Embedding 2 - Google's multimodal model capable of searching across text, images, video, and audio at once

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.