The Rundown AI homepage

Independent tool overview

Raven-1 at a glance

Raven-1 is Tavus's perception layer for real-time AI conversations, combining audio, video, and timing signals into natural-language descriptions of user context and intent.

Visit the official Raven-1 site ↗
Raven-1 product preview
Developer
Tavus
Product type
Real-time multimodal perception model
Inputs
Audio, video, and temporal dynamics
Output
Natural-language perception descriptions
Access
Included in Tavus conversations and APIs

Overview

What Raven-1 is

Raven-1 is not a standalone consumer app. It is the multimodal perception model built into Tavus conversations, where it processes vocal tone, prosody, facial expression, gaze, posture, gesture, visual context, and timing together. Its output is meant to help a downstream language model respond to how something was communicated—not only the transcript.

Instead of forcing each moment into a fixed emotion label, Raven-1 produces natural-language descriptions and updates them throughout a turn. Developers can also define perception events through an OpenAI-compatible tool schema, allowing an application to react when Raven detects a specified cue such as laughter, attention shifting, or rising frustration. These outputs remain machine inferences and should not be treated as objective readings of a person's inner state.

Use cases

Who Raven-1 is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Conversational video products

Give an AI video agent fresh context about visible and audible cues during a live interaction.

Adaptive coaching and training

Change explanations or offer help when the system infers confusion, hesitation, or disengagement, with human review and appropriate safeguards.

Customer support prototypes

Test whether multimodal cues can help an agent slow down, clarify, or escalate when an interaction appears to be deteriorating.

Perception-triggered workflows

Call application functions when narrowly defined cues occur, while keeping consequential decisions outside the model.

Capabilities

Core Raven-1 features

1

Audio-visual fusion

Combines tone, prosody, facial expression, posture, gaze, and other context in one representation rather than analyzing each channel independently.

2

Temporal modeling

Tracks how inferred emotional and attentional context changes at sentence-level granularity throughout a conversational turn.

3

Natural-language output

Produces nuanced descriptions that a downstream language model can consume directly instead of only numeric scores or fixed labels.

4

Rolling perception

Continuously refreshes its interpretation as a conversation progresses so responses can use recent context.

5

Visual-context analysis

Processes expressions, gaze, posture, gestures, surrounding context, and supported screen-sharing inputs.

6

Tool-call events

Lets developers define specified cues and receive callbacks through an OpenAI-compatible schema.

7

Tavus stack integration

Works alongside Sparrow-1 for conversational flow and Phoenix-4 for rendered facial behavior in Tavus's real-time video pipeline.

Process

How the Raven-1 workflow works

  1. Step 1

    Define a bounded use case

    Choose a specific interaction where perceptual context may improve the experience and document what the model must never decide.

  2. Step 2

    Prototype in Tavus

    Create a conversational experience through Tavus and confirm the required camera, audio, consent, and retention settings.

  3. Step 3

    Connect perception

    Use Raven's built-in context or configure narrowly scoped tool events for the application behavior you need.

  4. Step 4

    Test across people and conditions

    Evaluate different cultures, accents, lighting, cameras, disabilities, speaking styles, and levels of background noise.

  5. Step 5

    Add safe fallbacks

    Treat detections as uncertain signals, require confirmation for sensitive cases, and provide a clear path to a human.

  6. Step 6

    Monitor outcomes

    Measure false triggers, missed cues, user comfort, and downstream effects—not just perceived conversational realism.

Cost

Raven-1 pricing and free plan

Raven-1 is included as the perception and vision layer in Tavus developer plans rather than sold through a separate model subscription. Conversational minutes, concurrency, and overage rates vary by plan.

Basic

Free

Entry plan for testing Tavus's developer APIs and conversational video pipeline.

  • 25 conversational video minutes per month
  • 1 concurrent stream
  • 25 stock replicas
  • Raven-1 perception and vision included

Starter

$59/month

Paid developer plan for individuals, startups, and early product use.

  • 100 conversational video minutes
  • Up to 3 concurrent streams
  • $0.37 per additional conversational minute
  • 3 custom replica trainings per month

Growth

$397/month

Higher-volume plan for teams productionizing conversational experiences.

  • 1,250 conversational video minutes
  • Up to 10 concurrent streams
  • $0.32 per additional conversational minute
  • Conversation recordings
  • 7 custom replica trainings per month

Enterprise

Custom

Volume plan with negotiated scaling, concurrency, support, and service terms.

  • Volume discounts
  • Custom concurrency limits
  • White-label options
  • Enterprise support and SLAs

Pricing checked . Check current pricing at the source ↗

Assessment

Raven-1 strengths and limitations

Where it stands out

  • Uses vocal, visual, and timing information that disappears in a text transcript
  • Natural-language descriptions can preserve more nuance than a single emotion category
  • Rolling updates are designed for live conversation rather than offline-only analysis
  • Tool-call support makes perception usable in application logic
  • Integrated into Tavus's complete real-time conversational video stack
  • A free developer tier makes small prototypes possible

What to consider

  • Raven-1 is not available as an independent consumer tool or clearly priced standalone API
  • Emotion and intent cannot be observed directly; outputs are probabilistic interpretations of behavior
  • Performance can vary with culture, disability, neurotype, language, lighting, camera position, noise, and individual expression
  • Camera and microphone analysis creates significant consent, privacy, retention, and security obligations
  • It should not be the sole basis for healthcare, hiring, education, safety, or other consequential decisions
  • Tool-triggered automations can magnify false detections unless confirmation and rate limits are built in
  • Tavus's published latency and capability figures are vendor-reported and should be tested in the target environment

Compare

Raven-1 alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Marketing

D-ID

Offers AI avatars and conversational visual agents for customer-facing and training experiences.

Explore D-ID

Marketing

Synthesia

A more established avatar-video platform for teams prioritizing scripted training and business content over real-time perception.

Explore Synthesia

Questions

Raven-1 FAQs

What is Raven-1?

Raven-1 is Tavus's real-time multimodal perception model for combining audio, video, and timing signals into descriptions of user context, expression, and inferred intent.

Is Raven-1 a standalone app?

No. Tavus says Raven-1 is available across Tavus conversations and through its perception layer in Tavus APIs.

What does Raven-1 analyze?

Tavus lists signals such as vocal tone, prosody, facial expression, gaze, posture, gestures, visual context, and temporal changes during a conversation.

Does Raven-1 label emotions?

Its stated approach is to produce natural-language descriptions rather than reduce each moment to a single categorical emotion label.

How much does Raven-1 cost?

There is no separate Raven-1 price. It is included in Tavus developer plans, which currently range from a free Basic tier to $59 Starter, $397 Growth, and custom Enterprise pricing.

Can Raven-1 trigger application actions?

Yes. Tavus documents OpenAI-compatible tool-call events for developer-defined cues, but applications should require confirmation before sensitive or irreversible actions.

Can Raven-1 know how someone truly feels?

No model can directly know a person's internal state. Raven-1 infers context from observable signals, and those inferences can be wrong or biased.

Bottom line

Our Raven-1 verdict

Raven-1 is an ambitious perception layer for developers building live video agents, especially where tone and visible context affect conversational quality. Its value depends less on impressive demos than on careful validation: teams need to measure false interpretations, obtain meaningful consent, and prevent uncertain emotional inferences from becoming consequential decisions.

Visit Raven-1 website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.