Personalized assistants
Carry durable user preferences, history, relationships, and current account context across many conversation threads.
Independent tool overview
Zep is context-engineering infrastructure for AI agents. It turns chat, business events, documents, and other text or JSON into temporal context graphs, then retrieves relevant, provenance-linked memory for prompts and agent tools.
Visit the official Zep site ↗
Overview
Zep is built for developers and organizations that need an AI agent to remember more than the current conversation. Applications send user messages and business data to Zep, which extracts entities and time-aware facts into a Context Graph. The application can then request a ready-to-use context block or search the graph directly.
The main distinction is temporal memory: when new information supersedes an old fact, Zep preserves when each fact was valid instead of treating every statement as equally current. This is useful for assistants that need evolving customer preferences, account state, support history, transactions, decisions, or other changing relationships.
Zep Cloud is the commercial managed service, with self-serve and enterprise deployment options. Graphiti is Zep's separate Apache-2.0 open-source framework for building temporal knowledge graphs. Choosing Graphiti means operating the graph database, model integrations, indexing, scaling, and governance yourself; it is not a free self-hosted edition of the current managed Zep platform.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Carry durable user preferences, history, relationships, and current account context across many conversation threads.
Unify conversations with CRM, ticket, product-usage, billing, and other business events that change over time.
Track both current and superseded relationships when a flat chat buffer or static document-retrieval system loses important temporal context.
Use hosted ingestion, graph construction, retrieval, observability, and production deployment controls without assembling the entire graph stack internally.
Capabilities
Extracts entities, relationships, facts, and source episodes while recording when facts became valid or were invalidated.
Stores session messages and returns a context string, relevant facts, and recent messages for the next model call.
Adds text or structured business data and searches nodes or edges when teams need more control than the opinionated Memory API provides.
Maintains a graph across all sessions for an individual user or creates group graphs for shared organizational and domain knowledge.
Builds token-conscious context blocks from relevant summaries, facts, entities, episodes, and other context types, with customizable templates on supported plans.
Supports custom entity and edge types plus extraction instructions so graph construction can reflect a product's domain.
Enterprise options include managed cloud, customer-managed encryption keys, bring-your-own-cloud deployment, audit controls, retention policies, and data-access controls.
Provides a separately operated Python framework and MCP server for teams that want to build temporal knowledge graphs on supported graph databases.
Process
Step 1
Decide whether each graph represents a user, account, team, or topic, and specify which applications and roles may read or write it.
Step 2
Provision identifiers and add enough authorized identity metadata for messages and business data to map to the correct subject.
Step 3
Send ordered chat turns plus relevant text or JSON episodes from approved sources, including timestamps and provenance where available.
Step 4
Use the Memory API for a prepared context block or the Graph API for custom search, and pass retrieved content to the model as untrusted data rather than privileged instructions.
Step 5
Test fact extraction, invalidation, retrieval quality, latency, isolation, deletion, prompt-injection resistance, and cost with representative workloads before expanding access.
Cost
Zep uses ingestion credits based on Episode size. An Episode up to 350 bytes costs one credit, with another credit for each additional 350 bytes or part; retrieval, storage, graph storage, memories, and users are unmetered. Prices below are the monthly self-serve rates shown on September 1, 2026. Annual billing advertises a 17% saving, and enterprise pricing is negotiated.
$0 for prototyping
A limited monthly allowance for testing Zep before choosing a paid plan.
$125/month
The entry self-serve production plan with automatic credit top-ups.
$375/month
Higher-volume self-serve access with more graph customization and operational features.
Custom
Negotiated capacity, security, support, and deployment for production workloads.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
Choose LangChain when you want a broad open-source framework for composing agents and prefer to assemble memory, retrieval, storage, and evaluation components yourself.
Explore LangChain →Data Analysis
Choose Pinecone when the core requirement is managed vector search and conventional RAG rather than an automatically constructed temporal knowledge graph.
Explore Pinecone →Business Operations
Choose n8n when the primary problem is connecting applications and orchestrating AI workflows; it can coordinate external storage but does not replace Zep's specialized temporal memory layer.
Explore n8n →Questions
Zep ingests conversations and business data, constructs temporal context graphs, and returns relevant memory or graph-search results that an application can provide to an AI agent.
Not primarily. Zep combines semantic and full-text retrieval with graph relationships, time-aware facts, source episodes, and context assembly. A vector database is a simpler alternative when document similarity search is the main need.
The current hosted Zep platform is a commercial service. Graphiti, the temporal knowledge-graph framework associated with Zep, is open source under Apache 2.0 and must be deployed and operated separately.
As reviewed September 1, 2026, prototyping includes 10,000 free credits per month. Flex is $125 per month with 50,000 credits, Flex Plus is $375 per month with 200,000 credits, and Enterprise is custom-priced. Ingestion credits scale with Episode size.
Zep documents official SDK quickstarts for Python, TypeScript, and Go. Its API can work with an agent framework or a custom application.
No. Zep's documentation recommends sending recent messages alongside the retrieved long-term context because newly ingested information may not yet appear in the graph.
Zep advertises SOC 2 Type II certification and offers HIPAA Business Associate Agreements on Enterprise, plus BYOK and BYOC options. A team must still perform its own legal, security, privacy, data-flow, access, and retention review before sending sensitive data.
Yes. Entity extraction, relationship resolution, temporal invalidation, and retrieval are model-assisted processes. Test representative cases, preserve provenance, add access checks, and require human review wherever an incorrect fact could cause material harm.
Bottom line
Zep is a strong fit when an AI product needs durable, evolving context across conversations and business systems and the team wants a managed memory layer rather than a hand-built retrieval stack. Its temporal graph, source provenance, high- and low-level APIs, and enterprise deployment controls are meaningful advantages. It is more infrastructure—and potentially more cost and governance work—than a simple chatbot or vector-search project needs, so validate extraction quality, latency, prompt-injection defenses, deletion behavior, and credit consumption with a production-shaped pilot.
Visit Zep website ↗
Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.