Private-codebase specialization
Train an agent on internal APIs, conventions, and repositories while retaining control of the data and resulting weights.
Independent tool overview
SERA is Ai2's open family of repository-level coding agents, released with model weights, training data, specialization code, and a Claude Code-compatible CLI.
Visit the official SERA site ↗
Overview
SERA stands for Soft-Verified Efficient Repository Agents. It is a family of open coding models from Ai2 designed for repository-level tasks such as debugging, code generation, review, maintenance, and explanation. Current releases include 8B, 14B, and 32B variants, plus variants trained from an earlier teacher model.
The more distinctive part of SERA is the reproducible data-generation and fine-tuning recipe. Teams can generate agent trajectories from their own repositories and specialize a smaller open model to internal APIs, conventions, and code patterns while controlling where the code and model run. Ai2 also provides a proxy that can launch SERA behind Claude Code through Modal, an existing vLLM endpoint, or self-hosted inference.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Train an agent on internal APIs, conventions, and repositories while retaining control of the data and resulting weights.
Inspect the released models, datasets, generation pipeline, verification metadata, and evaluation recipe.
Run an open repository agent behind a private vLLM endpoint instead of sending code to a closed hosted model.
Use the familiar Claude Code client with a SERA model through Ai2's compatibility proxy.
Explore synthetic data and supervised specialization without building a large reinforcement-learning stack from scratch.
Capabilities
Provides downloadable SERA-8B, SERA-14B, and SERA-32B weights, including additional teacher variants.
Models are trained for multi-step software tasks that require inspecting, editing, and testing code across a repository.
The released pipeline can generate training trajectories from arbitrary personal or internal codebases.
Ai2 releases generation code, configuration, datasets, metadata, and training instructions rather than only model weights.
The data pipeline documents support for SWE-agent and mini-swe-agent workflows.
The SERA CLI runs a proxy so Claude Code can use a SERA model served through vLLM.
A single CLI route can provision the model on Modal for one session and clean up the app when Claude Code exits.
A separate command can create a shared long-running SERA endpoint with an automatically generated API key.
Teams can serve SERA directly with vLLM on their own GPU infrastructure and point the proxy at that endpoint.
Ai2's release update adds verification thresholds and metadata intended to make filtering and downstream training easier.
Process
Step 1
Decide whether the need is evaluation, daily coding assistance, research reproduction, or specialization to a private repository.
Step 2
Balance benchmark results, latency, GPU memory, context needs, and concurrency across the 8B, 14B, and 32B variants.
Step 3
Use a sandboxed development environment with scoped repository access, short-lived credentials, resource limits, and no production secrets.
Step 4
Use the SERA CLI with Modal, connect it to an existing endpoint, or host the model through vLLM.
Step 5
Test representative issues from the target codebase and measure correct patches, test results, regressions, security findings, latency, and cost.
Step 6
Generate training data from approved repositories, filter and verify samples, fine-tune a selected model, and retain a clean holdout evaluation set.
Step 7
Require code review, tests, security scans, and explicit approval before any SERA-generated change reaches a protected branch or deployment.
Cost
SERA's model weights, CLI, and training code are available without a software subscription under open licenses, but running and fine-tuning the models requires GPU infrastructure. Ai2's published dollar figures are examples from its experiments, not fixed plans or guaranteed budgets.
Free to download
Open SERA weights, training repository, datasets, and CLI.
Infrastructure cost
Run SERA on owned or rented GPUs through vLLM.
Usage-based cloud cost
The CLI can provision ephemeral or persistent SERA endpoints on Modal.
Variable training and teacher cost
Generate synthetic trajectories and fine-tune on a target codebase.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
Mistral's open agentic coding model family offers another self-hostable route with its own CLI ecosystem.
Explore Devstral 2 →Coding
A managed coding-agent experience for teams prioritizing frontier model quality and a supported end-user workflow over open weights.
Explore Claude Code →Coding
Google's open-source terminal agent offers a lower-friction hosted-model workflow with high free usage allowances.
Explore Gemini CLI →Questions
SERA is Ai2's family of Soft-Verified Efficient Repository Agents: open coding models and a training recipe for repository-level software tasks.
Current model cards list SERA variants based on 8B, 14B, and 32B Qwen 3 models, plus variants trained from an earlier GLM teacher.
The weights, code, and CLI are free to download under their listed licenses. Users still pay for storage, GPU inference, cloud hosting, teacher-model calls, and any fine-tuning work.
Yes. Ai2's SERA CLI provides a compatibility proxy that lets Claude Code use a SERA model served on Modal, an existing endpoint, or self-hosted vLLM.
The released pipeline supports generating training trajectories from arbitrary repositories and fine-tuning a model to those patterns. Organizations must still protect secrets, verify data rights, and evaluate the specialized model independently.
Requirements depend on model size, context, precision, and concurrency. The SERA-14B model card describes one 80GB A100 or H100 for its documented 32K setup, while the 32B weights and high-throughput deployments require more capacity.
The current model card reports 49.5% plus or minus 1.9 for SERA-32B, 41.7% plus or minus 0.5 for SERA-14B, and 31.7% plus or minus 0.9 for SERA-8B at 32K context across three seeds. These are benchmark results, not guarantees on real repositories.
No. Ai2's own model cards describe the models as unsafety-tuned research artifacts. Use sandboxing, least privilege, tests, security review, and mandatory human approval.
Bottom line
SERA is one of the more complete open releases for teams that want to study or specialize a repository agent instead of treating coding AI as a closed API. Its openness is the opportunity and the burden: users gain control over data and weights, but also inherit GPU operations, evaluation, sandboxing, and the safety work that a managed service would normally absorb.
Visit SERA website ↗
NousCoder-14B - Nous Research's competitive olympiad programming model

Codex App - OpenAI's new Mac app interface for managing agents with its Codex agentic coding assistant

GLM-4.7 - Z AI's new SOTA open-source model with advanced coding and reasoning capabilities

Qwen3-Coder-Next - Alibaba's open-weight, small hybrid model for agentic coding

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.