Self-hosted coding agents
Run an open-weight coding model within infrastructure your team controls, subject to its license and acceptable-use terms.
Independent tool overview
Laguna S 2.1 is a large open-weight mixture-of-experts model built for agentic coding and long-horizon software work. It combines roughly 8B active parameters per token with a 1M-token context window, but self-hosting still requires substantial storage, memory, and inference expertise.
Visit the official Laguna S 2.1 site ↗
Overview
Poolside positions Laguna S 2.1 as the middle model in its Laguna coding family: larger and more capable than Laguna XS 2.1, but smaller than Laguna M.1. Its main audience is engineering teams that want to run a coding model inside their own infrastructure or through a compatible inference provider.
The model has 118B total parameters with about 8B activated for each token. It supports tool-oriented, multi-step coding workflows, interleaved reasoning between tool calls, and a 1,048,576-token context window for large repositories and long agent trajectories.
The weights are available in BF16 and several quantized formats, including FP8, NVFP4, INT4, GGUF, and MLX variants. Poolside's main BF16 repository is about 235 GB, so the model is not a lightweight local download despite its relatively low active-parameter count.
Poolside releases Laguna S 2.1 under OpenMDW-1.1, which it describes as permissive for commercial and noncommercial use. Teams should still review the license, acceptable-use terms, security controls, and deployment dependencies before adopting it in production.
Use cases
The strongest fit depends on the job you need the product to complete, not the size of its feature list.
Run an open-weight coding model within infrastructure your team controls, subject to its license and acceptable-use terms.
Use the 1M-token context window for codebase exploration, cross-file changes, and long agent sessions.
Deploy through supported runtimes such as vLLM, SGLang, TensorRT-LLM, llama.cpp, or Ollama.
Test Poolside's model and quantized variants against your own repositories, tools, and software-engineering tasks.
Capabilities
The model is post-trained for multi-step software tasks and agent workflows rather than simple code completion alone.
A 1,048,576-token context window can accommodate large codebases and extended tool-use histories, although usable performance still depends on the serving stack.
Laguna S 2.1 has 118B total parameters while activating roughly 8B per token to balance capability and inference cost.
Its chat template supports interleaved thinking around tool calls and per-request reasoning control.
Poolside publishes BF16, FP8, NVFP4, INT4, GGUF, MLX, and speculative-decoding assets for different deployment targets.
Poolside describes OpenMDW-1.1 as allowing use, modification, and commercial products without separate permission.
Process
Step 1
Select BF16 for maximum fidelity or a supported quantized build that fits your available hardware and serving runtime.
Step 2
Follow the official model-card instructions for a compatible inference engine and secure the endpoint before connecting tools.
Step 3
Give the model tightly scoped repository, terminal, test, and search tools with explicit approval boundaries.
Step 4
Measure correctness, latency, cost, security behavior, and long-context reliability on representative internal tasks.
Step 5
Keep human review, logs, secret filtering, sandboxing, and rollback paths around any agent that can modify or execute code.
Cost
Poolside publishes Laguna S 2.1's weights without a separate model download fee. Your real cost depends on storage, accelerators, inference software, operations, and any third-party hosting; Poolside does not list a simple public per-token price on the release or model card.
No model download fee
Download the official checkpoints and run them in your own environment under OpenMDW-1.1.
Infrastructure-dependent
Pay for the accelerators, storage, bandwidth, observability, and engineering needed to serve the model.
Provider-specific
Compatible third-party inference services may offer Laguna S 2.1 with their own prices, limits, and retention policies.
Pricing checked . Check current pricing at the source ↗
Assessment
Compare
The right alternative depends on the specific output, workflow, controls and budget your project requires.
Coding
Consider Mistral's open coding-model family if you want another self-hostable agentic coding option.
Explore Devstral →Business Operations
Consider DeepSeek for a broader open-model ecosystem with hosted API access and competitive coding capabilities.
Explore DeepSeek →Consumer
Consider Claude Sonnet 5 if you prefer a managed frontier model and do not need to operate the weights yourself.
Explore Claude Sonnet 5 →Questions
Laguna S 2.1 is Poolside's 118B-total-parameter, open-weight mixture-of-experts model for agentic coding and long-horizon software work.
Yes, if your workstation or server can handle the selected checkpoint. Quantized variants make local deployment more practical, but this is still a large model with significant storage and memory requirements.
The official model card lists a 1,048,576-token context window. Actual speed, memory use, and task quality at long contexts depend on your serving setup and prompt.
The official weights can be downloaded without a separate model fee under OpenMDW-1.1. Hardware, hosting, bandwidth, and engineering still create real costs.
It is best described precisely as an open-weight model under Poolside's OpenMDW-1.1 license. Poolside says the license permits commercial and noncommercial use and modification; review the license and acceptable-use terms for your case.
Test it on representative repositories and agent tasks, then measure correctness, tool use, long-context reliability, latency, infrastructure cost, and security behavior before production use.
Bottom line
Laguna S 2.1 is a compelling option for teams that want a capable coding-agent model they can deploy and govern themselves. Its long context, quantized releases, and low active-parameter count are attractive, but the 118B checkpoint still demands serious infrastructure and careful agent security.
Visit Laguna S 2.1 website ↗
LM Bionic - LM Studio's local-first agent that codes and edits docs using open models

Cursor Router - Cursor's auto-router picking the cheapest capable model per task

Canva Code 2.0 - Canva's vibe-coding tool with drag-and-drop editing

MAI-Cyber-1-Flash - Microsoft's new cybersecurity model for securing codebases

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.
Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.
To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.
Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.
Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.