The Rundown AI homepage

Independent tool overview

Router at a glance

Ramp Router is an active, U.S.-only beta LLM gateway at router.com. It gives developers OpenAI Responses, Chat Completions, and Anthropic Messages-compatible endpoints for supported models across multiple providers, then adds fallbacks, cost-aware Flex selection, benchmark routing, coding-agent escalation, shadow evaluations, prompt caching, and request-level spend logs. The routing layer is free through 2026; users pay the published token price for model calls and receive $26 in introductory model credits.

Visit the official Router site ↗
Router product preview
Status
Active beta
Availability
Individual developers and teams in the U.S.
Ramp account required
No
API formats
Responses, Chat Completions, and Anthropic Messages
Routing fee
Free through 2026
Introductory credit
$26 in model credits
Content retention
One year by default; future recording can be disabled

Overview

What Router is

Router sits between an application and model providers. Instead of maintaining separate integrations for OpenAI, Anthropic, xAI, Fireworks, and other supported catalogs, a team sends requests to Router's API and receives a consistent response shape. A direct model can remain pinned, a prioritized model list can provide fallback, or a Router-managed strategy can choose among eligible options.

Its cost controls operate at several layers. Cost-efficient routing keeps the same model and provider but selects lower-priced Flex capacity when observed latency is acceptable. Benchmark aliases distribute traffic among a changing weighted set of models for a workload. NVIDIA NeMo Switchyard can keep routine coding-agent turns on an efficient model and escalate harder or error-heavy turns. Shadow mode copies sampled production requests to candidate models so teams can compare cost and latency before changing the primary route.

Router is not merely an invisible pass-through. It authenticates and screens requests, records operational metadata, can bill through shared provider credentials, and—by default—stores model inputs, outputs, and tool calls for one year. Accounts can disable future content recording after the setting propagates, but that does not delete existing archives. That data path, the U.S.-only processing footprint, and the rapidly changing beta strategies require the same diligence as adopting any other production AI infrastructure provider.

Use cases

Who Router is best for

The strongest fit depends on the job you need the product to complete, not the size of its feature list.

Multi-model application teams

Use one integration to evaluate and call models from several providers without changing the application for every model release.

High-volume inference cost control

Move eligible traffic to Flex capacity or lower-cost models when measured quality and latency make the savings worthwhile.

Coding-agent workloads

Use Switchyard to route routine turns to an efficient model and escalate turns with errors, difficult tools, failed tests, or deeper context.

Production model evaluation

Shadow a sample of real requests to candidate models and compare their outputs, cost, and latency without changing the user-facing answer.

Reliability through fallbacks

Define model candidates so an eligible request can continue when the preferred provider fails, rate-limits, or times out.

Teams needing per-request economics

Inspect the model, provider, credential source, service tier, tokens, latency, cost, and fallback attempts behind each call.

Capabilities

Core Router features

1

Multi-provider API

Exposes OpenAI Responses, OpenAI Chat Completions, and Anthropic Messages surfaces through one gateway and account.

2

Direct model selection

Pin a concrete supported model and inspect its published input, output, context, reasoning, and Fast-tier details.

3

Ordered fallbacks

Provide a candidate list or use an alias so Router can try another eligible model when the first fails before response output begins.

4

Cost-efficient Flex routing

Automatically selects a lower-priced Flex service tier for the same provider and model when recent performance suggests latency will remain acceptable.

5

Benchmark routing

Uses Router-owned aliases with weighted model mixes that can follow current benchmark results without requiring an application release.

6

Switchyard for coding agents

Scores signals such as error severity, test results, tool patterns, context compaction, and turn depth to choose an efficient or frontier tier.

7

Shadow models

Mirrors a chosen percentage of eligible Responses API traffic to one to three candidate models for non-blocking cost and latency comparison.

8

Prompt caching

Reuses provider-cached prompt prefixes where supported to reduce repeated-input cost and latency.

9

Spend controls and logs

Tracks usage and cost per request and documents controls for budgets, maximum cost, timeouts, and debugging.

10

Bring your own key

Supports encrypted, write-only provider keys for OpenAI, Fireworks, and xAI, with configurable shared-key fallback.

Process

How the Router workflow works

  1. Step 1

    Inventory request classes

    Separate deterministic extraction, summarization, coding, tool use, reasoning, and user-facing generation because each needs different quality and latency tests.

  2. Step 2

    Create a key and disable unnecessary recording

    Before sending sensitive traffic, decide whether one-year content storage is acceptable and turn off future recording if it is not required.

  3. Step 3

    Start with a pinned model

    Change the base URL while keeping the existing model and response contract stable, then validate streaming, tools, errors, retries, and accounting.

  4. Step 4

    Add explicit fallbacks

    Choose candidates with compatible context windows, tools, output behavior, policy, and data handling; test a simulated outage and rate limit.

  5. Step 5

    Set latency and spend bounds

    Use provider timeouts, header timeouts, maximum-cost controls, usage alerts, and application-level circuit breakers.

  6. Step 6

    Collect a quality baseline

    Score representative prompts with deterministic tests, human review, business outcomes, safety checks, latency percentiles, and full cost.

  7. Step 7

    Test routing strategies gradually

    Enable Flex, benchmark aliases, Switchyard, or shadowing on a narrow low-risk cohort and compare against the pinned baseline.

  8. Step 8

    Monitor route drift

    Re-read the model catalog, watch alias candidates and deprecations, inspect fallbacks, and rerun evaluations whenever models or routing logic change.

Cost

Router pricing and free plan

Router charges no routing markup through 2026. Requests served with Router's shared credentials are billed at the published per-token rate for the selected model or service tier. New accounts receive $26 in promotional model credits. Requests served through a user's own supported provider key are billed by that provider instead, while shadow candidate calls are currently non-billable.

Routing layer beta

$0 through 2026

No separate Router markup during the published launch period.

  • Available to U.S. individuals and teams
  • A future routing price has not been announced
  • No Ramp finance-product account required
  • Enterprise features are still coming soon

Shared-key model usage

Provider list price

Router bills the input and output tokens, service tier, long context, and other billable features of the model that actually serves the request.

  • Rates vary by model
  • Fast and Flex tiers can differ from base rates
  • Cached tokens can have separate pricing
  • GET /v1/models is authoritative for the account's callable catalog

Launch credits

$26 free credits

Introductory balance for trying supported model calls through Router.

  • Subject to offer terms
  • Not an ongoing monthly allowance
  • Consumed by billable model usage
  • Check the account for balance and expiration

GPT-5.6 Sol launch promotion

50% off through September 18, 2026

A time-limited discount announced with the Router launch.

  • Applies only under the offer terms
  • Expires after the stated date
  • Do not model long-term unit economics from the promotional rate
  • Confirm the discount in request cost logs

Bring your own key

Billed by provider

Router does not charge for requests served using the user's OpenAI, Fireworks, or xAI key.

  • Provider bills the account directly
  • Router still estimates list cost in its logs
  • Anthropic and Google BYOK are not currently supported
  • A shared-key fallback is billed by Router unless disabled

Shadow model calls

$0 during beta

Candidate responses generated through Shadow Mode are currently not billed to the user.

  • Primary call keeps normal billing
  • Requires content recording
  • Unavailable on keys with BYOK credentials
  • Agreement scoring is marked coming soon

Pricing checked . Check current pricing at the source ↗

Assessment

Router strengths and limitations

Where it stands out

  • OpenAI- and Anthropic-compatible surfaces can reduce the initial migration to a base-URL and key change.
  • The public model catalog exposes rates, context, output, reasoning, and accelerated-tier availability.
  • Flex routing can lower cost without switching the requested model or provider.
  • Benchmark aliases and Switchyard address workload fit and turn difficulty rather than always sending every request to one frontier model.
  • Shadow mode evaluates candidates on real traffic without changing or delaying the primary response.
  • Fallbacks reduce dependence on a single provider's availability and rate limits.
  • Request logs make the route, credential, retries, token use, latency, and cost visible.
  • BYOK preserves direct provider billing for three supported providers.
  • There is no routing markup during 2026, and the initial $26 credit lowers evaluation cost.
  • Ramp says the underlying infrastructure already carries its own production AI traffic at very large scale.

What to consider

  • Router is a beta available only to individual developers and teams in the United States.
  • Pricing after the free-routing period ends in 2026 has not been announced.
  • Content recording is enabled by default and retains model inputs, outputs, and tool calls for one year unless future recording is disabled.
  • Turning off recording does not delete existing content archives, and operational metadata can still be retained.
  • All models currently run on U.S.-based provider infrastructure; residency choices are limited and provider-dependent.
  • BYOK is limited to OpenAI, Fireworks, and xAI, with no current Anthropic or Google Vertex AI key support.
  • Shared-key fallback is enabled by default for a saved provider key and can create a Router charge after a key or provider failure unless disabled.
  • Cost-efficient, benchmark, shadow, and Switchyard routing are beta strategies whose behavior and model mixes can change.
  • Automatic model switching can alter tone, tool behavior, safety decisions, output schemas, and quality even when the API surface remains stable.
  • Benchmark aliases use the smallest context and output limits across candidates and must be refreshed from the catalog rather than hard-coded.
  • Shadow mode sends real recorded prompts to additional providers and is unavailable when BYOK keys are attached.
  • The compatibility layer rejects unsupported semantics such as log probabilities or multiple completions instead of translating them.
  • The current provider and model catalog is smaller than some established gateways, and several homepage integrations are still marked coming soon.
  • Ramp's savings, latency, reliability, and scale figures are vendor-reported and should be reproduced on the buyer's own traffic.

Compare

Router alternatives

The right alternative depends on the specific output, workflow, controls and budget your project requires.

Coding

Helicone

Choose Helicone when provider-agnostic LLM observability, prompt analytics, evaluations, and gateway controls are the primary need.

Explore Helicone

Coding

Together AI

Consider Together AI for direct hosted inference, fine-tuning, and deployment of a broad open-model catalog rather than a neutral multi-provider gateway.

Explore Together AI

Agents

MiniMax Agent

Use MiniMax Agent when you want a finished web and desktop work agent rather than infrastructure for routing your own application's model calls.

Explore MiniMax Agent

Questions

Router FAQs

What is Ramp Router?

Ramp Router is a beta LLM gateway at router.com that provides one API for calling, evaluating, routing, and tracking supported models across multiple providers.

How much does Router cost?

The routing layer is free through 2026. You pay the published token price of the model and service tier serving each request, and new users receive $26 in model credits.

Do I need to be a Ramp customer?

No. Router is separate from Ramp's finance products, and the beta is open to individual developers and teams in the U.S.

Does Router always choose a cheaper model?

No. Direct requests can remain on the selected model. Flex routing changes only the service tier, while benchmark aliases and opt-in Switchyard can route among models under their specific rules.

Which API formats does Router support?

It supports OpenAI Responses, OpenAI Chat Completions, and Anthropic Messages formats. Some features that cannot be translated faithfully are rejected.

Does Router store prompts and responses?

Yes by default. Router records inputs, outputs, and tool calls for one year. You can disable future content recording after the setting propagates, but that does not delete existing archives.

Does Router support zero data retention?

Router offers U.S.-hosted model options associated with zero-data-retention provider policies, but Router's own content recording is a separate account setting that is enabled by default.

Can I use my own provider API keys?

Yes for OpenAI, Fireworks, and xAI. BYOK calls are billed directly by the provider; Anthropic and Google Vertex AI currently use Router's shared credentials.

What is Shadow Mode?

Shadow Mode mirrors sampled Responses API requests to one to three candidate models in the background. Only the primary answer reaches the app, while Router records results for comparison.

What happens when a model provider fails?

Router can retry another service tier, credential source, or configured fallback model before output begins. Teams should test exact fallback behavior and disable unwanted shared-key billing.

Bottom line

Our Router verdict

Ramp Router is an unusually well-documented new gateway with practical cost levers: free routing, transparent model rates, Flex-tier selection, real-traffic shadowing, coding-agent escalation, and detailed request economics. It is most compelling for U.S. teams already spending enough on inference to justify continuous routing evaluation. The beta should not be dropped blindly into sensitive production traffic: turn off content recording if the one-year archive is unacceptable, control shared-key fallbacks, pin a model for the first migration, and prove quality and latency before allowing strategies to move traffic automatically.

Visit Router website ↗
The Rundown University

AI training for the future of work.

Get access to all our AI courses, hundreds of real-world AI use cases, live expert-led workshops, an exclusive network of AI early adopters, and more.

AI Courses

Get unlimited access to all of our current & upcoming industry-specific AI courses for the duration of your subscription.

Daily Guides

To keep up with the rapid pace of AI, our team publishes AI implementation guides daily. Our library contains 300+ practical use cases to automate real-world work.

Workshops

Join weekly, live, interactive sessions with industry leaders who are at the forefront of AI for hands-on implementation guidance and exclusive insights.

Community

Network with an exclusive community of AI-first professionals who are working smarter with AI. Learn how early adopters are using AI in their work and businesses.