GATEWAY // PRICING & LIMITS

Limits that match your workload.

Start without a subscription and move to production controls when you need higher ceilings.

Loop savings calculator

Estimate exposure from runaway sessions

$35/mo

1 session25 sessions

Estimate uses the supplied $35 per prevented incident assumption multiplied by sessions per week. This is a planning estimate, not a guarantee or measured savings.

Developer

Free

$0/mo

For development and initial workloads.

  • Up to $15/mo monitored spend
  • 3× sliding-window loop detection
  • Hourly and daily budget caps
  • OpenAI-compatible proxy endpoint

Production

Pro

$29/mo

For production agents that need higher budget ceilings.

  • Unlimited monitored spend
  • $50/hr and $200/day burst ceilings
  • Emergency Slack and Discord webhooks
  • Priority European routing
Spend ceilings are enforced by the gateway using account configuration. Review actual limits and usage in your dashboard before routing production traffic.

Implementation details

Technical FAQs

What latency does the gateway add?

The proxy is designed for under 35ms of average overhead. Actual latency depends on network distance, provider response time, and deployment conditions.

Are prompts retained?

Prompt bodies are processed for loop detection in memory and are not written to disk. Request metadata and usage logs do not include prompt text.

How are upstream API keys protected?

Upstream credentials are encrypted with AES-256 (Fernet) before database insertion. Protected key material is hashed for lookup; raw upstream keys are not logged.

Which error codes can the gateway return?

HTTP 400 indicates an infinite loop was stopped, HTTP 429 indicates a budget ceiling was reached, and HTTP 502 indicates an upstream provider error.

Which agent frameworks are compatible?

Any client or framework that can send OpenAI-compatible requests to a custom base URL can use the gateway, including CrewAI, LangChain, LangGraph, AutoGen, and direct REST clients.