Route, monitor, and optimize LLM requests across providers. Built-in billing, quota enforcement, key management, and a Mission Control dashboard for your whole team.
Moat sits between your app and your LLM providers — a single endpoint that handles auth, billing, quotas, analytics, and multi-provider routing.
One API endpoint, any model. Route to Google Gemini, Groq, Anthropic, OpenAI, or any OpenAI-compatible provider. Automatic failover between providers.
Every request logged with tokens, latency, cost, model, and user. Granular dashboards show spend by team, model, and time period.
BYOK (Bring Your Own Key) or use shared pools. Per-org API keys with rotation, scoped to specific models and rate limits.
Per-org request limits, token budgets, and rate limiting. Hard and soft limits with configurable overage policies. Never blow your budget.
Bundled ops dashboard with 23 panels: team roster, agent orchestration, deploy guardian, prompt management, alerts, and health monitoring.
Register AI agents alongside human operators. Brain vs Muscles routing sends expensive reasoning to big models and volume work to fast models.
Point your OpenAI SDK at Moat. Everything else stays the same.
Sign in with email and password. Moat provisions your org, generates an API key, and assigns a tier with quota limits.
Set base_url to Moat's proxy endpoint. Use any OpenAI-compatible client — Python, Node, curl.
Request any supported model. Moat routes to the right provider, logs usage, checks quotas, and returns the response.
Open Mission Control to see real-time usage, costs, latency, and team activity across every model and member.
Every other LLM proxy just routes requests. Moat gives your agents persistent memory across sessions — so they pick up exactly where they left off.
Register AI agent projects, track tasks and decisions, take state snapshots, and auto-generate resume prompts. Your agents never lose context again.
Three-tier memory system: working memory (curated context injected into every call), episodic logs (automatic extraction), and long-term archive with decay.
Memory extraction happens via heuristics after every response — under 15ms, no extra LLM call. Curation uses Gemini Flash Lite for ~$0.001 per run.
CRON-based autonomous automation: schedule memory curation, context decay, cross-agent messages, webhooks, and LLM prompts — all running on a 60-second tick loop.
Every plan includes the Moat proxy and Mission Control dashboard. Upgrade when you need more calls, users, or agents.