darwin
Keep working when your model's limit runs low.
One API key for every model. Darwin keeps your best model for hard work, moves easy requests to cheaper ones as its limit runs down, and runs its own GPU model when everything else is spent.
Your best model keeps the hard work
Darwin grades every request before it goes out and sends it down your list of models: the hardest to the top, the easy ones to cheaper models as the top one's limit runs low. A conversation stays on one model unless that model runs out.

Your best model
Plans, hard bugs, long refactors. Darwin keeps these here until its limit is nearly gone.
Claude Fable 5.1, GPT-6 Astra

A mid-size model
Most coding work and tool calls once the top model's budget is past half.
Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 3.1 Pro

A small, fast model
Renames, summaries, lookups and other easy requests.
Claude Haiku 4.5, Gemini 3.8 Flash, DeepSeek V4.1 Flash

Darwin's GPU
Qwen3 30B on Darwin's own hardware. No key needed, and it catches everything when the rest are out.
Paid with burned $DARWIN
A model of its own
Darwin runs Qwen3 30B on its own GPUs. It answers with no provider key at all, and when every model on your list is out, the work goes there instead of stopping. Every reply names the model that wrote it.
Access is paid for by burning $DARWIN from your wallet. Each burn turns into credit for the API and the MCP server, and credit buys upgrades: Pro moves an agent to gpt-oss-120b with a 128K context, and warm hours keep a GPU running so it never waits for a cold start.
Setup is one setting
Connect a wallet or make a key in the app, then point the base URL in Claude Code, the OpenAI or Anthropic SDK, or any client that takes one at Darwin. Add your own provider keys whenever you want Darwin to route between your models.
Read the setup guideWhich models does it work with?
Anything you can reach with an Anthropic, OpenAI, Google or OpenRouter key. OpenRouter covers the rest: Grok, DeepSeek, Mistral, Qwen, Llama, Kimi, GLM and hundreds more. You put models in order in the dashboard and Darwin works down the list.
How does it decide how hard a request is?
By rules, not by another model reading your prompt: the length of the last message, words like plan, migrate or refactor against words like rename, summarise or format, and whether tools are attached.
How do I pay for Darwin?
By burning $DARWIN on Solana (5VnbrKp28Qs9CAH6PyZdxBNvLxZcbeX3YBsozgWnpump). Burning it from your wallet on the Credits page adds credit to your workspace, and the API and the MCP server only work while there is credit. Darwin's GPU model draws on that credit; routing to your own provider keys doesn't use it up.
Which limits does it watch?
The daily budget you set for your top model, counted from the usage each provider reports, and rate-limit errors. When the top model returns one, Darwin moves down the list until the retry time passes. It does not change the limits of a ChatGPT, Claude.ai or Gemini subscription.
Does a conversation stay on one model?
Yes. Darwin keeps a conversation on the model it started with and only moves it down when that model is out, because thinking blocks and prompt caches belong to one model.
What happens to my keys?
Provider keys are stored encrypted and only used to call that provider for you. Darwin keeps the model, token counts, cost and timing of each request, never the prompt or the reply.
Why not just lower the effort setting?
Try that first. A lower reasoning effort on one model often costs less and keeps a single cache. Darwin is for when that still runs out before the day does.
