AI model router · live

TRIAGE

Only critical work sees the surgeon.

Your most expensive model is a surgeon. You don't page a surgeon for a bandage. Triage reads every request before it leaves, gives it a severity tag, and sends it to the cheapest model on your ladder that can treat it.

One key for every model. Two settings in the client you already use. Every reply comes back with its tag, the model that answered and what it saved you.

$TRIAGE · Solana · contract address at launch
The request
Surgeon budget left100%
—
—
0.00
Not a mock-up. This panel imports lib/triage.js, the same file the server runs on every real request.
How it triages

Four tags, four models, one ladder.

Every request gets a severity score from plain, readable rules: how long the ask is, whether it needs judgement or just typing, code and stack traces, tools to drive, how deep the conversation is. The score picks the tag; the tag picks the model. You choose the four models.

Three rules

What makes it different.

01 · FROM REQUEST ONE

It saves from the first call.

Triage doesn't wait for your top model's budget to run low before it starts routing. A rename never sees the surgeon, even at 9am. Want everything on the surgeon while budget is plentiful? Flip to generous mode.

02 · ONLY UP

A conversation never gets demoted.

Once a conversation reaches a model, it stays there or moves up. A follow-up "thanks, now rename it" doesn't drop a hard debugging session onto a model that never saw the start, and the prompt cache stays warm.

03 · RECEIPTS

Every reply carries its chart.

Each response has headers with the tag, the score, the model and the reason. The app shows what each request cost and what it would have cost on the surgeon. The live board below shows the same for everyone.

Live board

Every request triaged, as it happens.

Tags, models, sizes and prices only. Never the prompt.

0
requests triaged
$0.00
saved vs. all-surgeon
$0.00
actually spent
tag mix
whentagmodeltokenssaved
No requests yet. The first one lands here the moment it's routed.
Setup

Two settings. Nothing else changes.

Make a key in the app, point your client's base URL at Triage, and send any model name. It's ignored: picking the model is the job. Streaming, tools and images work through both formats.

# OpenAI SDKs, Cursor, anything OpenAI-compatible
OPENAI_BASE_URL=https://triage.example/v1
OPENAI_API_KEY=tri_your_key

# Claude Code and Anthropic SDKs
ANTHROPIC_BASE_URL=https://triage.example
ANTHROPIC_API_KEY=tri_your_key
$TRIAGE

The token behind Triage.

$TRIAGE is Triage's token on Solana. Triage forwards your requests with your own OpenRouter key, so using it costs exactly what OpenRouter charges for the model that answered, and nothing more.

Credit comes from burning.

The house key pays for requests from workspaces that don't bring their own OpenRouter key. That credit is bought by burning $TRIAGE: you send tokens to the dead address from the app, your workspace tag rides along in the transaction, and once it's confirmed the workspace is credited at the token's market price. Bring your own key and Triage costs nothing; burning is only for house credit.

Questions

Fair questions.

Which models can be on my ladder?

Anything on OpenRouter's list with a fixed price: Claude, GPT, Gemini, DeepSeek, Qwen, Llama, Grok and the rest. The app shows the list with prices, read straight from OpenRouter, with the time it was read. Leave a rung empty and requests for it fall through to the next one.

How does it decide how hard a request is?

By rules, not by paying another model to read your prompt first. Words that signal judgement (debug, design, migrate, race condition, trade-offs) push the score up; words that signal typing (rename, summarise, translate, format) pull it down on short asks. Long asks, big context, code blocks, stack traces, tools and deep-reasoning requests add points. Every rule that fired is listed on the chart, both in the demo above and in the app.

What happens when the surgeon's budget runs low?

You set a daily dollar budget for the top rung. As it's spent, the bar to see the surgeon rises, so only the most critical work keeps going there. When it's gone, critical work goes to the doctor until midnight UTC.

Does a conversation stay on one model?

It can move up, never down. If a later turn is more critical it gets escalated; if it's less critical it stays where it was. That keeps context and prompt caching intact.

What happens to my OpenRouter key?

It's encrypted with AES-256-GCM before it's written, used only to forward your own requests, and the app only ever shows its first and last characters. Remove it and it's gone.

Does burning cost more than the credit is worth?

No markup: burned tokens are credited at the token's market price when the burn is verified, and house requests are charged at OpenRouter's listed price for the model that answered. The burn itself costs a normal transaction fee.

Why not just use a cheaper model?

Because some requests really are critical, and a cheap model on a hard bug costs more in retries than the surgeon would have. Triage keeps the surgeon for those and stops paying surgeon rates for everything else.