Questions
Fair questions.
Which models can be on my ladder?
Anything on OpenRouter's list with a fixed price: Claude, GPT, Gemini, DeepSeek, Qwen, Llama, Grok and the rest. The app shows the list with prices, read straight from OpenRouter, with the time it was read. Leave a rung empty and requests for it fall through to the next one.
How does it decide how hard a request is?
By rules, not by paying another model to read your prompt first. Words that signal judgement (debug, design, migrate, race condition, trade-offs) push the score up; words that signal typing (rename, summarise, translate, format) pull it down on short asks. Long asks, big context, code blocks, stack traces, tools and deep-reasoning requests add points. Every rule that fired is listed on the chart, both in the demo above and in the app.
What happens when the surgeon's budget runs low?
You set a daily dollar budget for the top rung. As it's spent, the bar to see the surgeon rises, so only the most critical work keeps going there. When it's gone, critical work goes to the doctor until midnight UTC.
Does a conversation stay on one model?
It can move up, never down. If a later turn is more critical it gets escalated; if it's less critical it stays where it was. That keeps context and prompt caching intact.
What happens to my OpenRouter key?
It's encrypted with AES-256-GCM before it's written, used only to forward your own requests, and the app only ever shows its first and last characters. Remove it and it's gone.
Does burning cost more than the credit is worth?
No markup: burned tokens are credited at the token's market price when the burn is verified, and house requests are charged at OpenRouter's listed price for the model that answered. The burn itself costs a normal transaction fee.
Why not just use a cheaper model?
Because some requests really are critical, and a cheap model on a hard bug costs more in retries than the surgeon would have. Triage keeps the surgeon for those and stops paying surgeon rates for everything else.