Darwin grades each request by rules, not with another model: words like plan, migrate or refactor push it up, words like rename, summarise or format push it down, and attached tools add a little.
Darwin says it grades requests and routes them across a model ladder, moving easier requests to cheaper models as the top model's budget declines, while keeping a conversation on one model unless it fails or runs out.
Testing was inconclusive: no request could be sent without workspace credit, so difficulty grading, budget transitions and conversation stickiness were not observed.
To settle it: Configure a ladder and budget in the app, send requests with varied wording and tools, and inspect the x-darwin-model and x-darwin-rung response headers.
Wallet test, inconclusive: The app and docs describe the claimed grading, budget-based routing, and keeping a conversation on one model unless it fails or runs out. Those descriptions are not runtime verification; no request was routed during this test.
Skeptic: plausible
Rule-based request grading, budget thresholds and model fallback are feasible routing behaviors, and the docs describe how they are intended to work. The page does not independently establish routing quality or cost savings.