What is rightmodeler?
rightmodeler is an open-source tool that checks whether cheaper AI models can handle each step of your agent. It replays the traces you already recorded through candidate models, has a judge from another vendor grade every answer, and reports cost savings, quality, sample size, and abstentions before any change reaches production.
Top Features:
- Trace autodetection: reads 12 trace formats, including LangSmith, Langfuse, and Braintrust exports.
- Step replay: runs each recorded step on cheaper models against shipped outputs.
- Pull requests: approved swaps become draft pull requests changing only model names.
Use Cases:
- Cost cutting: find agent steps where a smaller model gives the same answer.
- Migrations: test a new model on real traces before switching any production step.
- Reviews: attach quality scores and cost deltas to every proposed model change.
Who Can Use rightmodeler?
- AI engineers: right-size the models behind multi-step agents without running manual evals.
- Tech leads: keep LLM spending down while holding a clear quality floor.
- Platform teams: run model audits in CI beside your existing observability tools.
Pricing
- CLI (free): the MIT-licensed npm package runs with your own provider key.
- Self-hosted agent (free): clone the repo and run scheduled audits that open pull requests.
- Hosted agent (waitlist): a managed version is coming, and pricing is not yet published.
Pros and Cons
Pros:
- Evidence first: each recommendation shows sample size, quality score, and cost change.
- Honest abstention: no pull request opens when a candidate misses your quality floor.
- No gateway: it reads exported traces and never sits in your request path.
Cons:
- Setup needs: requires Node.js 24, a Git repo, and traces with token usage.
- Provider costs: replays call your model provider, so big audits add API spend.
- Early stage: the hosted agent and Crucible analytics suite are still in development.
FAQs:
1) Is it free?
Yes, it is MIT licensed, and you only pay your provider for replays.
2) Which traces does it read?
Claude Code, Codex, LangSmith, Langfuse, Braintrust, Phoenix, Helicone, and other common formats.
3) Does it sit in my request path?
No, it works from exported traces and never acts as a runtime gateway.
4) Does it change code automatically?
It opens a draft pull request editing only model names, and humans merge it.
5) What if no cheaper model qualifies?
It abstains and reports the missing evidence instead of forcing a risky swap.