Case study 04 / ChatGPT app and Codex plugin
Tierwise: which model each OpenAI call needs, before the old ones retire.
OpenAI retires models on a schedule, and most codebases call a bigger model than the job needs. Tierwise scans a codebase's OpenAI calls, works out the model, tier, and cache layout each call site needs, and opens a pull request once a person approves. It runs as a ChatGPT app and as a Codex plugin, and Jev, TypeSafe's calibrated judgment model, makes every judgment call inside rules that code enforces.
Repository available on request. Architected and shipped by me, directing a team of coding agents.
- 95.2%right tool from a plain-language request, against 73.8% for a keyword router
- 0wrong tools executed; every miss was a clarifying question
- 92.3%of arguments extracted, and it never guessed where it should ask
- 109 msmedian Jev decision (243 ms at p95)
- 78Jev decisions to analyze the demo codebase, for $0.0017
- 120tests, run with the evals and a secret scan before every deploy
The problem
Every model has a shutdown date, and every call site has a cheaper way to run that nobody has time to find.
A product with a few dozen OpenAI calls has no easy way to answer simple questions. Which calls are on a model that retires next month, and what should replace each one? Which could run on a smaller model without anyone noticing? Which could wait for Batch or Flex? Is the prompt laid out so the cache ever hits? The suggested replacement for a retiring model is often the expensive one: on the demo codebase, the suggested replacement for one call site would have cost hundreds of times more than the model it actually needed. Guessing wrong in the other direction costs quality, so I wanted a tool that makes the call site by site and shows its reasoning.
What made it hard, and how I solved it
Jev answers narrow, typed questions. Thresholds, budgets, and safety live in code, reviewed and tested.
- 01
Knowing which model is enough
A call site that summarizes tickets and one that writes customer replies need different models, and the difference is judgment, not a lookup.
One Jev call per call site answers typed questions: the task, its complexity, whether each of GPT-6 Luna, Sol, and Astra can handle it, whether it is user-facing, whether it needs structured output. Code picks the cheapest model whose probability of being enough clears the bar (0.7, or 0.8 when a person reads the answer).
The cheapest model that is enough, call site by call site.
- 02
Everything else that sets the bill
Model choice is only part of it. Sync, Batch, or Flex, and whether the prompt is laid out to hit the cache, move the cost as much.
Jev chooses the execution tier and the cache layout, and code applies them only when policy allows: a nightly digest can wait for Batch, a user-facing reply cannot, and a static block moves to the front only when it is long and stable enough to cache.
Batch where it can wait, cache where it is stable.
- 03
Picking the right tool from a vague question
People ask in their own words. A keyword router sends “any deadlines I should worry about?” to the wrong place.
Jev picks the tool with a confidence. Code runs it at 0.6 or above, shows the top two when it is close, asks one clarifying question, or declines. Arguments come from closed sets and stay unset below 0.5, so Tierwise asks instead of guessing.
Ask, don't guess.
- 04
Spending money and writing code safely
Replaying real requests costs money, and opening a pull request changes someone's code.
Anything that spends or writes needs a signed approval: scoped to one action, expiring, single-use, and capped in dollars, checked in constant time. Jev can lower a cap or demand a separate approval for a riskier change, but it can never waive the receipt.
Jev can tighten a cap, never waive one.
- 05
Keys and prompt injection
People paste API keys, and scanned code or reports can carry instructions.
A key-shaped string ends the request in code before any model sees it. Jev screens the input for injection and credential requests and the output for leaks, and only redacted state is ever sent to it.
A key never reaches a model.
- 06
When the judge is down
A tool that stops working whenever its decision model does is not a tool.
Every decision point has a defined fallback: keep the current model, require strict approval, or show a menu instead of guessing. Tests cover each confidence branch, including Jev being unavailable.
A defined fallback at every decision.
How it works
Two surfaces share one decision layer. The Codex skill works in your repository; the ChatGPT app reads the report it publishes.

One request in ChatGPT
- Code. Screens for credentials. A key-shaped string ends the request; nothing is sent anywhere.
- Jev. In parallel: an input guard for injection and credential requests, and tool selection with the report as context.
- Code. Runs the tool at 0.6 or above, fans out to the top two, asks one clarifying question, or declines.
- Jev. Extracts arguments from closed sets when the tool needs a call site or a model. Below 0.5 they stay unset.
- Code. Runs the tool on the report.
- Code, then Jev. Screens the output: redaction, then a leak check.
- Result. Text for ChatGPT, structured content, and the decision trace the widget draws.
Measured on Cloud Run: about 0.7 seconds end to end with four live Jev decisions.
The nine calls Jev makes
- Which tool a request needs, and its arguments
- What each call site's workload is
- Which model is enough: GPT-6 Luna, Sol, or Astra
- Sync, Batch, or Flex
- How to lay out the prompt for the cache
- Which requests to replay, and how the outputs grade
- How risky each change is to approve
- How to split the pull request
- Whether the input or output is unsafe
Jev never executes anything and can never loosen a safety rule. It sees only secret-redacted state.


Evals, and a gate that blocks deploys
The same evals run before every deploy. They blocked the first one, on a type error in the eval runner.
| Tool selection | 95.2% on 42 paraphrased requests, against 73.8% for a keyword router; 0 wrong tools executed; every out-of-scope request declined |
|---|---|
| Arguments | 92.3% of 26 closed-set arguments; the two misses were left unset, so Tierwise asked |
| Calibration | expected calibration error 0.074 on the choices it executed |
| Jev overhead | 109 ms at p50 and 243 ms at p95 per routing call; about $0.05 per thousand routed requests |
| Whole analysis | 78 Jev decisions for $0.0017 on the demo codebase |
| Projected cost | about 75% lower on the demo workload, from about $4.4k to $1.1k a month (projected on a synthetic demo workload) |

What didn't work
My first plan was to route every single request across all three models. On labeled examples it put only about 80% of requests on a model that was good enough, and Jev's “the small model is enough” signal barely separated the cases (AUC 0.59). So Tierwise plans per call site, routes per request only inside the band that analysis finds, with verify-then-escalate, and treats every saving as projected until a replay on real outputs proves the cheaper model is no worse.
Results
Live on Cloud Run, with the Codex skill opening a draft pull request under a signed receipt.
- End to end in Codex. On the demo codebase the skill opened a draft pull request under a signed receipt. The same receipt was rejected when reused, and Jev required a separate approval for the riskier group of changes.
- Fast on the live service. A question with four live Jev decisions answers in about 0.7 seconds; the savings report loads warm in about 160 ms.
- Deployed carefully. Cloud Run that scales to zero, OAuth 2.1 with PKCE, a dedicated service account, secrets in Secret Manager, and reports in private storage that expires after 30 days.
- A gate on every deploy. Typecheck, 120 tests, the evals, and a secret scan must pass first. The tests cover every decision point across its confidence branches, the OAuth flows end to end, and approvals.
- TypeScript
- MCP
- ChatGPT app
- Codex plugin and skill
- OpenAI API
- TypeSafe Jev
- OAuth 2.1
- Cloud Run
- Vitest