All work

Case study 04 / ChatGPT app and Codex plugin

Tierwise: which model each OpenAI call needs, before the old ones retire.

OpenAI retires models on a schedule, and most codebases call a bigger model than the job needs. Tierwise scans a codebase's OpenAI calls, works out the model, tier, and cache layout each call site needs, and opens a pull request once a person approves. It runs as a ChatGPT app and as a Codex plugin, and Jev, TypeSafe's calibrated judgment model, makes every judgment call inside rules that code enforces.

Repository available on request. Architected and shipped by me, directing a team of coding agents.

A 22-second demo, recorded in a local host that renders the same widget ChatGPT does, not in ChatGPT. The workload is a synthetic demo codebase.
01 / The problem

The problem

Every model has a shutdown date, and every call site has a cheaper way to run that nobody has time to find.

A product with a few dozen OpenAI calls has no easy way to answer simple questions. Which calls are on a model that retires next month, and what should replace each one? Which could run on a smaller model without anyone noticing? Which could wait for Batch or Flex? Is the prompt laid out so the cache ever hits? The suggested replacement for a retiring model is often the expensive one: on the demo codebase, the suggested replacement for one call site would have cost hundreds of times more than the model it actually needed. Guessing wrong in the other direction costs quality, so I wanted a tool that makes the call site by site and shows its reasoning.

02 / What made it hard

What made it hard, and how I solved it

Jev answers narrow, typed questions. Thresholds, budgets, and safety live in code, reviewed and tested.

  1. 01

    Knowing which model is enough

    A call site that summarizes tickets and one that writes customer replies need different models, and the difference is judgment, not a lookup.

    One Jev call per call site answers typed questions: the task, its complexity, whether each of GPT-6 Luna, Sol, and Astra can handle it, whether it is user-facing, whether it needs structured output. Code picks the cheapest model whose probability of being enough clears the bar (0.7, or 0.8 when a person reads the answer).

    The cheapest model that is enough, call site by call site.

  2. 02

    Everything else that sets the bill

    Model choice is only part of it. Sync, Batch, or Flex, and whether the prompt is laid out to hit the cache, move the cost as much.

    Jev chooses the execution tier and the cache layout, and code applies them only when policy allows: a nightly digest can wait for Batch, a user-facing reply cannot, and a static block moves to the front only when it is long and stable enough to cache.

    Batch where it can wait, cache where it is stable.

  3. 03

    Picking the right tool from a vague question

    People ask in their own words. A keyword router sends “any deadlines I should worry about?” to the wrong place.

    Jev picks the tool with a confidence. Code runs it at 0.6 or above, shows the top two when it is close, asks one clarifying question, or declines. Arguments come from closed sets and stay unset below 0.5, so Tierwise asks instead of guessing.

    Ask, don't guess.

  4. 04

    Spending money and writing code safely

    Replaying real requests costs money, and opening a pull request changes someone's code.

    Anything that spends or writes needs a signed approval: scoped to one action, expiring, single-use, and capped in dollars, checked in constant time. Jev can lower a cap or demand a separate approval for a riskier change, but it can never waive the receipt.

    Jev can tighten a cap, never waive one.

  5. 05

    Keys and prompt injection

    People paste API keys, and scanned code or reports can carry instructions.

    A key-shaped string ends the request in code before any model sees it. Jev screens the input for injection and credential requests and the output for leaks, and only redacted state is ever sent to it.

    A key never reaches a model.

  6. 06

    When the judge is down

    A tool that stops working whenever its decision model does is not a tool.

    Every decision point has a defined fallback: keep the current model, require strict approval, or show a menu instead of guessing. Tests cover each confidence branch, including Jev being unavailable.

    A defined fallback at every decision.

03 / How it works

How it works

Two surfaces share one decision layer. The Codex skill works in your repository; the ChatGPT app reads the report it publishes.

Tierwise architecture. Left, the Codex skill in the repository: scan finds OpenAI calls, analyze asks Jev one fan-out question per call site, a human-only approve step issues a signed receipt, replay spends under a budget meter, and pr opens draft pull requests. Center, the Jev decision layer: typed questions become probabilities and code policy decides tool selection, arguments, workload, routing, tier, cache, replay, approvals, pull requests, and guards, with every answer in a decision ledger. Right, the ChatGPT app on Cloud Run: the MCP server, its read-only tools, the widget, and the platform with OAuth 2.1.
Jev decides, code stays in control. Every Jev answer is logged as a decision record: the question, the probabilities, the confidence, the rule that applied, and the action.

One request in ChatGPT

  1. Code. Screens for credentials. A key-shaped string ends the request; nothing is sent anywhere.
  2. Jev. In parallel: an input guard for injection and credential requests, and tool selection with the report as context.
  3. Code. Runs the tool at 0.6 or above, fans out to the top two, asks one clarifying question, or declines.
  4. Jev. Extracts arguments from closed sets when the tool needs a call site or a model. Below 0.5 they stay unset.
  5. Code. Runs the tool on the report.
  6. Code, then Jev. Screens the output: redaction, then a leak check.
  7. Result. Text for ChatGPT, structured content, and the decision trace the widget draws.

Measured on Cloud Run: about 0.7 seconds end to end with four live Jev decisions.

The nine calls Jev makes

  1. Which tool a request needs, and its arguments
  2. What each call site's workload is
  3. Which model is enough: GPT-6 Luna, Sol, or Astra
  4. Sync, Batch, or Flex
  5. How to lay out the prompt for the cache
  6. Which requests to replay, and how the outputs grade
  7. How risky each change is to approve
  8. How to split the pull request
  9. Whether the input or output is unsafe

Jev never executes anything and can never loosen a safety rule. It sees only secret-redacted state.

The decision trace panel for one call site, reply-draft: the workload classification with probabilities for each model, the routing decision to start on GPT-6 Luna and escalate to Astra, the execution tier, and the prompt cache decision, each with its confidence. Labelled sample and projected.
The decision trace. Every call is shown with its question, probabilities, confidence, and the rule that turned it into an action.
The savings dashboard for a synthetic demo codebase, labelled sample and projected: eight call sites, each with its current and planned model, tier, and cache, and three shutdown deadlines.
The plan, per call site. Illustrative data: a synthetic demo codebase, with every saving projected until it is replayed.
04 / Evals

Evals, and a gate that blocks deploys

The same evals run before every deploy. They blocked the first one, on a type error in the eval runner.

What was measured
Tool selection95.2% on 42 paraphrased requests, against 73.8% for a keyword router; 0 wrong tools executed; every out-of-scope request declined
Arguments92.3% of 26 closed-set arguments; the two misses were left unset, so Tierwise asked
Calibrationexpected calibration error 0.074 on the choices it executed
Jev overhead109 ms at p50 and 243 ms at p95 per routing call; about $0.05 per thousand routed requests
Whole analysis78 Jev decisions for $0.0017 on the demo codebase
Projected costabout 75% lower on the demo workload, from about $4.4k to $1.1k a month (projected on a synthetic demo workload)
The Tierwise eval report: tool selection by Jev against a keyword router in ChatGPT and Codex, the projected monthly cost of the demo workload under four strategies, routing sufficiency against cost across thresholds, routing outcomes by strategy, and the release gates, all passing.
The eval report. Release gates include tool accuracy, wrong-execute rate, decline recall, argument accuracy, plan sufficiency, and Jev cost per thousand requests. All figures are on a synthetic demo workload.
05 / What didn't work

What didn't work

My first plan was to route every single request across all three models. On labeled examples it put only about 80% of requests on a model that was good enough, and Jev's “the small model is enough” signal barely separated the cases (AUC 0.59). So Tierwise plans per call site, routes per request only inside the band that analysis finds, with verify-then-escalate, and treats every saving as projected until a replay on real outputs proves the cheaper model is no worse.

06 / Results

Results

Live on Cloud Run, with the Codex skill opening a draft pull request under a signed receipt.

  • End to end in Codex. On the demo codebase the skill opened a draft pull request under a signed receipt. The same receipt was rejected when reused, and Jev required a separate approval for the riskier group of changes.
  • Fast on the live service. A question with four live Jev decisions answers in about 0.7 seconds; the savings report loads warm in about 160 ms.
  • Deployed carefully. Cloud Run that scales to zero, OAuth 2.1 with PKCE, a dedicated service account, secrets in Secret Manager, and reports in private storage that expires after 30 days.
  • A gate on every deploy. Typecheck, 120 tests, the evals, and a secret scan must pass first. The tests cover every decision point across its confidence branches, the OAuth flows end to end, and approvals.
  • TypeScript
  • MCP
  • ChatGPT app
  • Codex plugin and skill
  • OpenAI API
  • TypeSafe Jev
  • OAuth 2.1
  • Cloud Run
  • Vitest