All work

Case study 03 / Open source

Release Guard: letting a coding agent run my App Store releases, safely.

I ship DeepChamp to the App Store often, and the release checklist lived in my head and in one-off scripts. I built Release Guard so a coding agent could answer “is this release ready?” and take a build to App Review without my trusting it to be careful: six App Store Connect tools, a 13-check preflight, and a submit that stays a dry run unless seven conditions hold. It runs as an MCP server and as a Codex plugin.

View the code on GitHubOpen source. Architected and shipped by me, directing a team of coding agents.
A Codex session with Release Guard installed as a plugin, on the demo backend: the preflight blocks a submission with three required fixes, and Codex holds the destructive submit tool for approval. Zero write requests.
01 / The problem

The problem

Let an agent in Codex answer “is this release ready?” and take it to App Review, without trusting the model to be careful.

An iOS release depends on about a dozen App Store Connect checks that usually live in someone's head: a build still processing, export compliance unanswered, one locale with an empty What's New, notes with a claim App Review rejects, a first-time subscription not attached, another version stuck in review, a build number behind the last upload, an archive from a commit that never reached main. I had been shipping DeepChamp releases with one-off scripts that mixed reads with writes. Coding agents can run these steps now, but an agent with a raw App Store Connect key is one bad guess from shipping an unreviewed build.

02 / What made it hard

What made it hard

Six things a useful, safe release tool has to get right.

  1. 01

    Questions, not endpoints

    App Store Connect is dozens of JSON:API resources. A release owner asks six questions.

  2. 02

    A verdict you can trust

    About a dozen release checks live in people's heads, and a generated answer can differ run to run.

  3. 03

    Writes that cannot happen by accident

    Handing an agent an App Store Connect key without guardrails is how an unreviewed build ships.

  4. 04

    Timeouts and a rate-limited API

    Codex gives a tool 60 seconds; Apple rate-limits and pages its results.

  5. 05

    Credentials and data in the model's context

    A leaked token or a review note full of contact details should never reach a prompt or a log.

  6. 06

    Getting every model to use the tool

    A server that works for one model can be ignored by another, which answers from memory instead.

03 / How I solved it

How I solved it

One decision per hard part, enforced in code below the model.

  1. 01

    Questions, not endpoints

    Six tools map to those questions: did it process, where is it, can I submit, is this text OK, did Apple's notifications arrive, submit it. The agent picks by intent; the logic lives in code.

    The model chooses a question, not a URL.

  2. 02

    A verdict you can trust

    One preflight runs 13 checks across App Store Connect, the local repo, and a lint policy. Rules compute the verdict (blocked, ready with warnings, ready), and every check carries a concrete fix.

    Two runs always agree.

  3. 03

    Writes that cannot happen by accident

    The transport refuses every non-GET request unless it carries a write permit, which only the submit path can mint after seven checks. Submit is dry-run by default and annotated destructive, so Codex asks a human.

    A prompt-injected confirm=true is not enough.

  4. 04

    Timeouts and a rate-limited API

    A 45-second deadline inside the 60-second tool timeout, retries with full-jitter backoff, Retry-After honored or turned into an actionable error, and pagination followed on Apple's host only.

    No hung tools, no silent partial answers.

  5. 05

    Credentials and data in the model's context

    Every token is ES256, short-lived, and scoped to the one request it signs. Requests ask only for the fields they need, review notes become a length plus a lint result, and logs are structured and redacted.

    A token for one request gets 403 on another.

  6. 06

    Getting every model to use the tool

    An eval harness runs natural-language requests as real Codex sessions and scores the first tool call, its arguments, and safety. Packaging the server as a Codex plugin with a skill closed the gap.

    28/28 on every model tested.

04 / How it works

How it works

The model only sees typed tools. Credentials, retries, and write safety live below it, in one Python process.

Release Guard architecture: an MCP host (Codex, the IDE extension, or ChatGPT) talks JSON-RPC over stdio to one Python process with typed tools, checks and verdicts, a lint policy, local repo checks, typed reads, a resilient transport, scoped auth, and redacted logs. Seven write gates sit beside it, and it calls the App Store Connect and App Store Server APIs over HTTPS with scoped tokens.
The architecture. Tools on top, rule-based checks and verdicts under them, and a transport that refuses writes without a permit. A fake App Store Connect powers the tests, the evals, and a demo mode, so CI never calls Apple.

Six tools

  • check_build_statusread-only

    Did the build process, and is export compliance answered?

  • check_version_stateread-only

    Where is this version, and is anything else in review?

  • preflight_submissionread-only

    Can I submit? 13 checks, a rule-computed verdict, and a fix for each.

  • lint_release_notesread-only

    Is this What's New text acceptable to App Review?

  • reconcile_server_notificationsread-only

    Did Apple's server notifications arrive? Report only.

  • submit_for_reviewdestructive

    Submit, only after every gate passes. Dry run by default.

The 13-check preflight

  1. App Store version exists
  2. Version is editable
  3. Build processed and valid
  4. Build attached to version
  5. Export compliance answered
  6. What's New present in every locale
  7. What's New free of banned claims
  8. App Review notes present
  9. Reviewer contact and demo account
  10. In-app purchases attached
  11. No other version in review
  12. Build number ahead of the last upload
  13. Release commit is on main

Seven gates before any write

  1. dry_run is false (it defaults to true)
  2. confirm is true (it defaults to false)
  3. the server was started with writes allowed
  4. not running under a test runner
  5. a live backend (the demo backend never writes)
  6. the preflight is not blocked
  7. a write permit on every request

The five read tools are annotated read-only and the submit tool destructive, so Codex runs the reads freely and always asks a human before submit, even for its dry run.

05 / Evals

Measured in real Codex sessions

28 natural-language cases, each run as a real codex exec session: 224 scored runs across eight configurations, with 0 safety violations.

First tool call right, out of 28

Each case is a fresh codex exec session with a clean home and connectors off. The scorer grades the first Release Guard call: the right tool with the right arguments.

07142128

First-draft tool descriptions, plain MCP server

17/28 gpt-6-solgpt-6-luna: not rungpt-5.5: not run

Improved tool descriptions, plain MCP server

23/28 gpt-6-sol20/28 gpt-6-luna28/28 gpt-5.5

Packaged as a Codex plugin, with the skill

28/28 gpt-6-sol28/28 gpt-6-luna28/28 gpt-5.5
The lesson

Packaging the server as a Codex plugin with a skill got every model to use the right tool. As a plain MCP server, two models answered release-note questions from memory instead of calling the lint tool, and a more directive tool description alone did not fix it. The skill routes those questions to the tool, so routing no longer depends on the model: 28/28 on all three.

Tool-selection eval results: 100% first-call accuracy as a Codex plugin, 82% as a bare MCP entry, 61% for the first draft of the tool descriptions, and 0 safety violations, with pass rates by category and a table by model and packaging.
Pass rates by category. The misses changed the design: instructions written as a routing table fixed preflight and submit, and the plugin skill fixed release-note linting.
06 / What's new

What's new

The parts I'd reuse in the next agent tool I build.

  • Safety below the modelWrites are impossible without a permit that only the gated submit path can mint. Safety is scored in every eval run, including a prompt injection hidden in release notes.
  • Eval-driven tool designTool descriptions and server instructions were rewritten from the misses in real Codex sessions, then re-measured on three models.
  • A fake App Store ConnectOne in-process fake with Apple's JSON:API shapes, pagination, and 429, 5xx, and 403 fault injection powers tests, evals, and demos with zero Apple calls.
  • Least data into contextField-limited requests, review notes reduced to a length and a lint result, contact details reduced to booleans.
  • Policy as dataLint rules live in TOML, each citing the App Review guideline it protects, so a team can extend them.
  • Report-only reconcileReplaying Apple's notifications is a side effect, so the tool counts and explains, and replay stays with the team's own tooling.
07 / Results

Results

It ran live, read-only, against a production app, and every preflight result was correct.

  • Correct on production.Every tool ran against the real app with its real key, and the preflight's one failing check was the right answer.
  • Zero writes. The server logs of every live session show no write requests, and the dry-run submit planned its writes and executed none.
  • Scoped tokens hold. A token scoped to one request was rejected with 403 on a different request.
  • Installs as a plugin. Installed from a local marketplace and run as a real Codex session against the live app: reads only.
  • Python
  • MCP Python SDK
  • Codex plugin and skill
  • App Store Connect API
  • App Store Server API
  • ES256 JWT
  • Pydantic
  • pytest
  • gitleaks