Case study 03 / Open source
Release Guard: letting a coding agent run my App Store releases, safely.
I ship DeepChamp to the App Store often, and the release checklist lived in my head and in one-off scripts. I built Release Guard so a coding agent could answer “is this release ready?” and take a build to App Review without my trusting it to be careful: six App Store Connect tools, a 13-check preflight, and a submit that stays a dry run unless seven conditions hold. It runs as an MCP server and as a Codex plugin.

- 6tools, each answering a question a release owner actually asks
- 13preflight checks across App Store Connect, the repo, and a lint policy
- 7conditions that must all hold before a single write request can leave
- 112unit and MCP contract tests against an in-process fake App Store Connect
- 224scored runs as real Codex sessions: 28 cases, 8 configurations
- 0safety violations in every run, including a prompt injection
The problem
Let an agent in Codex answer “is this release ready?” and take it to App Review, without trusting the model to be careful.
An iOS release depends on about a dozen App Store Connect checks that usually live in someone's head: a build still processing, export compliance unanswered, one locale with an empty What's New, notes with a claim App Review rejects, a first-time subscription not attached, another version stuck in review, a build number behind the last upload, an archive from a commit that never reached main. I had been shipping DeepChamp releases with one-off scripts that mixed reads with writes. Coding agents can run these steps now, but an agent with a raw App Store Connect key is one bad guess from shipping an unreviewed build.
What made it hard
Six things a useful, safe release tool has to get right.
- 01
Questions, not endpoints
App Store Connect is dozens of JSON:API resources. A release owner asks six questions.
- 02
A verdict you can trust
About a dozen release checks live in people's heads, and a generated answer can differ run to run.
- 03
Writes that cannot happen by accident
Handing an agent an App Store Connect key without guardrails is how an unreviewed build ships.
- 04
Timeouts and a rate-limited API
Codex gives a tool 60 seconds; Apple rate-limits and pages its results.
- 05
Credentials and data in the model's context
A leaked token or a review note full of contact details should never reach a prompt or a log.
- 06
Getting every model to use the tool
A server that works for one model can be ignored by another, which answers from memory instead.
How I solved it
One decision per hard part, enforced in code below the model.
- 01
Questions, not endpoints
Six tools map to those questions: did it process, where is it, can I submit, is this text OK, did Apple's notifications arrive, submit it. The agent picks by intent; the logic lives in code.
The model chooses a question, not a URL.
- 02
A verdict you can trust
One preflight runs 13 checks across App Store Connect, the local repo, and a lint policy. Rules compute the verdict (blocked, ready with warnings, ready), and every check carries a concrete fix.
Two runs always agree.
- 03
Writes that cannot happen by accident
The transport refuses every non-GET request unless it carries a write permit, which only the submit path can mint after seven checks. Submit is dry-run by default and annotated destructive, so Codex asks a human.
A prompt-injected confirm=true is not enough.
- 04
Timeouts and a rate-limited API
A 45-second deadline inside the 60-second tool timeout, retries with full-jitter backoff, Retry-After honored or turned into an actionable error, and pagination followed on Apple's host only.
No hung tools, no silent partial answers.
- 05
Credentials and data in the model's context
Every token is ES256, short-lived, and scoped to the one request it signs. Requests ask only for the fields they need, review notes become a length plus a lint result, and logs are structured and redacted.
A token for one request gets 403 on another.
- 06
Getting every model to use the tool
An eval harness runs natural-language requests as real Codex sessions and scores the first tool call, its arguments, and safety. Packaging the server as a Codex plugin with a skill closed the gap.
28/28 on every model tested.
How it works
The model only sees typed tools. Credentials, retries, and write safety live below it, in one Python process.

Six tools
check_build_statusread-onlyDid the build process, and is export compliance answered?
check_version_stateread-onlyWhere is this version, and is anything else in review?
preflight_submissionread-onlyCan I submit? 13 checks, a rule-computed verdict, and a fix for each.
lint_release_notesread-onlyIs this What's New text acceptable to App Review?
reconcile_server_notificationsread-onlyDid Apple's server notifications arrive? Report only.
submit_for_reviewdestructiveSubmit, only after every gate passes. Dry run by default.
The 13-check preflight
- App Store version exists
- Version is editable
- Build processed and valid
- Build attached to version
- Export compliance answered
- What's New present in every locale
- What's New free of banned claims
- App Review notes present
- Reviewer contact and demo account
- In-app purchases attached
- No other version in review
- Build number ahead of the last upload
- Release commit is on main
Seven gates before any write
- dry_run is false (it defaults to true)
- confirm is true (it defaults to false)
- the server was started with writes allowed
- not running under a test runner
- a live backend (the demo backend never writes)
- the preflight is not blocked
- a write permit on every request
The five read tools are annotated read-only and the submit tool destructive, so Codex runs the reads freely and always asks a human before submit, even for its dry run.
Measured in real Codex sessions
28 natural-language cases, each run as a real codex exec session: 224 scored runs across eight configurations, with 0 safety violations.
First tool call right, out of 28
Each case is a fresh codex exec session with a clean home and connectors off. The scorer grades the first Release Guard call: the right tool with the right arguments.
Packaging the server as a Codex plugin with a skill got every model to use the right tool. As a plain MCP server, two models answered release-note questions from memory instead of calling the lint tool, and a more directive tool description alone did not fix it. The skill routes those questions to the tool, so routing no longer depends on the model: 28/28 on all three.

What's new
The parts I'd reuse in the next agent tool I build.
- Safety below the modelWrites are impossible without a permit that only the gated submit path can mint. Safety is scored in every eval run, including a prompt injection hidden in release notes.
- Eval-driven tool designTool descriptions and server instructions were rewritten from the misses in real Codex sessions, then re-measured on three models.
- A fake App Store ConnectOne in-process fake with Apple's JSON:API shapes, pagination, and 429, 5xx, and 403 fault injection powers tests, evals, and demos with zero Apple calls.
- Least data into contextField-limited requests, review notes reduced to a length and a lint result, contact details reduced to booleans.
- Policy as dataLint rules live in TOML, each citing the App Review guideline it protects, so a team can extend them.
- Report-only reconcileReplaying Apple's notifications is a side effect, so the tool counts and explains, and replay stays with the team's own tooling.
Results
It ran live, read-only, against a production app, and every preflight result was correct.
- Correct on production.Every tool ran against the real app with its real key, and the preflight's one failing check was the right answer.
- Zero writes. The server logs of every live session show no write requests, and the dry-run submit planned its writes and executed none.
- Scoped tokens hold. A token scoped to one request was rejected with 403 on a different request.
- Installs as a plugin. Installed from a local marketplace and run as a real Codex session against the live app: reads only.
- Python
- MCP Python SDK
- Codex plugin and skill
- App Store Connect API
- App Store Server API
- ES256 JWT
- Pydantic
- pytest
- gitleaks