All AI engineering

Case study 01/Live on the App Store

DeepChamp: a production AI research platform, end to end.

A native iOS sports research and analytics app on 168 cloud services. Multi-model AI with evidence gates, server-owned real-time voice, a signed entitlement pipeline, and a warehouse that turns product signals into gated decisions, with Jev making calibrated calls in real time.

Product film excerpt. Muted.
01 / The problem

The problem

People ask about games that are happening now and expect an answer they can trust, by text or by voice. Doing that well is six hard problems at once, and each one can quietly break the others.

What made it hard

  1. P1

    Noisy multi-source data, real-time decisions

    Scores, rosters, injuries, weather, and news arrive from many providers at different speeds, with gaps and conflicts, while the game is still live.

  2. P2

    Trust and calibration of AI answers

    A fluent answer is not a correct one. Answers must rest on current evidence and be honest about how sure they are.

  3. P3

    Real-time voice

    Voice must answer in conversational time, with tools and live data, without putting a provider key on the phone or letting sessions run unmetered.

  4. P4

    Monetizing within App Store rules

    Subscriptions must run through StoreKit and meet App Store guidelines, while the paywall still needs honest experimentation.

  5. P5

    Growth measurement through ATT and SKAN blindness

    With App Tracking Transparency most installs cannot be tied to a campaign, and SKAdNetwork reports arrive late, coarse, and anonymous.

  6. P6

    Entitlements across App Store, web, and identity

    People subscribe on iPhone or the web, reinstall, switch accounts, and Apple's notice can arrive before the app has linked who they are.

02 / The solution

How I solved each one

Each sub-problem got its own component with a narrow contract, and every component reports its evidence to the same warehouse. The model never owns a threshold, a weight, or an action.

System architecture

Requests flow from the app through authenticated services to models and data. Jev judges; code decides. Every signal lands in the warehouse, which feeds gated decisions back into the product.

iOS appSwiftUI app583 Swift files, MVVM modulesWidgets + Live Activitiesplus an iMessage extensionStoreKit 2 + Superwallverified purchases, holdoutsFirebase Auth + App Checktokens on every requestauthServicesResearch agentgrounded search, enrich, reasonResearch assistantstreamed, evidence admitted firstLive voiceWebRTC, server-owned sessionSports data gatewaynormalized, cached, age-stampedEntitlement ledgersigned JWS, dedupe, repairJobs + schedules37 jobs, 120 schedulescallsModels + judgmentOpenAIvoice, conversation, expert tiersGeminigrounding, structuring, LiveBedrockopt-in reasoning chainTypeSafe Jevtyped, calibrated judgmentsApple, Superwall, AppsFlyersigned webhooks, exportseventsData + decisionsFirestoreoperational store, ledgersBigQuery warehouseDataform: 342 definitionsConstraint enginebinding constraint, weeklyDecision matrixscale, hold, or pull backDecision dashboardversioned snapshotsRankings and gated decisions flow back into the product
Solves P1

Noisy multi-source data, real-time decisions

Before

Each feature calls providers directly: duplicated requests, no shared cache, and no way to know how old a number is.

After

One gateway normalizes every provider, caches by volatility, and stamps every answer with its source and age.

  • A sports data gateway with one normalized schema, cache lifetimes from 30 seconds for live data to 24 hours for final scores, per-query locks, and a provider-wide rate cap.
  • Allowlisted provider paths and validated query schemas reject bad calls before they cost anything.
  • Every response reports source, fetch time, age, and whether stale data was served. Missing data returns an explicit partial state and is never counted as a zero.
  • A real-time service pushes live scores over server-sent events, with refresh work tied to open requests so idle time costs nothing.
Solves P2

Trust and calibration of AI answers

Before

One model call with open web search can cite anything and sound equally certain about everything.

After

Evidence is admitted before reasoning. Jev scores data quality and grounding, and code decides whether to escalate, soften, or withhold.

  • Web search runs only after licensed data shows a fact is unverified, and a result counts only with a completed search call and URL citations.
  • Jev answers up to 48 typed questions in one call within a 1.2-second budget: routing lane, data quality, grounding, and veto checks.
  • The composite (data quality 0.55, grounding 0.45) is computed in code. Any veto at 0.75 or above withholds; vetoes are never averaged away.
  • The expert model tier runs only after evidence is admitted and lane confidence clears 0.85. A model's own label can never unlock it.
Solves P3

Real-time voice

Before

A session opened from the phone needs a provider credential and cannot be metered, steered, or cut off by the server.

After

The server owns every session. The phone streams audio over WebRTC and can only mute, unmute, or hang up.

  • The phone posts its WebRTC offer to the backend, which creates the provider session and attaches a private server-side control channel.
  • A Firestore transaction reserves daily voice seconds with a lease before the session starts. Sessions cap at 10 minutes and unused time is refunded.
  • Tools, model routing, and writes run on the server. A write needs an on-screen confirmation of an opaque proposal ID.
  • Gemini Live streams read-aloud narration as 24 kHz PCM. The phone authenticates with Firebase and App Check, never with a model key.
Solves P4

Monetizing within App Store rules

Before

Paywall changes ride along with app builds, and experiment results are read before anyone checks that the split was fair.

After

Server-driven paywall experiments with holdouts, and a warehouse that will not call a winner until integrity gates pass.

  • StoreKit 2 accepts only verified transactions. Each purchase carries an app account token, so Apple never sees the raw user ID.
  • Superwall experiments log holdout outcomes with experiment and variant IDs; offer parameters are decided server-side and cached per placement.
  • Per-variant readiness gates: a sample-ratio check (standardized residual of 3.29 or less), crossover of 1% or less, 90% identity and purchase linkage, a minimum sample, and 7-day maturity.
  • Until every gate passes, results are labeled descriptive.
Solves P5

Growth measurement through ATT and SKAN blindness

Before

Every platform claims credit for the same purchase, and noise in a small sample reads like a trend.

After

One owner per conversion signal, a one-to-one rule for paid credit, and a warehouse that labels what is decision-grade.

  • A dedicated consent screen sets what each SDK may do. AppsFlyer is the only owner of SKAdNetwork conversion values.
  • Server-side conversion events use one canonical event ID per transaction and an atomic single-send claim, through a transactional outbox.
  • Paid credit needs a one-to-one match of client identity, Apple account token, attribution install, and campaign. SKAN stays directional; conflicts are quarantined.
  • A Dataform warehouse on BigQuery: 342 definitions, 149 assertions, and 43 registered sources, each with a freshness target.
Solves P6

Entitlements across App Store, web, and identity

Before

Access follows whichever signal arrived last, so one out-of-order notification can grant or revoke the wrong thing.

After

Only signed facts count, the newest signed state wins, and conflicting evidence is quarantined instead of guessed.

  • App Store Server Notifications V2 are verified against Apple's root certificates before anything is written. Sandbox payloads never grant production access.
  • Each notification is deduplicated by UUID into a transactional ledger. Local processing and forwarding are tracked separately, so Apple's retries finish whichever part is incomplete.
  • Older signed dates are rejected, and an older period can never overwrite a newer paid one.
  • A UUIDv5 app account token, derived the same way on iOS and the server, beats stored history. Scheduled repair re-links purchases that arrived early and can re-verify with the App Store Server API.
  • Apple expiry never removes active web access: each provider's entitlement is evaluated on its own.
03 / Innovation

What's innovative

The pattern that runs through everything: turn judgment into typed questions, let a calibrated model answer them fast, and keep the weights, thresholds, and actions in code where they can be tested.

Signature technique

Jev: calibrated decisions in real time

Jev is TypeSafe's fast, calibrated judgment model, which I use in my own products. DeepChamp never asks it for prose. It asks typed questions (a score on a rubric, a choice among options, or a yes-or-no) and gets back a probability distribution with a confidence. Code owns the weights, thresholds, and actions, and a deterministic path takes over on any timeout, rate limit, or malformed answer.

How a Jev decision is made

A compact snapshot becomes typed questions. Jev returns probabilities with a confidence; code applies the weights and gates. Any failure drops to the deterministic path, so the product never waits on a model.

Data snapshotcompact state, no personal dataTyped questionsscoreposition on a rubricchoiceone of N optionsyes / noP(true)Jev (TypeSafe)probability + confidenceCode-owned weightscomposite in codeConfidence gatelow: soften or skipActionrank a boardroute answergate a budgettimeout, rate limit, malformedDeterministic path: the rule-based order or router takes overJev informs the decision. Thresholds and actions stay in tested code.The Jev model version is pinned, because every threshold is tuned against it.
  • Ranking every board

    Rankings, news, and the live board are judged in batches of four. Weights: deterministic 0.34, attention 0.28, semantic 0.26, choice mass 0.12. Below 0.55 confidence a score shrinks toward neutral; below 0.35 it is ignored.

  • Assistant routing

    Up to 48 questions in one call, in 1.2 seconds. For the assistant, Jev can withhold escalation or soften speech. It cannot escalate on its own.

  • Revenue actions and constraints

    Each candidate action is scored on growth effect, evidence strength, time to impact, production risk, and whether it is the binding constraint. Priority is computed in code, and Jev is asked again only when the data snapshot changes.

  • Decision matrix for budget moves

    Seven stages (evidence, delivery, activation, conversion, economics, retention, momentum) end in hard gates: fresh data, integrity, sample size, economics floors, and a decisive posture. Output: scale, hold, or pull back, in shadow mode behind a separate operating check.

  • Onboarding funnel rails

    A typed contract (drop-off yes-or-no, readiness score, continue-or-paywall choice) with a 0.60 confidence floor and a 2-second budget. The deterministic rail decides first, so onboarding never waits on a model.

  1. 02

    Server-owned real-time voice

    Speech runs over WebRTC to a realtime model, but the session, the tools, the quota, and the credentials all live on the server. A separate Gemini Live path streams narration.

    • Budget reserved in a Firestore transaction before the provider session opens
    • Idempotent close, refunds for unused time, conservative charging when usage is missing
  2. 03

    Signed-JWS entitlement pipeline

    Apple's signed notifications are verified, deduplicated, and written to a ledger in transactions. Repairs are scheduled, not manual, and ambiguous identity evidence is quarantined rather than assigned.

  3. 04

    Warehouse constraint engine

    Every week it rescores seven constraints, from checkout completion and attribution to source health and support capacity, and marks the top open one as binding. It grants no spend authority by itself.

    • Spend-response model: least squares of daily new users on paid spend over 14 days, refit on 28
    • Decision-grade only with 10+ days, R-squared of 0.5 or more, and a positive slope; always labeled observational
  4. 05

    Experimentation with holdouts

    Superwall paywall experiments with holdouts, read through sample-ratio, crossover, linkage, sample-size, and maturity gates before any result is treated as more than descriptive.

  5. 06

    Multi-model orchestration

    Nine models from four providers, each with one job: OpenAI tiers for voice, conversation, and evidence-gated expert reasoning; Gemini for grounded research, structuring, and Live narration; Bedrock as an opt-in reasoning path that falls back to Gemini, then to grounded context; Jev for judgment.

  6. 07

    Deploy gates that run the tests

    38 service deploy scripts run pytest and abort on the first failure. Firestore rules and indexes ship in the same change as the code that needs them, and fork-safe services reset every lock and cache after the server forks.

04 / Result

The result

A live product where every answer, ranking, and growth decision can explain what it rests on.

  • 01

    A live product

    DeepChamp is on the App Store, with text and voice research, widgets, Live Activities, and an iMessage extension.

  • 02

    Decisions with evidence

    Rankings, routing, and growth calls each carry a Jev probability, a confidence, and a code-owned threshold, with a deterministic fallback.

  • 03

    Safe to change

    17,000+ tests, 149 warehouse assertions, and deploy scripts that refuse to ship on a failing test.

  • 04

    Entitlements that heal

    Signed, deduplicated, replayable purchase facts, with identity conflicts quarantined instead of guessed.

View on the App Store

06 / Stack

Stack and ownership

I lead the architecture and build hands-on across every layer: the SwiftUI app, Python and TypeScript services, the data platform, and release engineering, working with a small team and a fleet of AI coding agents.

  • SwiftUI
  • StoreKit 2
  • Python
  • TypeScript
  • Cloud Run
  • Firestore
  • BigQuery
  • Dataform
  • OpenAI
  • Gemini
  • Bedrock
  • TypeSafe Jev
  • WebRTC
  • Superwall
  • AppsFlyer

Counts measured on September 28, 2026 from the DeepChamp iOS and web repositories and the production cloud project. Business metrics are intentionally not published.

Private repository. Walkthrough on request.