All writing

2026-09-28 / 7 min read

Debugging tool calls across distributed systems: a field guide

Four incidents from a production AI app: a provider outage, signed webhooks we rejected ourselves, an enforcement flag that nearly locked out paying customers, and a latency fix hiding a timeout. Each ends with what I do now.

In production, a single tool call crosses a model, your server, a queue, a vendor and a database, and often a webhook comes back the other way. When something breaks, the symptom shows up far from the cause, and the most convincing hypothesis is usually wrong. What follows are four incidents from DeepChamp, a sports research and analytics app I build and run. Each is written up as symptom, hypothesis, evidence and fix, followed by the guardrail that came out of it.

1. The provider that said 503 to every batch

Symptom. An internal decision panel showed "assessment incomplete" and no verdicts. The panel scores seven stages of our growth funnel with one batched call to a decision model, then makes a synthesis call.

Hypothesis. Throttling. That path had been rate-limited before.

Evidence. The logged error text said otherwise. The AI gateway in front of the model was returning "service temporarily unavailable" for every batched call and for most synthesis calls. Every retry is a paid call, so our retry policy retries automatically only on explicit throttles. Faced with a wall of 5xx errors, it left the panel empty. The same investigation turned up a bug in our error classifier. "Delay was aborted" is the SDK's own 429 back-off being cut short by our timeout, which makes it a rate limit, not a slow model.

Fix. We made three changes:

  • Call the model provider directly. The gateway adapter stays as a disabled fallback behind an explicit flag.
  • When the batched call doesn't return every verdict, re-ask each stage on its own. Auth and budget errors are the exception, because re-asking can't fix them. One bad answer now costs one stage instead of seven.
  • Cache only complete results, keyed on the input snapshot and given a TTL, and have concurrent page loads share one in-flight evaluation.
// Only complete matrices are cached; a throttled or failed run must be retried.
if (matrix.telemetry.failedCalls > 0 || matrix.synthesis === null) return false;

A related ranking job now reuses its last result while the input fingerprint is unchanged and retries failures at most once an hour. During an outage it shows the last good result, labeled STALE.

What I do now. Sort every error into one of five classes (throttle, unavailable, auth, budget or invalid response) and give each class its own retry policy. Batch to save cost and fall back to per-item calls to limit the blast radius. Never cache a partial result, and always make staleness visible.

2. Signed notifications we rejected ourselves

Symptom. Subscription events stopped reaching our ledger. A downstream vendor that gets App Store events only through our forwarder was missing them too.

Hypothesis. The obvious suspects: a delivery problem on Apple's side, or our endpoint being down.

Evidence. Apple's notification history API lists every notification along with each delivery attempt. It showed our endpoint rejecting every retry of a few dozen notifications over two days. The endpoint was up; it was failing verification. A deploy had shipped without Apple's public root certificates. A broad *.pem ignore rule, written to keep private keys out of uploads, had excluded the public trust roots too. Verification crashed before anything was recorded or forwarded. Apple retries a failed notification only a few times over several days, then gives up.

Fix. We restored the roots and added a deploy preflight. It inspects the upload list the deploy tool will actually send, confirms that both root certificates are present and match pinned fingerprints, and rejects anything that looks like a credential. We also changed the handler:

  • It records every verified notification by its UUID before any side effect.
  • It treats local processing and downstream forwarding as independent legs.
  • It acknowledges Apple only after the handoff succeeds. Otherwise it returns 503, so Apple retries.

Then we pulled the original signed payloads from Apple's history and replayed them through the same endpoint. Signature verification plus UUID dedupe made the replay safe, and the reconcile report dropped to zero missing.

What I do now. The sender's history is the source of truth, so reconcile against it daily. Ours checks the last 14 days and replays anything missing. Make handlers idempotent on the sender's id so that replay is routine, and test the artifact you deploy, not the repo.

3. The flag that would have locked out paying customers

Symptom. None yet, and that was the point. We were about to switch on a flag that made the server enforce subscription expiry for the premium voice feature.

Hypothesis. Any account whose stored expiry had passed had lapsed.

Evidence. A dry run listed every account the flag would deny. We checked each subscription lineage against Apple's signed current status through the App Store Server API. Most of those accounts had genuinely lapsed, been refunded or run out of billing grace. A handful, though, were paying renewers whose expiry dates had gone stale. Two writers were to blame:

  • A one-time identity repair had taken the first stored purchase row and skipped renewals. When the repair code was later fixed, accounts it had already processed never ran again.
  • The webhook forwarded plan upgrades without applying them, so the billing periods of upgraders went stale too.

The logs confirmed that no paying renewer had been denied anywhere yet.

Fix. Repairs now write the billing period from Apple's verified current state, and only when the signed token maps to exactly that user. They never fall back to stored rows. We backfilled the affected accounts, and the webhook now applies upgrades. The flag shipped with a self-heal step. Before the server denies an account that's free only because its expiry has passed, it re-verifies with Apple. It logs one of two events, denied or denial_prevented, and any "prevented" line means a writer is still stale.

What I do now. Before an enforcement flag goes on, dry-run it against production and check every would-be denial against the source of truth. Ship enforcement with a self-heal check and a metric that exposes drift in the writers.

4. Two paths, one reply

Symptom. Answers over iMessage felt slow to start.

Hypothesis. The model.

Evidence. Per-stage timing on test messages showed 2.2 to 2.7 seconds between storing the inbound message and starting the turn. Of that, 1.7 to 2.0 seconds was the hop through the task queue, and the model hadn't even started yet.

Fix. We now start the turn on the webhook request itself. Our serverless platform gives CPU only to open requests, so the turn runs on a worker thread while the request stays open. The queued task is still scheduled 15 seconds out as a safety net. A transactional claim lets exactly one path run the turn. The other path either sees completed, or gets already_running (a 429) and the queue retries later. Every send carries an idempotency key derived from the inbound message id, so even a double run can't text twice. We rolled it out behind an off/allowlist/on flag, starting with a test phone. In a dry run, the turn started about 2 ms after ingest.

A timeout was hiding in the first version. It held the request open for up to 120 seconds, which is exactly the platform's request timeout. Our bound and the platform's would have raced, and when the platform wins, you get its error, not yours. We caught it before it shipped:

def max_wait_seconds() -> float:
    """Under the service's 120 s request timeout (the bridge itself gives up after 85 s)."""
    try:
        return max(5.0, min(110.0, float(os.environ.get(MAX_WAIT_ENV) or 100)))
    except ValueError:
        return 100.0

What I do now. Every wait must be strictly shorter than whatever contains it, and a test should pin that bound. Duplicate execution paths need both a claim and idempotency keys. On serverless, any work that has to finish must finish while a request is still open.

The field kit

The same few tools did most of the work across all four incidents:

  • One trace id per turn across client, server and provider, with hashed turn ids and per-phase timings instead of transcripts.
  • Bounded error codes in logs. "Temporary error" is a symptom, not a cause.
  • A dry-run mode on every repair and enforcement tool, and a read-back after every write.
  • Replay from the sender's history, made safe by deduping on the sender's ids.
  • Deploys that mirror the serving config. A redeploy with default flags once silently turned a feature off for three days. Now we read the live revision's settings and pass them explicitly.

The bug is rarely in the model. It's in the handoffs, and the way to find it is to make every handoff leave a record.