All writing

Field note 01 / Decision dashboard

A gateway 503 on every call, and a panel with no verdicts

How a batched model call through a failing gateway blanked a seven-stage decision panel, and the per-stage fallback and snapshot caching that fixed it.

  1. 01

    Symptom

    The decision dashboard showed an incomplete assessment. Every batched Jev call through the AI gateway came back as service temporarily unavailable, and the whole seven-stage panel paused with no verdicts.

  2. 02

    Diagnosis

    • SDK retries were off, and a 5xx response was not classified as a throttle, so nothing retried or backed off.
    • One batched request carried all seven stages, so a single failure emptied the entire panel.
    • The failure was in the gateway, not the model: the same questions succeeded against the provider directly.
  3. 03

    Fix

    • Made the direct provider the default path, with the gateway still available behind a single configuration switch.
    • When a batch comes back incomplete, re-ask the missing stages one by one, except on 401 and 402, where a retry cannot help.
    • Cached the new growth-action ranking by an input fingerprint: no calls while the snapshot is unchanged, retries at most hourly, and the last good result stays on screen marked stale.
What I do now

Degrade by the stage, not by the page

  1. Classify 429 and 5xx as retryable, honor retry-after, and back off exponentially.
  2. Never let one batched call own every answer on a screen; re-ask only the missing parts.
  3. Keep the provider path switchable by configuration, not by a deploy.
  4. Fingerprint inputs and skip paid calls when nothing has changed.
  5. Show the last good answer with a stale label instead of an empty panel.
  6. Do not retry authentication or billing errors.