Case study 03 / Working system, broadcasting 24/7
Auto DJ: an AI that analyzes, plans, mixes, and grades a set.
Point it at a music library. It measures every track, turns the measurements into words for a calibrated judgment model, plans a set that can never break a harmonic rule, renders each transition the way a DJ would, and grades the result from the audio before anything ships.
Private repository. Walkthrough on request. Architected and shipped by me, directing a team of coding agents.

- 0.3 msmedian beat offset between decks after closed-loop sync, down from 1.5 ms
- 92%key agreement with Mixed In Key within a legal move, on 113 tracks
- ~18klegal song pairs scored to plan one five-hour loop
- 28 to 3transitions at risk of drifting out of sync, first planner versus the loop search
- 0.823mean transition grade over 216 rendered transitions, up from 0.810
- 106tests over about 20,900 lines of Python in 89 modules
The problem
Turn a folder of songs into a set a working DJ would sign off on: every blend in time, in key, and on phrase, with taste, for hours, and proven to sound right.
A DJ makes a hundred small calls a minute on top of millisecond timing: when to bring the next song in, where its bass takes over, which song should come next, and whether the room is still with them. Auto DJ has to make every one of those calls, for a five-hour set, with nobody at the decks.
What made it hard
Six things have to be right at once, and each one breaks the mix on its own.
- 01
Millisecond timing
Kicks more than about 5 ms apart flam. A neural beat tracker lands within tens of milliseconds.
- 02
Harmony
Clashing keys ruin a blend even when the beats are perfect, and single key detectors often disagree.
- 03
The shape of a song
Mixing on phrase means knowing where the verse, build, and drop start, to the bar.
- 04
Taste
Among the technically legal ways to mix two songs, a few feel right. That is hard to write as a rule.
- 05
Order over hours
Hundreds of songs, every neighbour legal, one tempo, and still a shape to the night.
- 06
Knowing it worked
A plan can look right and still flam, clash, or stack two basslines once it is rendered.
How we solved it
One fix per hard part, each owned by a component with a narrow contract and its own test.
- 01
Millisecond timing
Fit one straight-line grid through the tracker's beats, nudge it onto the kick transients, then render each blend, measure it, and correct it until it locks.
Median offset between decks: 0.3 ms.
- 02
Harmony
Three key profiles vote. The Camelot rule (same key, one step, or the relative key) is enforced in the candidate query and checked again on the keys playing in the blend bars.
No song pair is ever planned outside the rule.
- 03
The shape of a song
Segment on 4, 8, and 16-bar phrase lines, find drops where the song gets suddenly louder and fuller, and let Jev label each section.
Cues and blends land on real phrase lines.
- 04
Taste
Jev chooses among options that already pass every rule, and its answer counts for 30%.
Judgment where it helps, never where it can break something.
- 05
Order over hours
Score every legal pair once, seed the loop with the hardest-to-place songs first, then search the whole cycle at once.
Transitions at risk of drifting fell from 28 to 3.
- 06
Knowing it worked
Grade every rendered transition from its own audio. A transition that will not lock within 2 ms is banned and the set is re-planned.
Nothing ships on the plan alone.
How it works
Nine stages, one plan. The same plan drives the audio render, the Rekordbox cues, and the grade, so they can never disagree.
How it works, end to end
Audio becomes measurements, measurements become words for Jev, and hard rules bound every choice. The gold badges mark where Jev makes a decision. A transition that fails the render gate sends its pair back to the planner, never to be used again.
Analyze: a grid you can mix on
A neural beat tracker finds beats and downbeats. Auto DJ fits one straight line through them, snaps the tempo to the whole BPM a producer used, and nudges the grid onto the kick drum. Three key detectors vote on the key. Stem separation says, bar by bar, whether drums, bass, and vocals are playing.
Structure: where the song goes
Sections start on phrase lines, so the segmentation prefers boundaries every 4, 8, and 16 bars. The drop is where the song suddenly gets louder, fuller, and busier. Jev reads the measurements as words and labels each section.
Plan: legal, then good
Candidates come from a query that only returns key-legal, tempo-safe songs. Each legal pair gets a value from key, tempo, phrase fit, and energy, plus Jev's choice. For a five-hour loop, every legal pair is scored once and the whole cycle is searched at once, shaping the night toward a target energy curve.
Render: mix the way a DJ does
Volume, a three-band EQ, and one reverb, nothing more. In the house drop-swap, the outgoing song's bass is gone within a quarter bar, the incoming song's bass arrives on its drop downbeat, so two basslines never play at once. Each deck is rendered, measured against the other, and corrected until the beats lock.

Rendering in the cloud
Fast, cheap, and self-cleaning. A laptop cannot render an hour-long mix and its video quickly, and the cloud has to be cheap and never leave anything running.
Plan locally, render in the cloud
The laptop plans in seconds; one Spot VM does the heavy work and deletes itself; only the finished audio and video come back. A janitor, a hard run limit, and a budget kill switch make sure nothing is left running.
Speed
The laptop computes the mix plan and the edit list in seconds; a single Spot VM does everything heavy. The mix renders in parallel on every core, cut into chunks that join only where one song plays alone, so the result matches a one-at-a-time render. The video renders as roughly 30-second chunks, each one ffmpeg graph encoded once and uploaded the moment it is done. The chunks join without re-encoding, and the audio is muxed bit-identically.
Cost
Spot machines only, chosen from measured price-performance on the same 20-minute slice. Sources live in a same-region bucket, so reads are free, and only the finished video comes back, because download is the one real cost. Each source is color graded once and cached. Scratch storage deletes itself after 14 days.
Failure handling
If Spot reclaims the VM, a new one takes over and skips every chunk already uploaded. A heartbeat reports progress, CPU, and memory; a VM that goes silent for 15 minutes is replaced, and DONE and FAILED markers end every run. Every chunk's frame count is checked, the total must match the timeline exactly, and the video passes an automated QA before it is delivered.
No lingering spend, and least privilege
VMs delete themselves, with a hard run limit as a backstop, and a janitor runs before and after every render. An audit command lists everything billable. A budget alert triggers a function that deletes every render VM and blocks new renders for the month. Rendering runs in a dedicated project under its own service accounts with narrow custom roles, fully isolated from the 24/7 stream. The toolchain is pinned and checksum-verified, with Python packages held to the lockfile and no container registry.
| Machine | vCPUs | 20 min of video in | Speed | Per hour of video |
|---|---|---|---|---|
| c3d-highcpu-180 | 180 | 45 s | 26.5× real time | $0.062 |
| c2d-highcpu-32 | 32 | 3.4 min | 6.0× real time | $0.103 |
| n4-highcpu-32 | 32 | 3.7 min | 5.4× real time | $0.129 |
| c4d-highcpu-16 | 16 | 4.5 min | 4.5× real time | $0.057 |
| c3d-highcpu-16 | 16 | 6.0 min | 3.3× real time | $0.044 |
- A five-hour video in 45 minutesOn one 32-vCPU Spot VM, 6.7× real time with the CPU 99% busy, every frame accounted for.
- A five-hour mix in 12 minutes225 songs rendered on 32 cores, then loudness, loop trim, and lossless and MP3 encodes.
- Identical where it mattersA 12-song mix rendered in the cloud gave the same sync verdicts, level rides, loudness, and peaks as the laptop.
How a mix gets graded
Two grades, for two decisions: one picks the set before anything is rendered, and one decides what ships after.
How a mix gets graded
Two grades. Before rendering, every legal pair gets a value that decides the set order: code-owned signals, Jev's choices, and the transition's own rule score, minus a risk term. After rendering, every transition is graded from its own audio, and the gate decides whether it ships.
Before rendering
Every legal pair gets a value. Code scores what it can measure: how smooth the key move is, how far the next song is stretched, whether the verses line up and the mix lands on a build, and how close the next song sits to the night's target energy. Jev adds the taste: which legal song should come next, and whether each run of three songs flows. The transition's own rule score and a sync-risk penalty finish it, and the loop search picks the order with the best total.
After rendering
Each transition is graded from the audio that was made: four bars of the outgoing song, the blend, and four bars of the incoming song into its drop. Beat lock and harmony weigh the most. Anything that fails the gate is banned from future plans and the set is planned again. Across 216 transitions, the whole-loop plan averaged 0.823, against 0.810 for planning one song at a time.
One thing we tested and dropped: asking Jev to re-score these measured facts added nothing. Its strength is judgment from words, so grading the audio stays in code.
The Jev decision system
Jev, TypeSafe's calibrated judgment model, makes the taste calls. It never sees an option the rules would reject, and low confidence never acts on its own.
Where Jev decides
Four decision points, each a typed question over options the rules already allow. Every answer comes back as probabilities with a confidence, and code decides how much it counts.
| Decision | Typed question | How the answer is used | When it is not trusted |
|---|---|---|---|
| Section labels | Choice among nine section types, plus yes-or-no: does the drop kick in here, does this lead into it | Pooled with the rule-based label: | A cue section under 0.2 confidence goes to a Needs review playlist |
| Key tie-break | Choice among the detectors' own candidates, never a free answer | Overrides the vote | Only at 0.6 confidence or more |
| Next song | Choice among the top 24 legal next songs | 30% of the pair value, next to 55% code and a lookahead | On any error the planner uses its own score |
| Flow of three | Score: is this three-song run weak, fine, or great | Steers five rounds of loop search away from weak runs | Unrated runs take the average |
| Transition window | Choice among up to four legal windows within reach of the best | 30% of the pick; the rule score keeps 70% | Fewer than two options or any error: the rule pick stands |
How confidence gates a decision
Measurements go in as words, Jev answers typed questions with probabilities and a confidence, and code decides what that answer may change. Low confidence never acts on its own: it falls back to the rule answer, sends the track to review, or leaves the plan as it was.
- Rules first. Every option Jev sees already passes the key, tempo, bass, and sync rules.
- Code owns the weights.Jev's share is 30% or pooled at 70/30, and it is set in code, not by the model.
- Proven by test. A property test swaps in a random critic and checks that every hard guarantee still holds.
- Cached and pinned. Every question and answer is cached by content, so a re-plan never pays twice.
Results
Mixes that lock, sets that stay in key, and a system that plays and broadcasts around the clock.

- A mastered mixLoudness matched to -12 LUFS under a -1 dBTP true-peak limiter, every blend locked or re-planned.
- Rekordbox cuesHot cues A to H and 8 and 4-bar countdowns to each drop, so a human can play the same set by hand.
- A 24/7 broadcastPre-rendered mixes loop to YouTube and Twitch behind ingest checks and a watchdog.
The songs can come from my AI music pipeline, which generates and quality-gates them before they are ever analyzed.