All SDK docs

How the Engine Works

A complete, measured walk through the SDK's drawing engine: the bus rule, the beam model, a frame command by command, the PIO program, the DMA ring and the dual-core split.

This page explains how a native ARM cartridge draws on a Vectrex: what the hardware asks for, the model the SDK uses to answer it, and the machinery that keeps the beam busy while your game is thinking. Every number on it was either measured on a console or produced by the SDK's own emitter. The last section shows how to reproduce them.


1. The problem in one paragraph

A Vectrex has no framebuffer. Its picture is drawn by an analog beam, steered by two integrators that a VIA 6522 controls: the VIA's DAC sets a rate, its Timer 1 decides how long the integrators run, and the beam traces a line while they do. A normal cartridge feeds 6809 code to the console's CPU, and the CPU writes those VIA registers. An SDK cartridge holds the 6809 in /HALT and writes the VIA itself, from an RP2350, over the same bus.

That bus does one access per E period. E runs at 1.5 MHz, so one access is 667 ns and a 50 Hz frame is 30,000 bus cycles. That number is the whole drawing budget: a faster CPU cannot make a bus cycle shorter.


2. The one bus rule

The 6800-family bus is not symmetrical. In the first half of the E period (E low) the master changes the address; in the second half (E high) the VIA decodes it, and address, chip select and R/W must not move; the write is latched on E's falling edge.

The cartridge does not see E. It sees ~E on cartridge edge pin 12, the inverted E, which free-runs at 1.5 MHz whatever the bus is doing. Its rising edge is E's falling edge, so one edge means two things: "the previous write has just landed" and "it is now safe to present the next one".

The 6800 bus rule: change during E low, hold through E high
E high: the VIA decodesaddress must be stableE~E (pin 12)A0–15, D0–7write nwrite n+1write n+2latchlatchlatch245 ns · inside E low ✓0 ns333 ns667 ns1000 ns1334 ns

The state machine waits for ~E to rise (= E falls: the previous write has just been latched), then presents the next address and data about 245 ns later, inside E low and about 88 ns before the VIA starts decoding. The value is held until the next falling edge latches it.

Measured on a console: presenting within 30 CPU cycles of the latch always draws; presenting within 70 gives a black screen. One E period is 100 CPU cycles at 150 MHz, so E low is the first 50. A setup-time violation degrades; a phase violation fails outright. That asymmetry is why everything below is organised around one edge of one clock.


3. The beam model

Speed × time

A stroke (dx, dy) becomes three numbers: a signed 8-bit rate for X, one for Y, and a duration, T1, in bus cycles. The integrators move the beam at the rate for as long as T1 runs. So the same length can be drawn fast and short or slow and long, and the model has to choose.

ramp_params in vectrex-draw/src/ramp.rs (Rust, shared byte for byte with the Debug Cart's firmware) makes that choice. With the default calibration, every stroke the SDK emits satisfies:

length (device units) ≈ rate × (T1 + 2.5) / 160
  • 160 is DRAW_SCALE, the divisor. Larger values give shorter strokes for the same numbers.
  • 2.5 is t1_tail_q8 = 640 (640/256): the ramp keeps running about 2.5 E cycles after T1 expires, because the 6522's one-shot counts t1 + 1.5 and the beam needs time to stop.
  • 127 is the DAC's limit. A rate can never exceed it, so a long stroke is bounded by its duration, not its speed.
  • 8 is the floor (MIN_T1): no ramp is shorter than 8 bus cycles. With the floor removed, 27% of one game's strokes got ramps of 3 to 7 cycles and its geometry broke.
How the beam model splits a stroke into time × speed
T1: how long the ramp runs (bus cycles)
0501001255075100stroke length (device units)
Rate: what the DAC feeds the integrator
0501001255075100stroke length (device units)DAC limit 127

What the SDK actually writes to the VIA for a chained horizontal stroke, read back from the emitted list (default calibration). T1 moves in steps of 8 bus cycles; inside each step the rate grows with the length until it nears the DAC's limit of 127, then T1 takes the next step and the rate drops. Every point satisfies length ≈ rate × (T1 + 2.5) / 160.

The rule the chart shows: T1 is the smallest step that keeps the rate inside the DAC. An earlier version capped the speed at 42 to gain brightness, which doubled T1 for 109 of 427 glyph strokes in one screen. One ramp of T1 = 16 does not travel what two ramps of T1 = 8 do, and the high-score table came out with its rows stacked. Brightness comes from Z (the intensity register), not from slowing the beam.

Sub-units and the debt

Coordinates travel through the SDK in 1/16 of a device unit (UVM2_SUBUNITS, UVM2_Q_BITS = 4), so a 2.5-unit stroke stays 2.5 units. Splitting a distance into an integer rate and an integer T1 still loses a fraction, so ramp_params_chain carries the remainder forward as a debt and pays it into the next stroke of the chain. Errors therefore do not accumulate along a figure. With input finer than 1/64 the input is already almost exact, and the correction starts adding error instead of removing it, so the debt is turned off.

Re-zeroing

Integrators drift. The SDK clamps the beam back to the centre (/ZERO) when a blanked jump is longer than 24 device units (uvm2_zero_jump), and always before a frame's first jump. The value primed into the zero reference belongs to the console, not the game: it is the zero field of the calibration (default 35, 0x23). A wrong value adds the same velocity to every ramp, which is why text on an uncalibrated console leans into diagonals. See Calibrating a console.


4. A frame, command by command

Here is everything the SDK sends to the VIA for one tiny frame: set intensity 100, jump to (−30, −30), then draw a triangle as three chained strokes.

Anatomy of a real frame: 52 commands, 639 bus cycles
frame overheadbrightnessblanked jumplit strokes
VIA setupclampZre-zerojumpstroke →stroke ↑stroke ↙closePort A · DACPort B · mux, /RAMPT1 · ramp timerSR · beam on/offPCR · /ZERO clampDDR, ACR/ZERO clamp onramp runningbeam lit0100200300400500600bus cycles (1 cycle = 667 ns)

Emitted by the SDK for: set intensity 100, jump to (−30, −30), then three chained strokes forming a triangle. Each mark is one VIA write (1 bus cycle); the thin line after it is the delay the list asks for before the next write. The bottom three rows are the beam's state, replayed from the list. Hover any write to decode it. 639 cycles is 426 µs, or 2% of a 50 Hz frame.

Reading it from left to right:

  1. VIA setup (6 writes). Port A becomes all outputs (the DAC). ACR = $98 makes Timer 1 a one-shot that drives PB7, which is /RAMP, and makes the shift register drive CB2, which is /BLANK. From here on, writing T1's high byte starts a ramp and writing SR switches the beam on or off, with no extra commands.
  2. Clamp. PCR = $CC asserts /ZERO: both integrators are held at the centre.
  3. Brightness. The DAC gets 100 and the mux routes it to Z for a few cycles, then lets go. The value stays in Z's sample-and-hold.
  4. Re-zero. Y rate 0, then the zero reference: 35 for about 7 cycles, then $FF for 4. The reference only charges partially, which is why the right value depends on each console's analog parts. Then PCR = $CE releases the clamp.
  5. The blanked jump. The Y rate goes into Y's sample-and-hold through the mux; the X rate stays on the DAC, which feeds X directly. T1 = 39 starts the ramp with the beam dark.
  6. Three chained strokes, 6 writes each: Y rate, mux on, mux off, X rate, T1 low, T1 high. The first one is preceded by SR = 1: beam on.
  7. Close. SR = 0 blanks the beam and /ZERO clamps it to the centre again, where it waits for the next frame.

Two things this makes visible:

  • A chained stroke is 6 writes and 15 cycles of settling, plus the ramp's own wait. The waits are real: 3 cycles for the DAC to settle, 9 for Y's sample-and-hold to follow it. Nothing here is padding.
  • 92% of this frame is waiting (587 of 639 cycles). Writes are 1 cycle each; the analog side needs time. This is why the list carries a delay field instead of the CPU spinning, and why the bus can carry other work in the gaps: digitised samples are injected right before each T1CL, where the timer has expired and the integrators are frozen (see Sound).

5. What things cost

What a stroke costs, by length
chained strokejump + stroke
0200400020406080100stroke length (device units)bus cycles per strokechained strokejump + stroke

Bus cycles per stroke, measured on the real emitter (16 strokes each, averaged). A chained stroke is always 6 commands; below about 6 units it costs the 32-cycle floor whatever its length. A stroke that needs a blanked jump first costs 14–21 commands and 2.4–3× the cycles. Hover for the numbers.

The consequences, in order of impact:

  1. Chain your strokes. A stroke that starts where the last one ended costs 6 commands. One that needs a jump first costs 14 to 21, and 2.4 to 3 times the cycles. The number of jumps is set by how many disjoint polylines you draw; reordering strokes can shorten jumps but not remove them. Measured on a real scene, naive nearest-neighbour reordering made it worse (709 → 882 jumps), because it broke up existing chains.
  2. Fewer, longer strokes. Below about 6 units every stroke costs the same 32-cycle floor, so splitting a line into short pieces is pure loss.
  3. Don't re-set an intensity you haven't changed. It costs three more writes and a mux cycle each time.

An empty frame (open, close, nothing drawn) costs 14 commands and 138 cycles. Everything else is your geometry:

Frame budget estimator
030,000 cycles = one 50 Hz frame
20,758
bus cycles
69%
of the budget
50
frames per second · ✓ on target

An estimate from the measured costs above: 138 cycles to open and close a frame, one blanked jump per figure (a figure is one connected polyline), and every other stroke chained. It leaves out brightness changes, extra re-zeros and text, so treat it as a floor. The bus runs at 1.5 MHz whatever the CPU does: when the list is longer than the budget, the frame rate drops.

When a list is longer than the budget the picture does not break; the frame just takes longer. A real arcade port draws much more than a menu: one Donkey Kong port's stream measured 69,718 cycles (46.5 ms), which caps it near 21 frames per second however fast the CPU is.


6. The command list

Core 0 never writes the bus while drawing. It records VIA writes into a list, and core 1 replays them. Each command is 24 bits:

The two formats: the command list, and what the PIO receives
Command (in the list, 3 bytes)delay: idle cycles after the write (0–4095)2312VIA reg118data70PIO word: write3126the bus word: address + data, as the pins expect them25110PIO word: one period of silence3120100PIO word: park for N periods3126N − 12521100

The list stores 3 bytes per command (a 64 KB buffer holds 21,845 commands instead of 16,384). The PIO reads bit 0 first: out shifts from the low end, so the sentinel has to live there. Bit 0 = 1 means drive the bus. Bit 0 = 0 means don't, and bit 1 then picks a single silent period or a park of N periods. One park word replaces about 34 words per vector, roughly 8,400 FIFO writes per frame. The bus word is 25 bits on the Debug Cart and 27 on the UVMC2 (drawn here at 25).

The 12-bit (register, data) field maps straight onto GPIO0–11 on the UVMC2, so the executor's inner loop does no shifting. The 12-bit delay is also why T1_TRANSPORT is 160: 4,095 is the most the executor can wait (2.73 ms), and a larger value would spill into the register field and corrupt the command with no warning.

Reads cannot be recorded. A read needs the data bus turned around mid-cycle. Every read in the SDK (the controllers, the PSG) happens between frames, while /ZERO holds the beam at the centre. Even there it is not free: with the input read removed entirely, stray bright vectors dropped from 4–5 per frame to 1, which is why /RAMP is held off during the whole conversion.


7. The pipeline

From your game to the beam
Core 0your gamerecords VIA writes3 bytes eachCommand listtwo buffers:frame n goesto buffer n & 1Core 1decodes eachcommand into1 or 2 PIO wordsBatch ring2 × 64 words,≈ 43 µs ofbus eachDMA ch 0paced by thePIO FIFO'sDREQPIO SMone word perE period, syncedon ~EVIA 6522DAC, mux,T1, SR →the beamframe_request++frame_done++core 1 only: it owns the bus pins exclusively

Core 0 never touches the bus. It publishes a finished list by bumping uvm2_frame_request; core 1 replays it and bumps uvm2_frame_done. Both are monotonic 32-bit counters, so no lock is needed. Core 1 turns each 3-byte command into PIO words; the DMA moves them into the PIO's FIFO as fast as the state machine drains it; the state machine puts exactly one word on the bus per E period. If the FIFO ever runs dry, the bus parks instead of glitching.

Core 1's loop, one frame

for (;;) {
    wait until uvm2_frame_request != served;   // nothing published: read input once per period
    uvm2_exec(buffer[served & 1], length);      // decode → PIO words → batch ring → DMA
    read buttons and axes;                      // between frames, beam clamped at centre
    drain the PSG queue; tick the music;        // by elapsed Vectrex time, not per frame
    pace to the frame period (TIMER0);          // 50 Hz = 20 ms, against the clock
    uvm2_frame_done = served;
}

Each step is timed into uvm2_stats (us_wait, us_exec, us_input, us_rest), so you can see on a running console where core 1's time goes. See Debugging.

Pacing is set against the clock, not against the list. An earlier version padded each frame by period − list cycles, which assumes nothing else happens between frames. Input, the PSG and the SD card all do, so frames ran 20.5–26.5 ms instead of 20. The pacer now targets "the previous frame ended at T, this one ends at T + 20 ms" and counts an overrun if it is already late.

The PIO program

vectrex-bus/src/bus_stream.pio is about a dozen instructions. The loop:

start:
    wait 0 gpio 31          ; ~E low
    wait 1 gpio 31          ; ~E rises = E falls: the previous write is latched
    nop [14]                ; 15 PIO cycles = 100 ns: phase calibration
present:
    pull noblock            ; next word, or X (the park word) if the FIFO is empty
    out y, 1                ; bit 0: sentinel
    jmp !y, no_write
    out pins, 25            ; address + data, inside E low (patched to 27 on the UVMC2)
    jmp start
no_write:
    out y, 1                ; bit 1: park or silence?
    jmp !y, silence
    out y, 24               ; N - 1
park:
    mov osr, x
    out null, 1
    out pins, 25            ; drive the park word...
    wait 0 gpio 31
    wait 1 gpio 31          ; ...once per E period
    nop [14]
    jmp y--, park
    jmp present             ; already in phase: don't burn another period
silence:
    out null, 25
    nop                     ; equalise the two paths

Every line is there for a measured reason:

  • The order of the waits decides which half of E the write lands in. With wait 1 before wait 0, the address changed 484 ns after the latch, inside E high, and the screen was black. Swapping them moves the presentation back by half a period (333 ns).
  • nop [14] is the remaining fixed offset. The path known to draw presents at ~245 ns; after the swap this one presented at ~132 ns, so 15 PIO cycles (100 ns) put it where the working path does. It is 14 and not 16 because the sentinel adds two instructions before the presentation, and those two cycles had to be given back. Adding an instruction to this loop changes this number.
  • pull noblock makes underrun the idle state. With an empty FIFO a plain out stalls with the pins holding the last word, and a VIA register left selected is re-latched on every E fall. That is harmless for a port register and disastrous for T1_HI, which restarts the ramp 1.5 million times a second. pull noblock instead loads X, which holds the park word: an address nothing on the Vectrex decodes ($8000 on the UVMC2, $C000 on the Debug Cart). A starved stream parks the bus. So DMA is an optimisation here, not a correctness requirement.
  • Silence is not the same as a write. The analog gaps inside a stroke are silence on a working bus. Emitting a write every period instead (2,390 writes became 17,314 words per frame) kept the figures but broke stroke chaining.
  • One park word for N periods. The ramp waits used to push about 34 park words per vector, around 8,400 FIFO writes per frame. A repeated park does the same with one word.
  • jmp present, not jmp start, at the end of a park. The park's last pass is already in phase. Jumping back to start waited for another E edge and lost one period per delay. On Major Havoc that was 2,882 delayed commands per frame, 30,000 cycles asked for and 32,882 spent: 9.6 of the 11.3 points it was missing from 50 Hz.
  • No label after the last instruction. A jmp to a label written after the final instruction assembles without error and points one slot past the program. Empty PIO memory reads 0x0000, which decodes as jmp 0: the state machine ran its preamble, made one pass and spun forever. Nothing that was measured (word layout, phase, preamble, branch timing) was wrong. Disassemble the .pio; don't trust the parser.

One out drives everything. The whole bus word goes out in a single out pins, N. On the UVMC2 the window is GPIO0–26: D0–D7 on 0–7, A0–A13 on 8–21, A14–A15 on 24–25 and R/W on 26, so N = 27; the two lines inside it that must not be driven (PB6 on GPIO22, and GPIO23) just get a direction bit of 0. On the Debug Cart the layout differs (A0–A15 on GPIO4–19, the bus direction on 20, D0–D7 on 21–28) and N = 25. The .pio is assembled once and the width is patched into out pins and out pindirs at install time, so the two can never disagree.

The DMA ring

Core 1 pushes words into a double batch ring, 2 × 64 words. 64 words is about 43 µs of bus: enough to amortise starting a DMA, short enough that a half-full batch never holds the drawing up. When a batch fills, or when the DMA channel is idle, it is handed to DMA channel 0, which writes into the PIO's TX FIFO paced by its DREQ: it only moves a word when the FIFO has room. The CPU fills one batch while the DMA drains the other.

The DMA's job is to remove spread, not offset. pull noblock never waits, so a word that arrives late is presented in the next period. Without the DMA, the moment a word reaches the FIFO depends on what the CPU is doing; with it, the FIFO is simply never empty while a list is playing.


8. Dual core: what the overlap buys

One core against two: what the overlap buys
core building frame nbus replaying frame n
one corecore 0build 0beam 0build 1beam 1build 2beam 2build 3beam 3build 4beam 4build 5beam 5build 6two corescore 0build 0build 1build 2build 3build 4build 5build 6build 7build 8build 9build 10core 1beam 0beam 1beam 2beam 3beam 4beam 5beam 6beam 7beam 80 ms40 ms80 ms120 ms160 ms200 ms
33.3 fps
one core · 30.0 ms per frame
50.0 fps
two cores · 20.0 ms per frame

A model, driven by the two numbers you set: how long your game takes to think and build a frame, and how long the beam takes to replay it. With one core they add up. With two, core 0 builds frame n+1 while core 1 replays frame n, so the frame costs whichever is longer. Neither number gets smaller: the beam is paced by the console's 1.5 MHz clock.

Dual core does not make drawing faster. The replay is paced by the console's clock, and 30,000 cycles is 20 ms by definition. What it buys is that the game's own time is no longer added to the beam's.

Measured on a Donkey Kong port:

configurationframe
single core, CPU-driven bus51 ms of drawing plus 46 ms of game logic, in series: 97 ms per frame
dual core, PIO + DMA40.25 ms per frame, of which the beam is busy for 39.51 ms (98%)

At 98% the beam, not the CPU, is the limit; the only way to go faster from there is to draw less. That is why uvm2_bus.h refuses to compile without UVM2_DUAL_CORE and UVM2_PIO_STREAM: three games once ran for weeks without either, and the build said nothing.

Who owns the bus

Core 1, exclusively, from the moment it starts. Two cores driving the same GPIOs means two writers on the VIA with no arbitration. Everything that touches the bus already went through one place (WAIT_RECAL: the replay, the controllers, the audio tick), and it moved to core 1 as a whole. Your game still reads buttons through the same calls, which answer from a cache refreshed between frames. The two exceptions:

  • PSG writes can happen at any moment, so they go through a 64-entry lock-free ring that core 1 drains between frames. When it overflows it drops the oldest entries, because for a sound chip the newest write is the current state.
  • Raw bus reads and writes are refused while core 1 owns the pins.

The handshake

extern volatile uint32_t uvm2_frame_request;  // frames core 0 has finished building
extern volatile uint32_t uvm2_frame_done;     // frames core 1 has finished replaying
// the buffer for frame n is n & 1

Two counters that only ever increase. A 32-bit load cannot tear, so no lock is needed; a dmb barrier before each increment makes sure the data is visible before the count. Core 0 waits before writing buffer n & 1 until frame n − 2 has been replayed, which is what the "two cores" row above shows when the beam is the slower side.

When the game is slower than the beam, core 1 has nothing new to draw. It keeps reading the controllers once per period, so a game that stops publishing still sees its buttons. By default the screen stays dark until the next list arrives: a game at 14 fps showed its picture for 22 ms out of every 70, and flickered. Build with -DUVM2_REPLAY_LAST and core 1 redraws the last list instead, holding its buffer so the game cannot overwrite it mid-replay; uvm2_stats.replays counts how often it did.


9. Two cartridges, two splits

The engine is the same code on both cartridges. Which core does what differs:

UVMC2Vectrex Studio Debug Cart
your game runs oncore 0core 1
the list is built bycore 0 (your game, through the SDK)core 0 (the BIOS), from operations your game records into a ring
the bus is driven bycore 1, through the 64-word batch ringcore 0, with one DMA for the whole list
how the game reaches the SDKlinked into the .um2svc syscalls into the BIOS

The Debug Cart fires the whole frame as a single DMA (vbus_list_fire) and returns at once, because there the core that drives the bus also builds the next frame. With 64-word batches it would wait for each batch and build and draw in series (measured: 40 fps in the BIOS menu with a 20 ms list). On the UVMC2 core 1 has nothing else to do, so the batches cost nothing and save the 98 KB that a whole-frame buffer takes.


10. Reproduce it

Nothing on this page needs a console to check, except the measurements on real games. The charts in sections 3 to 5 come from a small program that links the SDK's real emitter (uvm2_draw.c) and the real Rust beam model against a simulated bus, emits frames, and decodes the list the executor would replay:

Build it inside the starter kit:

cd sdk/uvm2-sdk
sh tools/build_host_tools.sh            # builds the Rust beam model and the SDK's own tools
cc -O2 -w -DUVM2_HOST -DUVM2_BENCH_NO_CORE1 -DUVM2_SUBUNITS -DUVM2_HZ=0 \
   -DUVM2_CMD_CAPACITY=65536u -I. -o beam_probe beam_probe.c uvm2_draw.c uvm2_hud.c \
   uvm2_text.c tools/uvm2_host_stubs.c \
   ../vectrex-draw/cabi/target/release/libvectrex_draw_cabi.a
./beam_probe > beam.json

-DUVM2_HZ=0 matters: with the 50 Hz lock on, every list is padded to 30,000 cycles and you would be measuring the padding.

To measure your own game, the SDK ships the same kind of tools: uvm2_list_count breaks a dumped frame down by VIA register, uvm2_anatomy decodes a chain command by command, beam_sim.py replays a list against an ideal beam, and on a console the HUD (buttons 1+4 for two seconds) or stats.py over SWD read the counters live. See Debugging and Host tools.