All SDK docs

Dual Core, PIO and DMA

How the SDK keeps the beam moving while the CPU works: the core split, the PIO bus program and the DMA feed.

Three mechanisms serve one goal: keep the beam moving while the CPU is doing something else. They are the baseline of the SDK, not options. The runtime refuses to compile without them.


Dual core

core 0   builds a command list into one of two buffers, publishes it
core 1   replays it over the bus, reads the controls, ticks the audio

What it buys, and what it does not

It does not make drawing faster. The replay is paced by the Vectrex's own 1.5 MHz clock: 30,000 bus cycles is a 50 Hz frame by definition, and no amount of CPU shortens a bus cycle.

What it buys is overlap. Single-core, a frame runs strictly in series:

[ game logic ][ build the list ][ replay ][ pace ]

so the game's time is added to the beam's. Split across cores, the logic for frame n+1 runs while the beam draws frame n, and the frame costs max(logic, replay) instead of their sum. One game measured with it switched off: 51 ms of drawing plus 46 ms of logic in series, 97 ms per frame; with it, 40.25 ms. See How the engine works for an interactive model.

Who owns the bus

Core 1, exclusively, from the moment it starts. Two cores driving the same GPIO would put two writers on the VIA with no arbitration.

Every bus contact goes through one place (the frame boundary): the replay, the button and axis reads, and the audio tick. All of them run on core 1. The game still reads buttons and axes through the same calls, which answer from a cache, so nothing changes on the game side.

Two exceptions:

  • PSG writes can happen at any moment, so they go through a small queue that core 1 drains at the frame boundary. Routing depends on which core is executing, not which function was called, so new call sites stay correct.
  • Raw bus reads and writes are refused while core 1 owns the pins.

The handshake

extern volatile uint32_t uvm2_frame_request;  /* frames core 0 has finished building */
extern volatile uint32_t uvm2_frame_done;     /* frames core 1 has finished replaying */
/* the buffer for frame n is n & 1 */

Two plain monotonic counters. A 32-bit load cannot tear, so no lock is needed.


PIO: one write per E period

The bus is driven by a PIO program of about 12 instructions (sdk/vectrex-bus/src/bus_stream.pio). Its loop:

start:
    wait 0 gpio 31          ; ~E low
    wait 1 gpio 31          ; ~E rising = E falling = the latch
    nop [14]                ; phase calibration
present:
    pull noblock            ; next word, or X (the park word) if the FIFO is dry
    out y, 1                ; sentinel bit 0
    jmp !y, no_write
    out pins, 25            ; address + direction + data, inside E low
    jmp start

Every instruction is there for a measured reason:

  • The order of the waits. With them swapped, the address changed in the middle of E high, the half where the VIA decodes. The VIA latched nothing: black screen. This order presents the address during E low, as the bus rule requires.
  • The nop [14]. A fixed offset of about 100 ns, measured with a scope against E, puts the presentation where a known-working path puts it. It is [14] because two sentinel instructions come before the presentation. If you add an instruction to this loop, this number changes.

Running dry is the idle state, not a hazard

If the FIFO empties, a plain out stalls with the pins holding the last word. A VIA address left selected is re-latched on every E fall: harmless for a port register, catastrophic for the timer's high byte, which restarts the ramp 1.5 million times a second.

pull noblock solves it in hardware. With an empty FIFO it loads the park word from X instead ($C000, which decodes to nothing on the Vectrex). A starved stream parks the bus instead of hammering the VIA, with no branch, no DMA and no CPU involvement. So DMA is an optimisation here, not a correctness requirement.

The sentinel bits

out consumes from the low bit, so the payload is shifted up by one and bit 0 says what the word is:

bit0 = 1             a write; payload in bits 1..25
bit0 = 0, bit1 = 0   one period of silence
bit0 = 0, bit1 = 1   PARK for N periods, with N-1 in bits 2..25

Silence matters. The analog gaps (beam-on delay, blank settle) are part of how a stroke is formed. Emitting a write every period instead turned 2,390 writes per frame into 17,314 words and broke stroke chaining. The repeated park replaces about 34 park words per vector (about 8,400 a frame) with one word.

A trap worth knowing: a jmp to a label written after the last instruction assembles without error and points past the program. Empty PIO memory decodes as jmp 0, so the state machine makes one pass and then spins forever, with no warning. Disassemble the .pio; do not trust the parser.


DMA

The DMA feeds the PIO FIFO so it never runs dry. Its job is to remove timing spread: a word that arrives late is presented in the next period, and the DMA keeps words arriving on time.

In dual core, core 1 pushes the list in 64-word batches and the DMA drains them. The large single-core list buffer (12,288 words, 98 KB) is never written in dual core, so dual-core games get UVM2_LIST_MAX=64 by default and those 98 KB stay free.


Turning it off to bisect

make uvm2 UVM2_PIO_STREAM=0            # drive the bus from the core over SIO
make uvm2 UVM2_DUAL_CORE=0             # single core
make uvm2 UVM2_STREAM_INSTALL_ONLY=1   # install the stream but draw over SIO

These exist to split a fault in half, not to ship. Anything measured with them is not comparable to the SDK's reference numbers.