All SDK docs

The Command List

How a frame is recorded as a list of VIA writes, what each command costs, and what happens when the list fills up.

Drawing on the Vectrex means writing VIA registers in time with a 1.5 MHz clock. Doing that one write at a time stalls the CPU on every edge and leaves the beam idle whenever the game is thinking. So the SDK is built around one idea:

Core 0 RECORDS a list of VIA writes. Core 1 REPLAYS it back to back.


The command

Each command is 24 bits, stored in 3 bytes:

bits 23..12   delay: bus cycles to idle AFTER this write (0..4095)
bits 11..8    VIA register select   -> A0-A3
bits  7..0    data byte             -> D0-D7

In code (uvm2_bus.h) a command is built as a 32-bit word and packed on the way into the buffer:

#define UVM2_CMD(reg, data, delay) \
    (((uint32_t)(delay) << 20) | ((uint32_t)(reg) << 16) | (((uint32_t)(data) & 0xFF) << 8))
 
#define UVM2_CMD_PACK(w)      ((w) >> 8)            /* 32 bits -> the 24 that matter */
#define UVM2_CMD_REG_DATA(v)  ((v) & 0xFFF)         /* lands on GP0-11 with no shift */
#define UVM2_CMD_DELAY(v)     ((v) >> 12)

The 12-bit register-and-data field lands directly on GPIO 0–11, so the executor's inner loop does no shifting at all. Storing 3 bytes instead of 4 matters because the image lives in 496 KB: a 64 KB buffer holds 21,845 packed commands against 16,384 unpacked.


Cost

One command = one bus cycle = 667 ns, plus its delay.

A 50 Hz frame is 30,000 bus cycles, and uvm2_exec() returns exactly how many a frame spent. Two numbers worth remembering:

  • A chained stroke costs 6 commands, whatever its length. A long stroke and a short one cost the same number of commands, because the length lives in the timer's value, not in the number of writes. In cycles it costs 32 at minimum, plus the ramp's wait.
  • A stroke that needs a blanked jump first costs 14–21 commands, and 2.4–3× the cycles. Averaged over a real frame, that comes to about 13 commands per segment.

That is why chaining strokes (each starting where the last ended) is the cheapest habit there is on this hardware, and why fewer, longer strokes beat more, shorter ones every time. See Drawing.

The host tool uvm2_list_count breaks a real frame down by VIA register. Most of the delay lives on the register that waits for the ramp to finish. See Host tools.


The frame

uvm2_frame_begin();        /* reset the list and the counters */
  /* ... uvm2_draw_* ...      records commands */
uvm2_frame_end();          /* publish for core 1; pace to UVM2_HZ */

A game normally never calls these directly: v_WaitRecal() does, and vpy_frame_begin() calls that.

Under dual core, uvm2_frame_end() publishes to a double buffer (the buffer for frame n is n & 1) and bumps uvm2_frame_request. Core 1 bumps uvm2_frame_done once it has replayed it. Both are monotonic 32-bit counters, so a torn read is impossible and no lock is needed; core 0 issues a memory barrier before publishing.


When the list fills up

stats.dropped counts commands thrown away because the list was full.

If dropped is not zero, nothing you see on screen is evidence of anything. A silently truncated frame looks like "the drawing is broken" and sends you to debug the wrong file. When it happens:

  • raise UVM2_CMD_CAPACITY (3 bytes per command, per buffer, two buffers under dual core);
  • or move the list to PSRAM with UVM2_CMDS_IN_PSRAM=1, at the cost of determinism;
  • or draw less.

The diagnostics HUD (hold buttons 1 and 4 for two seconds) shows dropped on screen as D. See Debugging.


Reads cannot be recorded

A read needs the data bus turned around mid-cycle, so it cannot go in the list. uvm2_via_read() runs directly, and every read in the SDK happens between frames, while /ZERO holds the beam clamped at the centre. Reading in the middle of a frame disturbs the integrators; that window was measured as the source of most stray bright vectors. See Input.