█▓▒░C64 READY.

Docs / Performance

Performance

This is a cycle-accurate C64: every master cycle it clocks the VIC-II, the 6510, two CIAs, the SID shadow voices and (when attached) the 1541's 6502. That per-cycle fidelity, not demo content, is what sets the cost, so the performance story is mostly about keeping the always-on per-cycle machinery cheap and allocation-free. Presentation switches live in src/switches.js; the rendering pipeline is summarised in the master overview ("Rendering pipeline") and detailed in the VIC-II §14.

1. Cost model

One PAL frame is 19 656 master cycles and must finish inside ~20 ms (50 Hz) to hold realtime. Sampling ten full demo runs (2 600 CPU-seconds of real workload) puts the budget here:

VIC-II 6510 orchestrator SID shadow 1541 CIA
65% 13% 9% 4% 4% 3%

Structurally, that breaks down as:

2. Standing optimizations

The VIC optimizations are permanent, with automatic live and diagnostic paths. WebGL and CRT presentation switches remain in src/switches.js: use ?NAME=0 in the browser or NAME=0 in Node harnesses.

The true PAL drive clock ratio and IEC read-side propagation are permanent parts of the hardware model. Disk writes respect each mounted image's write-protect state. See the 1541 drive and machine orchestrator. Mechanical timing experiments remain compile-time constants in drive1541.js; diagnostic console controls are documented in TESTING.

Assembly64 requests

Catalog work stays outside the emulation loop. Each tab permits three concurrent requests, spaces request starts by 250 ms, shares identical in-flight JSON calls, and honors Retry-After without automatic retries. Search pages are cached for 30 seconds; presets, categories, metadata and file lists for five minutes. The in-memory cache is bounded to 64 responses and 4 MiB. Binaries are streamed with size limits and are never cached by the API client or service worker.

D71 and D81 images use the virtual drive with TDE off. Their filesystem work runs on file/channel operations, without adding per-cycle drive emulation.

3. Throughput & footprint

The figures below are approximate and machine-dependent; treat them as orders of magnitude, not benchmarks.

4. Platform & mobile

Desktop V8 has generous heaps and an optimiser that erases most short-lived allocation, so garbage collection is a non-issue there. Phones are less forgiving:

5. Requirements

The SID audio ring is a SharedArrayBuffer, so the page must be cross-origin isolated. That requires two response headers on the document (and its assets):

Cross-Origin-Opener-Policy: same-origin
Cross-Origin-Embedder-Policy: require-corp

Without them SharedArrayBuffer is unavailable and the machine constructor throws. The Vite dev and preview servers set these headers automatically; a production host must send the same. See the machine orchestrator §7.

6. Measurement Harnesses

Never trust a performance comparison across a different thermal window. On an M1-class laptop, a heavy build, test run, or screenshot pass can change the next number enough to make a neutral change look like a large regression. Run a same-thermal A/B, back-to-back:

  1. Measure the candidate.
  2. Stash only the files under test.
  3. Measure the baseline.
  4. Pop the stash and measure the candidate again.

Two traps are worth naming, because both produce confident nonsense:

The desktop V8 harness is tools/perf-workloads.mjs:

node --expose-gc tools/perf-workloads.mjs <workload> <mode> [frames]

The demo files the workloads load come from test/external-assets.json (orbit-untold-prg, raster-time-demo, the c64stuff collection); edit the paths there or set the per-entry environment variable.

Workloads are idle, orbit, rastertime, comaload, and comarun; modes are time, prof, alloc, allocsites, mem, and all. time is useful for throughput, but V8 can escape-analyze short-lived allocations away. Use allocsites to see allocation a non-EA engine, such as JavaScriptCore on iOS, would still pay for.

Compact render history, sprite interval scheduling and immutable graphics fetches are permanent parts of the renderer. Tracing and diagnostic history access retain dense recording; live rendering retains per-cycle sprite dispatch. Cartridge and tracing lines use the RAM-reading graphics path. Both graphics sources use the combined color/foreground decoder. Active fetches record packed samples once; live/deferred rendering and corrections consume them without rereading graphics RAM. RAM/DMA and fetch-configuration writes no longer force those lines to render early. Collision observers retain catch-up. The disk-trap workload explicitly disables true-drive mode and rejects a run whose pending LOAD/RUN never completes.

An AC-powered candidate/baseline/candidate check of compact history, sprite intervals and the interior pixel-copy path measured these median frame times (Node v25.6.0, 9 batches of 300 frames each):

Workload Baseline Candidate, first / repeat
BASIC READY 6.135 ms 5.947 / 5.868 ms
Orbit Untold 10.218 ms 9.715 / 9.662 ms
Raster Time 7.012 ms 6.774 / 6.767 ms

These are modest throughput gains, about 3–5%, for the combined change, not isolated attribution to each optimization. Desktop WebKit idle measured 5.051 ms baseline against 4.996 / 4.999 ms candidate, effectively neutral. These figures do not establish a phone performance gain.

Next Round's first 6,000 PAL frames after RUN (true drive, SID 8580, loading excluded) were compared on frozen common sources with one change per candidate/baseline/candidate triplet. All 39 timed runs matched pixel and sampled-state hashes across Node 25.6.0 and desktop WebKit. AC power was used with low-power mode off; an inconsistent fetch triplet under background CPU load was repeated. Approximate reductions in mean frame time were:

Change and baseline Node/V8 Reduced-JIT WebKit
Sparse history vs dense history 2.1–3.0% faster 1.5–2.1% faster
Sprite intervals added to sparse history 5.3–7.3% faster 2.7–4.0% faster
Contiguous sprite copies, RAM-based renderer 3.7–4.2% faster 3.3–3.7% faster

These incremental gains are not additive. WebKit used JSC_useFTLJIT=false, verified in its option dump, with no extra warm-up beyond boot/loading. This limits the JIT tiers available; it does not emulate phone hardware or measure input latency. Input sampling, frame presentation scheduling and emulated cycle order are unchanged by contiguous copying.

V8 allocation sampling across all configurations measured 0.17–0.23 KiB of VIC allocation per frame, including runs with escape analysis disabled. There was no sustained allocation or deoptimization problem attributable to contiguous copying. Occasional missing-feedback bailouts occur as new sprite branches are reached during the demo; these are not JSC allocation or deoptimization results.

The combined package (sparse history, sprite intervals, contiguous copies and fetch-fed graphics using the combined decoder) was also measured together against dense history, per-cycle sprites and RAM-based graphics on the same source. Next Round used the same 6,000-frame window and candidate/baseline/candidate procedure:

Engine Reference mean Package first / repeat Mean frame-time reduction
Node/V8 10.375 ms 9.133 / 9.256 ms 10.8–12.0%
Reduced-JIT desktop WebKit 8.923 ms 8.197 / 8.090 ms 8.1–9.3%

All six runs matched every pixel and sampled-state hash. These package gains are measured directly, not added from the individual comparisons. WebKit's FTL tier was disabled in all three runs. The measurements exclude host input, presentation and audio synthesis, and do not establish actual phone latency. Fetch-fed graphics is permanent; tracing and cartridges retain the RAM-reading path.

The Safari/JSC harness is tools/jsc-perf.mjs plus tools/jsc-perf-harness.html:

npx playwright install webkit          # one-time
node tools/jsc-perf.mjs "label"

It starts a throwaway static server with the COOP/COEP headers required by the SID SharedArrayBuffer, launches Playwright WebKit, boots to READY, and times 9×300-frame idle batches through the real unbundled modules. Treat desktop WebKit time as only part of the story: allocation reductions can be neutral on a desktop core with an ample nursery and still remove visible jank on a phone.