Skip to content

IP28: Indigo2 IMPACT R10000, with the JIT fast path on tcache - #150

Merged
techomancer merged 5 commits into
techomancer:mainfrom
atomchild411:ip28-tcache
Sep 30, 2026
Merged

techomancer merged 5 commits into
techomancer:mainfrom
atomchild411:ip28-tcache

Conversation

@atomchild411

Copy link
Copy Markdown
Contributor

Replaces #145. The Indigo2 IMPACT R10000 (IP28) and the IMPACT board
model, now with the JIT's fast path built on tcache instead of beside it.
Everything IP28 is behind --features ip28; without it none of this code
is built and the R4400/R5000 machines are unchanged.

  1. mgras: an IMPACT board model in place of the ID stub. The display
    control bus devices, the raster engine, the command FIFO and GE port,
    host DMA and the interrupt lines. It implements GfxDisplay and
    presents a prebuilt frame the way GR2 does, so CI screenshots work
    headless.
  2. ip28: the IP28 machine and its R10000 CPU. The R10000 shadow cache
    and its CACHE ops, R10000 Config/TagHi/64-bit TagLo, a 64-entry JTLB,
    44-bit VA; the IP28's MC sizing (256 MB banks, IP28-only), low-memory
    alias, HPC3 board revision, Count/MC clock rates; ppmem always on, with
    a remap that no longer leaves the window faulting. No longer touches
    the cpu-tests harness, which is what made IP28: the Indigo2 IMPACT R10000 and the IMPACT board model #145's prebuilt stamp check
    fail.
  3. jitv2: tcache serves the R10000 through the window alone. The
    shadow cache implements the tcache hooks and reports a tagless
    geometry, and the inline path for it is tcache's own window path minus
    the L1-D tag match, LRU and dirty bit. IP28: the Indigo2 IMPACT R10000 and the IMPACT board model #145's separate window pointers,
    setter and byte-lane table are gone. The fast path needs
    --features ip28,tcache; without tcache the IP28 is correct but calls
    out for every access. Whether ip28 should imply tcache is your call.
  4. jitv2: tcache's gate tests the bitmap bit as a mask, in one load.
    This one changes your tcache code for every model: the gate reads
    MipsCore::ppmem_bitmap directly (one load instead of two), and tests
    bits & (1 << region) instead of (bits >> region) & 1. The second
    change is what matters: see the numbers below.
  5. monitor: ps2 mouse; status bar: 8254 clock ticks.

Measurements

Host: Apple M4 Pro (8 performance + 4 efficiency cores), macOS, aarch64,
16 KB pages. Build: --release --features jitv2,ip28,lightning,chd,
plus tcache where noted. Machine: indigo2_ip28, cpu = "r10000",
banks = [256, 256, 256, 256] (IRIX sees 768 MB), Solid IMPACT,
headless, no audio, [jitv2] threads = 4. Guest: IRIX 6.5.22m (64-bit).
Workload over ssh: an awk integer loop, an awk floating-point loop, tar
of /usr/include piped to sum, and a timed awk loop over a 4096-entry
array. Three runs each, same base (95b7ad0):

Build awk array loop whole workload
no tcache (every access calls out) 5.49 s 26.4 s
tcache, gate as on main (commit 3) 4.57-4.91 s 17.5-18.5 s
tcache, one load, shift test 4.57-4.79 s 17.5-18.2 s
tcache, one load, mask test (commit 4) 3.99-4.02 s 16.7-16.9 s
#145's separate direct path, for reference 3.99-4.08 s 16.7-17.0 s

Output is identical in every run. On this host the mask test is worth
about 15% on that loop; your comment's code-size measurement pointed the
same way. Not measured on x86 or on the R4400/R5000 tagged path, which
takes the same gate.

Testing

  • cargo test --lib with ip28,tcache,jitv2, tcache,jitv2,
    ip28,jitv2 and jitv2: all pass except
    index_store_tag_discards_the_line_instead_of_writing_it_back in the
    two tcache builds. It fails the same way on main with tcache; it is
    our test from cache: Index_Store_Tag must not write the line back #133 and I'll look at it separately.
  • IP28: the real PROM passes POST, and IRIX 6.5.22 boots to the 4Dwm
    desktop on IMPACT in about 26 s with and without tcache.
  • Indy: IRIX 6.5.22 (R4400, lightning,rex-jit,jitv2,chd,tcache) boots
    and runs integer and FP loops with the expected results.

🤖 Generated with Claude Code

atomchild411 and others added 5 commits September 29, 2026 22:17
src/mgras.rs answered the board-ID probe and nothing else. This is a
model of the Indigo2 IMPACT (MGRAS) graphics board that the PROM and the
IRIX X server draw on, as a module in src/mgras/:

- dcb.rs: the display control bus devices -- the VC3 timing and cursor,
  the XMAP, the colour maps and the gamma DAC.
- raster.rs: the raster engine -- lines (stippled too), rectangles
  and block fills, pixel logic ops, RGB and colour-index pixels, the
  overlay planes, window IDs -- plus its tests.
- mod.rs: the GIO interface, the command FIFO and geometry engine port,
  host DMA for pixel transfers, the three interrupt lines (FIFO, general,
  vertical retrace), and scanout into a finished frame.

The board presents like GR2 does: it implements `GfxDisplay` and hands
the renderer a prebuilt frame (`Rex3Screen::prebuilt`), keeping `rgba`
current so CI screenshots work headless. Register names follow OpenBSD's
impact(4) driver, or say what the register does.

`[impact]` accepts a board in the graphics slot only; a second head is not
modelled, and exp0/exp1 are refused rather than silently ignored.

Tested by the model's own unit tests here; it boots to the IRIX desktop on
the IP28 that the next commit adds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A machine profile, `indigo2_ip28`, with an R10000 CPU model, built only
with `--features ip28` (which implies ppmem). Without the feature the
profile and the model are refused at startup with a message naming it,
and none of their code is compiled: `MachineProfile::ip28()` is a
constant false, so the R4400/R5000 machines are unchanged.

The R10000 (`CpuModel::R10000`, `R10000ShadowCache`):
- A shadow cache, out of the data path: loads and stores go to memory,
  while tag and data arrays exist to answer CACHE ops and the PROM's
  diagnostics -- two ways per set, 64-bit tags with TagHi, the set's MRU
  bit, and the ECC bits Index_Store_Data carries.
- Cache ops 5/6/7 take their R10000 meanings (CacheBarrier,
  Index_Load_Data, Index_Store_Data), and TagLo is a 64-bit register.
- An R10000-format CP0 Config, a 64-entry JTLB (the arrays are sized to
  the largest model; R4400/R5000 still use 48), and 44 virtual address
  bits.

The IP28 board:
- The memory controller sizes banks the IP28 way (MEMCFG's size field
  counts units of the base shift, 256 MB banks, which only the IP28
  accepts), and the low-memory alias follows where RAM actually is.
- The HPC3 reports board revision 13, so IRIX does not take the machine
  for an early IP26 baseboard.
- Count runs at 97.5 MHz and the MC clock at the CPU clock, the rates
  IRIX assumes from the PROM's `cpufreq`; the generic ones made UST run
  at a third of real time.
- ppmem is always on. A MEMCFG write that moves no bank leaves the window
  alone, a real move bumps every generation counter, and cleared ranges
  are scrubbed to zero-filled memory rather than unmapped: IRIX rewrites
  MEMCFG during boot while the JIT's compile workers and DMA are running,
  and a stale reader used to fault.

Tracing goes through devlog (`log mips mask cp0`, `log l2c`), with a few
IRIS_IP28_* switches for the bring-up paths. `ip28.toml.example` is a
starting config.

The real IP28 PROM passes POST, including its secondary-cache
diagnostic, and IRIX 6.5.22 (64-bit) boots to the 4Dwm desktop on the
IMPACT from the previous commit, in about 27 s with jitv2.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The R10000's shadow cache keeps neither line data nor tags on the
load/store path: every access goes to memory. Under tcache, jitv2's
inline path for it is therefore tcache's window half alone.

`JitDcGeometry` gains `tagless`. The shadow cache implements the tcache
hooks (window base, inline bitmap, generation window) and, once both
windows are published, reports a tagless geometry; without tcache it
reports none and every access calls out, as before. For a tagless
geometry `emit_inline_mem_guard` skips the L1-D tag match, LRU update and
dirty bit and goes straight to tcache's gate, now `emit_tc_mapped` and
shared with the tagged path, then to `jit_tc_base + phys` with tcache's
byte lanes; a store's generation bump is tcache's own. `phys` is cut to
32 bits first, as the callout does. The R4400 and R5000 emit the same
code as before.

On the IP28 (IRIX 6.5.22), the same tree built with and without tcache:
a workload over ssh (awk, a floating-point loop, tar of /usr/include, an
awk array loop) 17.5-18.5 s with, 26.4 s without, with identical output.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two changes to the mapped-region test every inline tcache load and store
makes:

- The bitmap is read from `MipsCore::ppmem_bitmap`, where ppmem
  publishes it in the same statement as the cache's own copy, instead of
  through `MipsCore::jit_tc_bitmap`, a pointer to that copy: one load
  instead of two dependent ones. `jit_tc_bitmap` is gone.
- The bit is tested as `bits & (1 << region)` rather than
  `(bits >> region) & 1`. The code-size measurement already in the
  comment favoured this form; on a live workload it is faster too.

IP28, IRIX 6.5.22, an awk array loop, three runs each: 4.57-4.91 s
before, 4.57-4.79 s with the one-load change alone, 3.99-4.02 s with
both; the whole workload 17.5-18.5 s -> 16.7-16.9 s. Measured on an
aarch64 host; the R4400/R5000 tagged path takes the same gate.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`ps2 mouse <dx> <dy> [buttons]` pushes one mouse packet (buttons: 1 left,
2 right, 4 middle), so a desktop can be driven from the monitor or a
script without a window to click in.

The status bar's Hz now counts 8254 timer 0/1 interrupts as well. IRIX
keeps time with the 8254 unless the IOC's is known broken, in which case
it uses CP0 Compare, which was already counted; either way the number is
the kernel's tick rate, and the two never both run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@techomancer
techomancer merged commit 837b048 into techomancer:main Sep 30, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants