IP28: Indigo2 IMPACT R10000, with the JIT fast path on tcache - #150
Merged
Merged
Conversation
src/mgras.rs answered the board-ID probe and nothing else. This is a model of the Indigo2 IMPACT (MGRAS) graphics board that the PROM and the IRIX X server draw on, as a module in src/mgras/: - dcb.rs: the display control bus devices -- the VC3 timing and cursor, the XMAP, the colour maps and the gamma DAC. - raster.rs: the raster engine -- lines (stippled too), rectangles and block fills, pixel logic ops, RGB and colour-index pixels, the overlay planes, window IDs -- plus its tests. - mod.rs: the GIO interface, the command FIFO and geometry engine port, host DMA for pixel transfers, the three interrupt lines (FIFO, general, vertical retrace), and scanout into a finished frame. The board presents like GR2 does: it implements `GfxDisplay` and hands the renderer a prebuilt frame (`Rex3Screen::prebuilt`), keeping `rgba` current so CI screenshots work headless. Register names follow OpenBSD's impact(4) driver, or say what the register does. `[impact]` accepts a board in the graphics slot only; a second head is not modelled, and exp0/exp1 are refused rather than silently ignored. Tested by the model's own unit tests here; it boots to the IRIX desktop on the IP28 that the next commit adds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A machine profile, `indigo2_ip28`, with an R10000 CPU model, built only with `--features ip28` (which implies ppmem). Without the feature the profile and the model are refused at startup with a message naming it, and none of their code is compiled: `MachineProfile::ip28()` is a constant false, so the R4400/R5000 machines are unchanged. The R10000 (`CpuModel::R10000`, `R10000ShadowCache`): - A shadow cache, out of the data path: loads and stores go to memory, while tag and data arrays exist to answer CACHE ops and the PROM's diagnostics -- two ways per set, 64-bit tags with TagHi, the set's MRU bit, and the ECC bits Index_Store_Data carries. - Cache ops 5/6/7 take their R10000 meanings (CacheBarrier, Index_Load_Data, Index_Store_Data), and TagLo is a 64-bit register. - An R10000-format CP0 Config, a 64-entry JTLB (the arrays are sized to the largest model; R4400/R5000 still use 48), and 44 virtual address bits. The IP28 board: - The memory controller sizes banks the IP28 way (MEMCFG's size field counts units of the base shift, 256 MB banks, which only the IP28 accepts), and the low-memory alias follows where RAM actually is. - The HPC3 reports board revision 13, so IRIX does not take the machine for an early IP26 baseboard. - Count runs at 97.5 MHz and the MC clock at the CPU clock, the rates IRIX assumes from the PROM's `cpufreq`; the generic ones made UST run at a third of real time. - ppmem is always on. A MEMCFG write that moves no bank leaves the window alone, a real move bumps every generation counter, and cleared ranges are scrubbed to zero-filled memory rather than unmapped: IRIX rewrites MEMCFG during boot while the JIT's compile workers and DMA are running, and a stale reader used to fault. Tracing goes through devlog (`log mips mask cp0`, `log l2c`), with a few IRIS_IP28_* switches for the bring-up paths. `ip28.toml.example` is a starting config. The real IP28 PROM passes POST, including its secondary-cache diagnostic, and IRIX 6.5.22 (64-bit) boots to the 4Dwm desktop on the IMPACT from the previous commit, in about 27 s with jitv2. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The R10000's shadow cache keeps neither line data nor tags on the load/store path: every access goes to memory. Under tcache, jitv2's inline path for it is therefore tcache's window half alone. `JitDcGeometry` gains `tagless`. The shadow cache implements the tcache hooks (window base, inline bitmap, generation window) and, once both windows are published, reports a tagless geometry; without tcache it reports none and every access calls out, as before. For a tagless geometry `emit_inline_mem_guard` skips the L1-D tag match, LRU update and dirty bit and goes straight to tcache's gate, now `emit_tc_mapped` and shared with the tagged path, then to `jit_tc_base + phys` with tcache's byte lanes; a store's generation bump is tcache's own. `phys` is cut to 32 bits first, as the callout does. The R4400 and R5000 emit the same code as before. On the IP28 (IRIX 6.5.22), the same tree built with and without tcache: a workload over ssh (awk, a floating-point loop, tar of /usr/include, an awk array loop) 17.5-18.5 s with, 26.4 s without, with identical output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two changes to the mapped-region test every inline tcache load and store makes: - The bitmap is read from `MipsCore::ppmem_bitmap`, where ppmem publishes it in the same statement as the cache's own copy, instead of through `MipsCore::jit_tc_bitmap`, a pointer to that copy: one load instead of two dependent ones. `jit_tc_bitmap` is gone. - The bit is tested as `bits & (1 << region)` rather than `(bits >> region) & 1`. The code-size measurement already in the comment favoured this form; on a live workload it is faster too. IP28, IRIX 6.5.22, an awk array loop, three runs each: 4.57-4.91 s before, 4.57-4.79 s with the one-load change alone, 3.99-4.02 s with both; the whole workload 17.5-18.5 s -> 16.7-16.9 s. Measured on an aarch64 host; the R4400/R5000 tagged path takes the same gate. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`ps2 mouse <dx> <dy> [buttons]` pushes one mouse packet (buttons: 1 left, 2 right, 4 middle), so a desktop can be driven from the monitor or a script without a window to click in. The status bar's Hz now counts 8254 timer 0/1 interrupts as well. IRIX keeps time with the 8254 unless the IOC's is known broken, in which case it uses CP0 Compare, which was already counted; either way the number is the kernel's tick rate, and the two never both run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #145. The Indigo2 IMPACT R10000 (IP28) and the IMPACT board
model, now with the JIT's fast path built on tcache instead of beside it.
Everything IP28 is behind
--features ip28; without it none of this codeis built and the R4400/R5000 machines are unchanged.
control bus devices, the raster engine, the command FIFO and GE port,
host DMA and the interrupt lines. It implements
GfxDisplayandpresents a prebuilt frame the way GR2 does, so CI screenshots work
headless.
and its CACHE ops, R10000 Config/TagHi/64-bit TagLo, a 64-entry JTLB,
44-bit VA; the IP28's MC sizing (256 MB banks, IP28-only), low-memory
alias, HPC3 board revision, Count/MC clock rates; ppmem always on, with
a remap that no longer leaves the window faulting. No longer touches
the cpu-tests harness, which is what made IP28: the Indigo2 IMPACT R10000 and the IMPACT board model #145's prebuilt stamp check
fail.
shadow cache implements the tcache hooks and reports a tagless
geometry, and the inline path for it is tcache's own window path minus
the L1-D tag match, LRU and dirty bit. IP28: the Indigo2 IMPACT R10000 and the IMPACT board model #145's separate window pointers,
setter and byte-lane table are gone. The fast path needs
--features ip28,tcache; without tcache the IP28 is correct but callsout for every access. Whether
ip28should implytcacheis your call.This one changes your tcache code for every model: the gate reads
MipsCore::ppmem_bitmapdirectly (one load instead of two), and testsbits & (1 << region)instead of(bits >> region) & 1. The secondchange is what matters: see the numbers below.
ps2 mouse; status bar: 8254 clock ticks.Measurements
Host: Apple M4 Pro (8 performance + 4 efficiency cores), macOS, aarch64,
16 KB pages. Build:
--release --features jitv2,ip28,lightning,chd,plus
tcachewhere noted. Machine:indigo2_ip28,cpu = "r10000",banks = [256, 256, 256, 256](IRIX sees 768 MB), Solid IMPACT,headless, no audio,
[jitv2] threads = 4. Guest: IRIX 6.5.22m (64-bit).Workload over ssh: an awk integer loop, an awk floating-point loop,
tarof /usr/include piped to
sum, and a timed awk loop over a 4096-entryarray. Three runs each, same base (95b7ad0):
Output is identical in every run. On this host the mask test is worth
about 15% on that loop; your comment's code-size measurement pointed the
same way. Not measured on x86 or on the R4400/R5000 tagged path, which
takes the same gate.
Testing
cargo test --libwithip28,tcache,jitv2,tcache,jitv2,ip28,jitv2andjitv2: all pass exceptindex_store_tag_discards_the_line_instead_of_writing_it_backin thetwo tcache builds. It fails the same way on main with
tcache; it isour test from cache: Index_Store_Tag must not write the line back #133 and I'll look at it separately.
desktop on IMPACT in about 26 s with and without tcache.
lightning,rex-jit,jitv2,chd,tcache) bootsand runs integer and FP loops with the expected results.
🤖 Generated with Claude Code