Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion cpu-tests/Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -38,7 +38,7 @@ CFLAGS := $(ARCHFLAGS) -ffreestanding -nostdlib -nostdinc \
-fno-delete-null-pointer-checks -fomit-frame-pointer \
-O1 -g -std=gnu11 \
-Wall -Wextra -Werror -Wno-unused-parameter \
-Iharness
-Iharness $(if $(ONLY),-DCPUTEST_ONLY=$(ONLY))
ASFLAGS := $(ARCHFLAGS) -ffreestanding -nostdlib -Iharness -g -Wa,--fatal-warnings
LDFLAGS := -EB -T harness/link.ld -nostdlib --no-warn-mismatch

Expand Down
26 changes: 23 additions & 3 deletions cpu-tests/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,10 @@ make # -> build/cputest.elf
No root? `make toolchain-local` unpacks the same packages into
`~/.local/opt`; the Makefile finds them automatically.

`make ONLY=group_umode` (any `group_*` from `harness/tests.c`) builds a suite of
`identity` plus that one group - for iterating on a slow target such as an HDL
simulator. Do a clean build when switching between it and the full suite.

## Run

```sh
Expand Down Expand Up @@ -134,10 +138,26 @@ yet. Each is plain architecture rather than a part-specific quirk:
| `mem/load_then_trap` | a syscall, break, trap, overflow or reserved instruction right behind a load: exactly one exception, EPC on it, the load complete |
| `mem/load_then_more` | a CP0 read, a divide, a jump, and nearby loads and stores right behind a load |
| `fpu/trap_behind_a_load` | an FP trap right behind an integer load, while the load may still be waiting on its fill |
| `mem/load_then_load` | a load right behind a load that does not name its register - same line, two lines, the same word at two widths, three in a row, the first load's value as the third's base, the second's used at once, after a store, before a store, an LWL merge, KSEG1 against KSEG0 - with both lines cold, both warm, and each one warm alone |
| `mem/load_then_load_evict` | the second of two loads evicts the first's D-cache line, or finds its own line just written back to memory |
| `mem/load_then_store` | a store right behind a load, in the load's line and another, read back, in the same four cache states |
| `tlb/load_then_mapped_load` | two loads where one or both walk the TLB, and the second's value is used at once |
| `umode/*` (12 tests) | User mode the way IRIX 6 runs it: n32 processes with `Status.UX` = 1 under a 32-bit kernel (KX = 0), so every exception and ERET switches addressing mode. Each test runs a few words at kuseg `0x00400000` under three Status settings - KX = SX = UX = 1, all clear (IRIX 5), and UX alone (IRIX 6) - and records every exception: vector, Cause, EPC, BadVAddr, Context, XContext, EntryHi. Syscalls as the first and second instruction after ERET (libc's `_getuid`), a break, a pending interrupt, load and fetch TLB refills (XTLB vector when UX = 1; fixed up and retried), KSEG0 and misaligned loads (AdEL), and 64-bit operations with UX set |

The first results for them are from the `sgiindy_MiSTer` FPGA core presenting as
an R4600: all pass (2259 checks over 246 tests in its simulator, both with every
load stalling execute and with loads that stall only when they must).
an R4600: all pass (2409 checks over 250 tests in its simulator, both with every
load stalling execute and with loads that stall only when they must, including
behind another load or store). `mem/load_then_load` found that core releasing a
stalled load's execute hold on the completion of the load ahead of it.

The `umode` group came out of IRIX 6.5's installer dying on that core with `init
died (why = 3, what = 0xb)`: init's saved frame showed a TLB miss at the general
exception vector's own address, reported on the `syscall` in `_getuid`. The core
built the vector address with the kernel's 32-bit mode and fetched it under the
process's UX = 1. Before its fix `umode/syscall_second` reproduced init's frame
exactly; after it, 2752 checks pass over 267 tests on the FPGA, all but
`umode/fetch_miss_entry` (the core still takes an instruction-TLB miss on an ERET
target as a nested exception). IRIS passes all 12.

## Writing a test

Expand Down Expand Up @@ -179,7 +199,7 @@ and the CPU each did about it.

```
harness/ startup, exception vectors, CHECK macros, console
tests/ identity alu muldiv mem branch excep cp0 tlb fpu cache mips4
tests/ identity alu muldiv mem branch excep cp0 tlb umode fpu cache mips4
gen/ fpvectors.py — computes the FP expectation tables
run/ run-local.sh matrix.sh run-prom.sh bare.toml boot.toml
docs/ findings gotchas status oracle memory-map toolchain
Expand Down
11 changes: 7 additions & 4 deletions cpu-tests/docs/status.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,20 +16,23 @@
| `identity` | 5 | PRId, FIR, cache geometry, Config.K0, TLB size |
| `alu` | 29 | sign extension across the 32/64-bit boundary, overflow traps, shifts, logic, SLT |
| `muldiv` | 18 | mult/div in both widths, HI/LO, the unspecified cases |
| `mem` | 21 | load/store widths, the whole unaligned family at every offset, alignment faults, KSEG0/KSEG1 |
| `mem` | 24 | load/store widths, the whole unaligned family at every offset, alignment faults, KSEG0/KSEG1, loads and stores right behind a load |
| `branch` | 15 | every conditional, likely-nullification, link registers, delay slots, faults in delay slots |
| `excep` | 17 | traps, reserved instructions, coprocessor usability, EXL/ERET, vector selection |
| `cp0` | 21 | read-only registers, reserved-bit masks, 64-bit access, Count/Compare, LL/SC |
| `tlb` | 10 | entry round-trip over all 48, TLBP, every page size, real translation, V/D bits, ASIDs, refill |
| `tlb` | 11 | entry round-trip over all 48, TLBP, every page size, real translation, V/D bits, ASIDs, refill |
| `umode` | 12 | User mode under KX/SX/UX all set, all clear, and UX alone (IRIX 6 n32): syscalls, interrupts, TLB refill vectors and Context/XContext, address errors, 64-bit ops |
| `fpu` | 89 | see below |
| `cache` | 8 | geometry, tag round-trip, cached/uncached views, I-cache coherency |
| `mips4` | 13 | every MIPS IV addition — computes on R5000, must raise RI on R4400 |
| **total** | **246** | |
| **total** | **262** | |

Six of these (`excep/cp0_unusable_user`, `excep/cp0_usable_cu0`,
`mem/load_then_use`, `mem/load_then_trap`, `mem/load_then_more`,
`fpu/trap_behind_a_load`) were added after the 2026-09-11 run and are not in
its numbers; `identity/config_k0` was also re-enabled.
its numbers; `identity/config_k0` was also re-enabled. `mem/load_then_load`,
`mem/load_then_load_evict`, `mem/load_then_store`, `tlb/load_then_mapped_load`
and the whole `umode` group came later still and are not in them either.

### Inside `fpu`

Expand Down
9 changes: 9 additions & 0 deletions cpu-tests/harness/tests.c
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,7 @@ DECLARE_GROUP(group_branch);
DECLARE_GROUP(group_excep);
DECLARE_GROUP(group_cp0);
DECLARE_GROUP(group_tlb);
DECLARE_GROUP(group_umode);
DECLARE_GROUP(group_fpu);
DECLARE_GROUP(group_fpu_trap);
DECLARE_GROUP(group_fpu_denorm);
Expand All @@ -28,6 +29,12 @@ DECLARE_GROUP(group_cache);
DECLARE_GROUP(group_mips4);
DECLARE_GROUP(group_mips4_fp);

/* `make ONLY=group_umode` builds a suite of identity + that one group, for
* iterating on a slow target (a simulator). */
#ifdef CPUTEST_ONLY
DECLARE_GROUP(CPUTEST_ONLY);
const struct test_group *const all_groups[] = { &group_identity, &CPUTEST_ONLY };
#else
const struct test_group *const all_groups[] = {
&group_identity,
&group_alu,
Expand All @@ -37,6 +44,7 @@ const struct test_group *const all_groups[] = {
&group_excep,
&group_cp0,
&group_tlb,
&group_umode,
&group_fpu,
&group_fpu_trap,
&group_fpu_denorm,
Expand All @@ -49,5 +57,6 @@ const struct test_group *const all_groups[] = {
&group_mips4,
&group_mips4_fp,
};
#endif

const unsigned n_groups = sizeof(all_groups) / sizeof(all_groups[0]);
221 changes: 221 additions & 0 deletions cpu-tests/tests/mem/mem.c
Original file line number Diff line number Diff line change
Expand Up @@ -687,6 +687,224 @@ static void t_load_then_more(void)
}
}

/*
* A load right behind a load, and a store right behind a load, where neither
* names the first load's register. A core whose memory stage reads a
* synchronous cache RAM has the second access's address on that RAM while the
* first one is still being answered, so these are the sequences where it hands
* one access the other's word. Every sequence runs in four cache states: both
* lines cold, both warm, only the first warm, only the second warm - a fill,
* a hit, and each of the two mixed.
*
* t_load_then_load_evict adds the case where the second load evicts the
* first one's line, or finds its own line just written back.
*/
#define LL_WARM(first, second) \
do { \
q[0] = v0; q[1] = (u64)(s64)(s32)(unsigned long)&q[8]; q[2] = 0; \
q[3] = 0; q[4] = v4; q[5] = v5; q[6] = 0; q[7] = 0; q[8] = v8; \
SYNC(); \
dcache_wb_invalidate_range(q, 9 * 8); \
SYNC(); \
if (first) { u64 w = q[0] ^ q[1]; (void)w; } \
if (second) { u64 w = q[4] ^ q[5]; (void)w; } \
} while (0)

static void t_load_then_load(void)
{
volatile u64 *q = (volatile u64 *)(_scratch_start + 1536);
volatile u64 *q1; /* the same lines, uncached */
const u64 v0 = 0x0123456789ABCDEFull;
const u64 v4 = 0x5555AAAA3333CCCCull;
const u64 v5 = 0x0F1E2D3C4B5A6978ull;
const u64 v8 = 0x7766554433221100ull;
u64 r, s, t;
int pass;

q1 = (volatile u64 *)K1_PTR(q);

for (pass = 0; pass < 4; pass++) {
/* Same line, consecutive words. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%2)\n\tld $9, 8(%2)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q) : "$8", "$9");
CHECK_EQ_AT("same line first", pass, r, v0);
CHECK_EQ_AT("same line second", pass, s, (u64)(s64)(s32)(unsigned long)&q[8]);

/* Two lines. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%2)\n\tld $9, 32(%2)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q) : "$8", "$9");
CHECK_EQ_AT("two lines first", pass, r, v0);
CHECK_EQ_AT("two lines second", pass, s, v4);

/* The same word twice, at two widths and offsets. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "lw $8, 4(%2)\n\tlbu $9, 1(%2)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q) : "$8", "$9");
CHECK_EQ_AT("same word lw", pass, r, 0xFFFFFFFF89ABCDEFull);
CHECK_EQ_AT("same word lbu", pass, s, 0x23u);

/* Three in a row, each from its own line. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%3)\n\tld $9, 32(%3)\n\tld $10, 64(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero\n\tdaddu %2, $10, $zero" Z
: "=r"(r), "=r"(s), "=r"(t) : "r"(q) : "$8", "$9", "$10");
CHECK_EQ_AT("three first", pass, r, v0);
CHECK_EQ_AT("three second", pass, s, v4);
CHECK_EQ_AT("three third", pass, t, v8);

/* The first load's value is the base of the third instruction, which
* reads it while the second load is in the memory stage. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 8(%2)\n\tld $9, 32(%2)\n\tld $10, 0($8)\n\t"
"daddu %0, $9, $zero\n\tdaddu %1, $10, $zero" Z
: "=r"(r), "=r"(s) : "r"(q) : "$8", "$9", "$10");
CHECK_EQ_AT("chase behind second", pass, r, v4);
CHECK_EQ_AT("chase value", pass, s, v8);

/* The second load's value used at once: that one has to hold. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%2)\n\tld $9, 32(%2)\n\tdaddu $10, $9, $9\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $10, $zero" Z
: "=r"(r), "=r"(s) : "r"(q) : "$8", "$9", "$10");
CHECK_EQ_AT("second used first", pass, r, v0);
CHECK_EQ_AT("second used", pass, s, v4 + v4);

/* A store, then two loads: the first of them waits out READWAIT. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "sd %2, 16(%3)\n\tld $8, 16(%3)\n\tld $9, 40(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(OPAQUE((u64)v8)), "r"(q)
: "$8", "$9", "memory");
CHECK_EQ_AT("store load load first", pass, r, v8);
CHECK_EQ_AT("store load load second", pass, s, v5);

/* Two loads, a store, and the store read back. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%3)\n\tld $9, 32(%3)\n\tsd $8, 48(%3)\n\tld $10, 48(%3)\n\t"
"daddu %0, $9, $zero\n\tdaddu %1, $10, $zero\n\tdaddu %2, $8, $zero" Z
: "=r"(r), "=r"(s), "=r"(t) : "r"(q) : "$8", "$9", "$10", "memory");
CHECK_EQ_AT("load load store second", pass, r, v4);
CHECK_EQ_AT("load load store reload", pass, s, v0);
CHECK_EQ_AT("load load store first", pass, t, v0);

/* LWL merges with the register's old value, behind a load. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "lui $9, 0x1234\n\tori $9, $9, 0x5678\n\t"
"ld $8, 0(%2)\n\tlwl $9, 33(%2)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q) : "$8", "$9");
CHECK_EQ_AT("lwl first", pass, r, v0);
CHECK_EQ_AT("lwl merged", pass, s, 0x0000000055AAAA78ull);

/* Uncached (KSEG1) and cached (KSEG0) loads of the same memory. */
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%3)\n\tld $9, 32(%2)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q), "r"(q1) : "$8", "$9");
CHECK_EQ_AT("kseg1 then kseg0 first", pass, r, v0);
CHECK_EQ_AT("kseg1 then kseg0 second", pass, s, v4);
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%2)\n\tld $9, 32(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q), "r"(q1) : "$8", "$9");
CHECK_EQ_AT("kseg0 then kseg1 first", pass, r, v0);
CHECK_EQ_AT("kseg0 then kseg1 second", pass, s, v4);
}
}

/*
* The second load evicts, or is evicted by, the first. p16 is 16 KB above q:
* another line of memory at the same index in a direct-mapped D-cache of up to
* 16 KB (on a two-way cache the two simply share a set). Writing it in C
* leaves its line dirty in the cache, so the load of q that follows writes it
* back before filling, and the load of p16 right behind has to find the
* written-back word in memory.
*/
static void t_load_then_load_evict(void)
{
volatile u64 *q = (volatile u64 *)(_scratch_start + 1536);
volatile u64 *p16 = (volatile u64 *)(_scratch_start + 1536 + 16384);
const u64 v0 = 0x0123456789ABCDEFull;
const u64 w0 = 0x6B6B6B6B00000000ull;
u64 r, s;
int pass;

for (pass = 0; pass < 2; pass++) {
q[0] = v0; p16[0] = 0;
SYNC();
dcache_wb_invalidate_range(q, 8);
dcache_wb_invalidate_range(p16, 8);
SYNC();

/* p16's line cached and dirty; load q (write-back, fill), then p16. */
p16[0] = w0 + (u64)pass;
__asm__ __volatile__(A "ld $8, 0(%2)\n\tld $9, 0(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q), "r"(p16) : "$8", "$9");
CHECK_EQ_AT("evict q", pass, r, v0);
CHECK_EQ_AT("evicted p16", pass, s, w0 + (u64)pass);

/* And the other way round, with q's line dirty. */
q[0] = v0 + (u64)pass;
__asm__ __volatile__(A "ld $8, 0(%3)\n\tld $9, 0(%2)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(q), "r"(p16) : "$8", "$9");
CHECK_EQ_AT("evict p16", pass, r, w0 + (u64)pass);
CHECK_EQ_AT("evicted q", pass, s, v0 + (u64)pass);
}
}

/*
* A store right behind a load, not storing or addressing through the loaded
* register, read back afterwards - in the load's own line and in another, in
* the same four cache states as above.
*/
static void t_load_then_store(void)
{
volatile u64 *q = (volatile u64 *)(_scratch_start + 1536);
const u64 v0 = 0x0123456789ABCDEFull;
const u64 v4 = 0x5555AAAA3333CCCCull;
const u64 v5 = 0x0F1E2D3C4B5A6978ull;
const u64 v8 = 0x7766554433221100ull;
const u64 x = 0x3C3C3C3CA5A5A5A5ull;
u64 r, s;
int pass;

for (pass = 0; pass < 4; pass++) {
LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%3)\n\tsd %2, 16(%3)\n\tld $9, 16(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(OPAQUE((u64)x)), "r"(q)
: "$8", "$9", "memory");
CHECK_EQ_AT("store own line value", pass, r, v0);
CHECK_EQ_AT("store own line reload", pass, s, x);
CHECK_EQ_AT("store own line memory", pass, q[2], x);

LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 0(%3)\n\tsw %2, 44(%3)\n\tld $9, 40(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(OPAQUE((u64)x)), "r"(q)
: "$8", "$9", "memory");
CHECK_EQ_AT("store other line value", pass, r, v0);
CHECK_EQ_AT("store other line reload", pass, s, (v5 & 0xFFFFFFFF00000000ull) | (x & 0xFFFFFFFFull));

LL_WARM(pass == 1 || pass == 2, pass == 1 || pass == 3);
__asm__ __volatile__(A "ld $8, 32(%3)\n\tsb %2, 3(%3)\n\tld $9, 0(%3)\n\t"
"daddu %0, $8, $zero\n\tdaddu %1, $9, $zero" Z
: "=r"(r), "=r"(s) : "r"(OPAQUE((u64)x)), "r"(q)
: "$8", "$9", "memory");
CHECK_EQ_AT("sb value", pass, r, v4);
CHECK_EQ_AT("sb reload", pass, s, (v0 & ~0x000000FF00000000ull) | 0x000000A500000000ull);
(void)v8;
}
}
#undef LL_WARM

static const struct test tests[] = {
TEST("mem/load_widths_sign", t_load_widths_and_sign, CPU_ALL),
TEST("mem/store_widths", t_store_widths, CPU_ALL),
Expand All @@ -709,6 +927,9 @@ static const struct test tests[] = {
TEST("mem/load_then_use", t_load_then_use, CPU_ALL),
TEST("mem/load_then_trap", t_load_then_trap, CPU_ALL),
TEST("mem/load_then_more", t_load_then_more, CPU_ALL),
TEST("mem/load_then_load", t_load_then_load, CPU_ALL),
TEST("mem/load_then_load_evict", t_load_then_load_evict, CPU_ALL),
TEST("mem/load_then_store", t_load_then_store, CPU_ALL),
};

const struct test_group group_mem = {
Expand Down
Loading
Loading