Skip to content

linux: a guest holds only the memory it asked for - #586

Closed
eKisNonos wants to merge 13 commits into
linux/go-guestsfrom
linux/guest-memory
Closed

eKisNonos wants to merge 13 commits into
linux/go-guestsfrom
linux/guest-memory

Conversation

@eKisNonos

@eKisNonos eKisNonos commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

This makes a Linux guest's memory behave the way Linux says it should. Until now the kernel filled any page a guest touched for the first time, so a PROT_NONE reservation read back as zeros and a pthread's guard page stopped nothing. mprotect(PROT_NONE) on a mapped page left it readable, a fork after mprotect gave the child the old protection, MAP_FIXED kept the pages it landed on, and mlock, msync and mincore weren't served at all.

This branch sits on top of linux/go-guests and should be merged after it. Against that base it is 11 commits.

How it works

The kernel no longer fills a page for a foreign guest on its own. Every page a guest is supposed to have (brk, the initial stack, ELF segments, file maps, anonymous maps, mprotect commits, mremap) was already mapped up front by the personality through MkPeerMap. So demand fill only ever handed a guest pages nobody had given it. Now a fault on such a page takes the normal user fault path: the thread dies, the supervisor gets the death notice, and the personality ends the process with 139. The kernel reports and the personality decides.

I looked at two other ways to do it: keeping a list of committed ranges in the kernel, or a per-guest fill policy. Both need a new syscall and new kernel state, and a range list would be a second copy of what the page tables already say. It would also have put lazily filled guest pages under the 16 MiB demand budget in demand_cap.rs, which a Go heap outgrows. The change I made is one check, in faults/demand_refuse.rs.

PROT_NONE on a page that is already mapped needed a way to say "no access", and the peer protection bits only had write and exec. There is now a PEER_PROT_NONE bit for MkPeerMap and MkPeerProtect; no new syscall. x86_64 has no present-but-no-access user page, so the page stays present with the user bit cleared. Any guest access faults, the bytes stay put for a later mprotect that opens the page again, fork can still copy it (MkPeerCopy doesn't look at the user bit), and teardown frees it like any other page. usercopy does require the user bit, so no native syscall can read or write such a page.

Commits

  • e0482b2 memory: fill no page for a foreign guest on its own
  • 1951525 linux: make PROT_NONE mean no access on pages a guest has
  • 187181d linux: record the protection mprotect sets on a backed span
  • 7232acf linux: replace what a MAP_FIXED mapping lands on
  • 259d2d9 linux: never lay a new mapping over one the guest holds
  • 137901e linux: move the break in whole pages, and give pages back when it drops
  • f60b819 linux: keep a mapping's protection and provenance through mremap
  • 7374720 linux: refuse an unaligned address where Linux refuses it
  • f87ffa0 linux: serve mlock, munlock, mlockall, munlockall, msync and mincore
  • 242bcf5 linux: build the memory proofs as one guest, so the test store loads
  • a42cf1f linux: split the memory files to 75 lines and write comments as blocks

The memproof commit is there because the guest test store has a 16777216 byte load budget. With five separate proof guests next to the ones already enrolled it came to 18039927 bytes and wouldn't load. As one guest it's 16565760 bytes.

What I found in munmap, mremap, brk and mmap

While in there I went through the other memory calls looking for the same kind of bug: pages left mapped, a region list that disagrees with the page tables, or a move that doesn't unmap.

  • MAP_FIXED over an existing mapping kept the old pages and protection and left two overlapping regions. Fixed in 7232acf.
  • A hint address without MAP_FIXED was treated as fixed, the mapping cursor could land on a mapping the guest had placed itself, and MAP_FIXED_NOREPLACE mapped right over things. Fixed in 259d2d9.
  • brk mapped from the unaligned old break, so every call pushed a duplicate region for the page it already held. There's no guest-visible effect beyond fork copying that page twice. I found this by reading the code. Fixed in 137901e.
  • Lowering the break left the pages mapped, so growing it again returned the old bytes instead of zeros, and brk could grow straight over a mapping in the heap area. Both fixed in 137901e.
  • mremap grew or moved a read-only or PROT_NONE mapping as read-write, dropped the mark saying file bytes were never proved (so the moved copy could then be made executable), accepted a span running across mappings with different protection, and left the new span mapped if a move failed halfway. All fixed in f60b819. The last one I found by reading and never triggered live.
  • munmap, mprotect, MAP_FIXED and the mmap offset quietly rounded an unaligned value where Linux says EINVAL. Fixed in 7374720.
  • madvise(MADV_DONTNEED) still answers 0 and keeps the bytes. Not fixed here.

mlock, munlock, mlockall, munlockall, msync and mincore are served now. Every page a guest holds is resident from the moment it's mapped and nothing is ever paged out, so the lock calls only need Linux's argument checks. msync has nothing to write back because every file mapping here is private. mincore answers from the region list.

Testing

Everything below was built and booted on the tree at a42cf1f (q35 under TCG, one vCPU). The kernel builds with 43 warnings, same as the base, and the personality with none. No boot had a triple fault or a reset, none logged a demand fill, and every guest finished with 0 unserved calls.

The proofs are one guest, /bin/memproof, run with the proof's name as its argument. I ran each one on host Linux first as the reference.

  • guardpage: a pthread recursing into its guard page. The process dies on SIGSEGV with exit 139 and never prints the line it would print if it ran past the guard. 11 calls served.
  • protnone: 7 of 7 parts. Reads and writes of a PROT_NONE mmap fault, a page closed with mprotect faults, the closed page below an opened one faults, and bytes survive a close and reopen. 82 calls.
  • protfork: 7 of 7. The child gets the protection the parent has now, for RW to R, RW to NONE, RW to R to RW, and one page closed out of three. 87 calls.
  • touchfork: 4 of 4. Bytes written into an opened part of a reservation reach the child, also after closing and reopening it; a page never opened faults in the child; MAP_FIXED PROT_NONE over a written page reads back zero. 54 calls.
  • memcalls: 32 of 32, covering the placement, brk, alignment and mremap fixes above plus the lock, msync and mincore calls. 131 calls.

The existing guests all still pass: gohello (293 calls), goconc (250), cthreads (87), threadfault (death line and exit 139, 7 calls), gopoll (3 parts in 3434 ms, 371 calls) and cwait (16 parts in 2291 ms, 345 calls).

For the negative controls I removed each change with a small sed, rebuilt, booted, and restored it afterwards. That took four builds, each removal breaking only its own parts:

  • Without the fill refusal and the PROT_NONE bit, guardpage prints "ran 66 KiB below the stack with no fault". protnone fails 4 of 7, protfork fails RW to NONE, and touchfork fails the never-opened page.
  • Without recording the protection and without the MAP_FIXED unmap, protfork fails 3 of 7 and touchfork reads the old byte after MAP_FIXED.
  • Without the placement, brk, mremap and alignment fixes and the new syscall table entries, memcalls fails 29 of 32. The three that still pass are the ones none of those removals touch, and the log shows 33 unserved lines for syscalls 26, 27, 149, 150, 151, 152 and 325.
  • Without the cursor skipping held spans, memcalls fails the 2 parts that depend on it.

The repo gates show nothing new (check_unreachable still lists the known has_children). Every one of the 11 commits passes cargo check for the kernel and the personality on its own, and the head was built in full and booted.

Kernel segments

This changes the kernel, so here are the x86_64 PT_LOAD segments. I built the same tree with the kernel files put back as they are on linux/go-guests (before) and as they are here (after). Two builds of "after" came out identical.

  • Segment 0 (R X, 0xffffffff80000000): filesz and memsz go from 0x1267ca to 0x12688a, so text grows 0xc0 bytes. sha256 before e30403b7ab233692a154df35a2335ca9165ea3aae71b6d2b9fc9d6c8a060d281, after 2be443f02448f1651dd753b69fc33be9bb656fa618ebed48150cec6ec69ad667.
  • Segment 1 (R, 0xffffffff80127000): 0x58fcec0 both before and after, new contents. sha256 before 7b0452a271ab0e09c61196bd3ae730d69e8e48728e5dd1ef508db6619e4fd8f4, after be4d7c923fc3ec7cb70826169468d611e4305d11d19cfc5d8e961b5906ddceab. This segment also holds the embedded capsule set.
  • Segment 2 (R W, 0xffffffff85a24000): filesz 0x6a000, memsz 0x6a12ab8, unchanged. sha256 98cca0db86e6c375bd6a6bce17392ff04c04cddc61d1a8da03dc24bc168fca0e both times.

This needs your sign-off.

Not done here

  • madvise(MADV_DONTNEED) answers 0 and keeps the bytes; Linux hands a private anonymous page back as zeros.
  • mremap with MREMAP_FIXED or MREMAP_DONTUNMAP answers EINVAL without a named line; Linux serves both.
  • A read(2) or write(2) buffer in a PROT_NONE or read-only page is still served, because MkPeerCopy ignores protection; Linux answers EFAULT.
  • msync with MS_INVALIDATE over an mlocked span answers 0 where Linux says EBUSY, since locks aren't tracked.
  • MAP_FIXED_NOREPLACE at address 0 is refused with EPERM like MAP_FIXED. I have no live proof of that, because host Linux running as root maps it, so there's no reference.
  • A guest's own SIGSEGV handler doesn't run for a memory fault; the process just ends.
  • A child that died on a signal shows up in wait4 as exit status 139 instead of WIFSIGNALED, and reaping one prints "kill refused: pid N outlives its process, errno 1".
  • Outside the memory code this touches call/spawn/fork_copy.rs, where fork now maps each span with its protection. It also makes mem public in call/mod.rs, adds one line each to abi/mod.rs and abi/nr.rs for the new abi/nr_mem.rs, and touches guest/mod.rs and serve/table_mem.rs.

The kernel demand-filled any user page a thread touched first, foreign
guests included. A Linux guest's PROT_NONE reservation read as zeros, and
a pthread's guard page, which musl leaves as the unopened bottom of a
PROT_NONE stack reservation, took a fresh page and guarded nothing: a
recursion ran straight through it.

Every page a Linux guest is meant to have is already mapped by its
supervisor with MkPeerMap: brk, the initial stack, the ELF segments, file
maps, anonymous maps and mprotect commits. Demand fill only ever gave a
guest pages nobody had mapped for it. Now a not-present fault in a
foreign guest is refused. The fault path ends the thread with -11, the
kernel posts the death to the supervisor as before, and the personality
ends the process with status 139. The page tables stay the one record
of what a guest holds; nothing new is kept in the kernel.

guardpage, a musl pthread recursing into its guard page, proves it: the
process ends on SIGSEGV with status 139 and never prints the line it
prints when it runs 64 KiB below the stack.
The peer protection bits had write and exec and nothing for no access:
MkPeerMap and MkPeerProtect with neither bit gave a present, readable
user page. So mprotect(PROT_NONE) on a mapped span, and a PROT_NONE file
mapping, left the pages readable, and the region list had no way to say
a backed span was closed.

MkPeerMap and MkPeerProtect now take PEER_PROT_NONE. The page stays
present with the user bit clear: every guest access faults, the kernel
still copies it at fork and frees it at teardown, and the bytes are
there again when a later mprotect opens it, as Linux keeps them. A
region records whether it allows access at all, and one function turns
a region's protection into peer bits for map, commit, mprotect and the
fork copy.

protnone proves it part by part in forked children.
mprotect changed the page protection of a backed span in the kernel but
left the region list with the protection the span was mapped with. Fork
maps the child from that list, so after mprotect RW to R, or RW to
PROT_NONE, the child got the span read-write: a write the parent could
not make succeeded in the child.

Every protection change on a backed span goes through protect_span, and
protect_span now records the new write, exec and access on exactly the
part it changed, splitting the region around it and keeping whether the
bytes were proved. Fork maps the child with what the parent has now.

protfork proves it: RW to R, RW to PROT_NONE, RW to R to RW, and one
page closed in the middle of three, each checked in a forked child.
A MAP_FIXED mmap over pages the guest already had kept them: MkPeerMap
skips a page that is present, so the old bytes and the old protection
stayed, and the region list held the old span and the new one on top of
each other. A program that wrote a page, mapped PROT_NONE over it with
MAP_FIXED and opened it again read its old bytes, where Linux gives
zeroes, and fork copied the doubled span twice.

MAP_FIXED now unmaps the target span just before the new pages go in,
as Linux replaces a mapping. It is done after every refusal the mapping
can meet, so a refused mapping leaves the old one in place.

touchfork proves it with the reservation cases around it: bytes written
into an opened part of a reservation reach a forked child, still do
after the part is closed and opened again in the child, a page never
opened faults in the child, and MAP_FIXED PROT_NONE over a written page
reads zero once opened.
mmap took any nonzero address as fixed, hint or not, and the mapping
cursor took the next address above the last mapping it chose without
looking at what the guest had mapped there itself. Either way a new
mapping could land on an old one: MkPeerMap skips a present page, so the
guest got the old bytes back as its new, zeroed mapping, and the region
list held both. MAP_FIXED_NOREPLACE was taken as a hint and mapped over
whatever was there.

Placement now follows Linux. MAP_FIXED is the exact address;
MAP_FIXED_NOREPLACE is the exact address or EEXIST when anything is
there; any other address is a hint, rounded up to a page and used only
when nothing is there. Otherwise the cursor chooses and skips every span
the guest holds, and moves past a mapping only when it chose it.
brk mapped from the old break itself, which is rarely page aligned, so
every call that grew the break pushed a second region for the page the
break already sat in, and fork copied that page once per call. A lower
break only moved the number: the pages above it stayed mapped, and
growing again handed back the old bytes where Linux gives zeroes. And a
break that grew into a mapping the guest had placed in the heap area
with MAP_FIXED mapped over it.

The break now grows from the page above the old one, is refused, as
Linux refuses it, when that span meets a mapping, and a lower break
unmaps the whole pages above it.
mremap gave the part it grew read-write whatever the mapping was, so a
PROT_NONE or read-only mapping grew a writable tail. A move read the old
bytes and failed with EFAULT on a reservation, which has none; restored
only read-only or read-write on the copy, so a PROT_NONE mapping came
back readable; left the new span mapped when the copy failed; and
dropped the mark that the bytes came from a file nothing proved, so the
moved copy could then be made executable with mprotect, past the check
that refuses it where the file was mapped. It also accepted an old span
running across mappings with different protections.

The grown part and the moved copy are now held like the mapping they
came from: its protection, its backing and its provenance. A reservation
grows or moves as a reservation with no bytes to copy. A move that fails
unmaps what it made. An old span that is not one mapping is refused with
EFAULT, as Linux refuses it. The move goes to a free span found the same
way mmap finds one.
munmap and mprotect rounded an address that was not on a page boundary
down to one, and so acted on bytes below the address the guest named.
mmap did the same for a MAP_FIXED address, and took a file offset that
was not on a page boundary. Linux answers all four with EINVAL.

They now answer EINVAL. Only a length is rounded, and a hint address
without MAP_FIXED, as on Linux.
A guest calling mlock, mlock2, munlock, mlockall, munlockall, msync or
mincore got an unserved line and ENOSYS, where Linux answers each.

In this memory model every page a guest holds is resident from the
commit that maps it until it is unmapped, and nothing is paged out. So
the lock calls have nothing left to do but Linux's argument checks: the
span must be held (ENOMEM), the flags known (EINVAL), and RLIMIT_MEMLOCK
is already reported unlimited. Every file mapping is private, since
MAP_SHARED of a file is refused, so msync has nothing to write back and
answers its checks alone, ENOMEM for a span with a hole included.
mincore answers from the region list, one byte per page: 1 for a backed
page, whatever its protection, and 0 for a reservation page, which has
no frame.

memcalls proves these and the placement, brk, alignment and mremap
changes before it, one line per part.
With the gopoll, gopreempt and cwait guests enrolled beside the five
memory proof guests, the guest-test store no longer loaded:
nonos-store-pack counted 18039927 payload bytes against the vfs budget
of 16777216 and refused it, so no guest could boot. Each guest in the
store is its binary plus its certificate, manifest and a 327 KiB proof
trailer, and the five memory guests came to 1907576 bytes of it, 1683520
of them proofs.

The five proofs are now one program, memproof, whose first argument names
the proof: guardpage, protnone, protfork, touchfork or memcalls. Each
proof keeps its own file and is built with its main renamed, so it runs
exactly as it did as a program of its own. memcalls maps its own file for
the provenance part, which is now /bin/memproof. The ids 4982 to 4989 are
free again.
Nine files this branch touches ran past 75 lines, from 76 for map.rs to
224 for memcalls.c, and the comments added in the kernel, personality
and guest files were written with // where the rule is /* */.

Each long file is split along a line it already had: the three refusals
of a demand fill go to faults/demand_refuse.rs; perms_of moves into
peer_protect.rs; reserve and map_span leave mem_map.rs; the mprotect
walk, the mremap one-mapping check, the mmap kind and the cursor search
each get a file. The memory proofs share one file for printing, running
a part in a child and counting, and memcalls is split by call family.
call/mod.rs is back to its own export line and makes mem public, so
the lock, msync and mincore calls are named call::mem::... in the
table. No behaviour changes.
@eKisNonos
eKisNonos changed the base branch from main to linux/go-guests September 29, 2026 09:14
Fork mapped each span into the child with one MkPeerMap over the whole
span. The kernel refuses a peer call longer than 1 MiB, so forking a
guest that held a larger span, as a Go heap is, failed. Another branch
fixes that by mapping in pieces, but maps each piece with write and exec
only, so a span the parent closed with PROT_NONE comes back open in the
child.

Fork now maps and copies each span a megabyte at a time, and gives every
piece the protection the span has now, PROT_NONE included. The proof
helper also counts a fork that fails as a failed part; it used to wait
on nothing and report a pass. protfork gains two parts on a 2 MiB
mapping: read-write, with the child reading the last page, and closed
with PROT_NONE, where the child's read must fault.
A guest that turns the isolation proofs around: instead of showing a
mapping behaves, it attacks its own confinement through the syscalls it
has. It maps at the kernel half, wraps a span, asks for a write-and-
execute page, adds execute to a file it never proved, calls a number the
personality does not serve, and steps into an unmapped hole. Each part
passes when the machine refuses.

It is a memproof part, so it shares the one guest binary and its store
trailer. On NONOS all seven are refused. On native Linux two are allowed,
a write-and-execute mapping and adding execute to an unproven file, which
are the two the personality enforces and stock Linux does not.
@senseix21

Copy link
Copy Markdown
Collaborator

Reviewed at 9ac41dbda (not a draft, base linux/go-guests, 56 files, +1697/−249, last pushed 2026-09-29). 13 commits of its own from 24cafe103, the same fork point #584 and #585 use, so this is nine commits behind where linux/go-guests now is. Kernel share is small and real: +59/−18 in src/memory, +35/−18 in src/process.

Verdict: Comment. I did not find a defect in this change. The one thing that needs a decision is not in this PR's diff — it is that #585 edits the same function the other way, and the wrong resolution reopens the hole this PR closes.

This one found a real hole by running the attacker

e0482b2ef is the commit worth reading first, and its evidence is the kind that cannot be argued with:

a pthread's guard page, which musl leaves as the unopened bottom of a PROT_NONE stack reservation, took a fresh page and guarded nothing: a recursion ran straight through it.

The kernel demand-filled any user page a thread touched first, guests included. So a Linux guest's PROT_NONE reservation read back as zeros, and musl's thread guard page — the thing whose entire job is to fault — was quietly backed by a fresh frame the first time a recursion reached it. The guardpage guest proves both directions: the process now ends on SIGSEGV with status 139, and never prints the line it prints when it runs 64 KiB below the stack.

The fix is also the right shape. Rather than tracking guest reservations in the kernel, demand_refuse::refused simply declines to fill anything for a foreign pid, on the stated grounds that "every page a Linux guest is meant to have is already mapped by its supervisor with MkPeerMap" — so "the page tables stay the one record of what a guest holds; nothing new is kept in the kernel". I checked that claim where it is most likely to be false, the main stack: guest/layout.rs:37 maps a fixed STACK_SIZE = 1 << 20 and call/limits_table.rs:29 reports RLIMIT_STACK as the same 1 << 20. The stack does not grow on fault here the way Linux's does, but it is fully mapped and the guest can read back exactly how much it has. Nothing is silently smaller than advertised.

And then 9ac41dbda adds escape, which is an adversarial guest rather than a functional one: reach_kernel, wrap_span, wx_map, exec_escalate, forged_call, past_end, with "Each part passes when the machine refuses; one boot names any that got through." A test that fails when confinement works is the one worth checking in.

Important

1. fork_copy.rs is edited incompatibly by this PR and by #585, and only one version is safe.

a1112700a's message names the conflict without naming the branch:

Another branch fixes that by mapping in pieces, but maps each piece with write and exec only, so a span the parent closed with PROT_NONE comes back open in the child.

That branch is #585, and the accusation is correct. Both PRs fix the same bug — MkPeerMap refuses a span over 1 MiB, so forking a Go heap failed — and both rewrite copy_spans in call/spawn/fork_copy.rs:

  • here, mk_peer_map(child, span.at + done, take, span.peer_prot()), with the comment "Each piece gets the protection the span has now, PROT_NONE included".
  • on linux: process lifecycle and signals as Linux has them #585 at 15067eb12, mk_peer_map(child, span.at + done, take, prot_of(&span)), where prot_of sets only PEER_PROT_WRITE and PEER_PROT_EXEC. There is no PROT_NONE on that branch at all — PROT_NONE is introduced by this PR, at peer_guard.rs:44.

So a fork on #585 hands the child read access to every span the parent had closed. If #585 lands after this PR and the conflict resolves toward its side, the guard-page fix survives but the fork path reopens it.

There is a second, quieter half to the same conflict. The two versions disagree about what happens to an unbacked reservation in the child:

#585's comment is describing demand fill — the exact behaviour e0482b2ef removes. Merged together without care, that path produces a child whose first touch of a reservation kills the thread, in code written on the assumption that it would be filled.

Nothing to change in this PR. But these two cannot be merged by resolving a textual conflict, and this is the side to keep.

2. Nine failing lanes, six causes, all inherited.

Checked individually rather than assumed, and identical to the rest of the stack: crypto-proofs (the crypto_proofs mirror includes against_pedersen.rs but not against_root.rs, its only non-test caller), proof-crates (capsule_linux_proofs) (mutation_tests.rs:35, 0xDEB5_EEDu64 under unusual_byte_groupings), runnable-proofs (usage_tests.rs:98 and :111, from #567's help.rs restructure), evidence-manifest (stale EVIDENCE.json), abi-contracts (the assumption register), and build/production-build/benchmark all from tcb-budget: 134013 against baseline 132664, delta +1349. Above the trunk's +1237 that is +112, nearly all of it #582's first 36 commits; this PR's own ring-0 addition is the ~50 lines of demand_refuse.rs and perms_of.

hygiene here is the single inherited capsule_wallet_nonos/.../fact.rs:18 false positive — unlike #585, this PR adds none of its own.

Runs 36570749917 and 36570750018 are cancelled and account for all thirteen of the 4–5 second failures in gh pr checks. The live set is the nine above, from 36570756002 and 36570756427.

Minor

  1. Same stale base as linux: a guest's sockets are the family's own, on 127.0.0.0/8 #584 and linux: process lifecycle and signals as Linux has them #585 — forked at 24cafe103, nine behind linux/go-guests. Less consequential here than on linux: a guest's sockets are the family's own, on 127.0.0.0/8 #584, since nothing in those nine commits fixes something this PR lists as outstanding, but item 1 is a reason to do the merge early rather than at retarget time.
  2. PROT_NONE is added as a prot bit at peer_guard.rs:44 and is now accepted by MkPeerMap as well as MkPeerProtect, since peer_map.rs switched to peer_protect::perms_of. That is almost certainly intended — a supervisor wants to map a reservation closed — but abi/syscalls.toml documents the prot argument only as u64, so the new bit is not written down anywhere a reader of the ABI would find it.

Questions

  1. perms_of maps PROT_NONE to PagePermissions::READ with the USER bit clear, which makes every guest access fault while the frame keeps its bytes. That is a neat way to get Linux's semantics out of the hardware. What distinguishes such a page from a genuine kernel page for anything that walks a guest's tables — teardown, the fork copy, or an accounting pass that counts a guest's resident pages? The frame is reachable from ring 0 by design, so the question is whether anything decides "this is kernel memory" from the USER bit alone.
  2. demand_refuse::refused calls is_foreign(pid) on every not-present user fault in the system, not only for guests, and that is an RwLock read plus a linear scan of a Vec (foreign/registry.rs:40). I could not measure a cost — see below — and the ordering makes it safe. Is the table expected to stay small enough that the scan never matters, or is a per-PCB flag the eventual home for this?

Verified correct

  • The order of the checks in refused is load-bearing for more than its stated reason. !layout::in_user_space(addr) returns before is_foreign does anything, so a kernel-half fault never reaches FOREIGN.read(). That matters because registry::insert takes FOREIGN.write() and then Vec::push, which can allocate: if a kernel-heap fault could re-enter is_foreign, a spin RwLock would deadlock against its own writer. It cannot, because the kernel-half test comes first. I also checked clear, which calls guests_of (taking and releasing a read) before taking its write rather than nesting the two.
  • PROT_NONE pages really can be filled by the supervisor. The claim "the kernel copies into a page whatever its protection, so the bytes still go in" is what makes the fork path correct for a closed span, and it holds: sys_peer_copy resolves the frame with translate_in_asid and hands the physical address to chunk_copy, consulting neither the USER bit nor write permission (peer_copy.rs:45-51). So a child's closed span is filled and then unreachable, which is what the parent had.
  • No measurable boot cost from the new fault check. userspace_entry_ms across the stack: linux: personality, store install, desktop and kernel fixes #567 2531, linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole #582 2565, linux: a guest's sockets are the family's own, on 127.0.0.0/8 #584 2548, linux: process lifecycle and signals as Linux has them #585 2524, this PR 2508 — the fastest of the five. Single samples under TCG, so this is not a benchmark, but it does rule out the obvious regression from putting a lock on the demand-fault path.
  • The boot is clean. nonos-benchmarks-36570756427-1/boot-log.json: zk_attest_ok: 24, zk_attest_fail: 0, fatal: 0, panic: 0. Worth more here than elsewhere in the stack, because refusing demand fill for guests is exactly the change that would show up as capsules failing to come up if the "already mapped by its supervisor" claim were wrong for any of them.
  • The refactor in a42cf1f43 is a move, not a rewrite. perms_of relocated from peer_map.rs to peer_protect.rs with the PROT_NONE arm added; the existing READ|USER / WRITE / EXECUTE construction is unchanged, and peer_map.rs now imports it rather than keeping a second copy.
  • The refusal is reported, not silent. A refused fault ends the thread with -11, the kernel posts the death to the supervisor through the existing notice path, and the personality turns it into status 139 — so a guest that touches a page it does not hold looks to its parent exactly like a Linux SIGSEGV, rather than hanging or faulting in a loop.
  • boot-smoke, boot-proofs, extraction, kani, lean, verus, attestation, attestation-attack, adversarial, reproducible, supply-chain, trust-chain, trust-ledger, proof-crates-kani, symbol-scan, section-size, proof-coverage and build-x86_64-capsules all pass in the live runs.

CI at this head

Nine real failures, six causes, every one inherited from #567 or #582 and none introduced here. Everything else passes or is cancelled noise from runs 36570749917 and 36570750018.

The work in this PR is done. What is left is a stack decision: item 1 says this version of copy_spans is the one to keep, and #585's is the one that quietly undoes e0482b2ef.

@eKisNonos

Copy link
Copy Markdown
Contributor Author

Superseded. This work is integrated into the 0.9.2 release and ships in the current tree. Closing as part of the 0.9.2 consolidation.

@eKisNonos eKisNonos closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants