Conversation
The kernel demand-filled any user page a thread touched first, foreign guests included. A Linux guest's PROT_NONE reservation read as zeros, and a pthread's guard page, which musl leaves as the unopened bottom of a PROT_NONE stack reservation, took a fresh page and guarded nothing. Every page a guest is meant to have is already mapped by its supervisor with MkPeerMap, so a not-present fault in a foreign guest is now refused: the fault path ends the thread and the supervisor is told, as before. A peer protection of PROT_NONE now maps the page present without the user bit, so the frame keeps its bytes for a later protection that allows access while every guest access faults. perms_of moves to peer_protect beside the bit it reads, and peer_map takes it from there. Ported from #586 onto this branch; without it this branch lands the lane with the hole still open.
PROT_NONE had no way to reach the kernel. protect_span sent only write and exec bits, so mprotect(PROT_NONE) on a backed span asked for prot 0 and the kernel mapped it READ|USER: the guest kept reading what it had just given up. A PROT_NONE mmap was worse, reserving the span and leaving the kernel to demand-fill it on first touch, which is what let a musl guard page take a real frame. A region now records access alongside write and exec, and peer_prot turns the three into the bits a peer call carries, PEER_PROT_NONE included. protect_span sends that and records it with set_prot, so the region list fork maps a child from no longer holds the old protection. fork_copy maps each piece with the protection its span has now, so a span the parent closed comes back closed in the child. A reservation keeps nothing mapped, as before, but is marked without access: a touch before a commit faults, which is what Linux does. Ported from #586 onto this branch, alongside the kernel half.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
PROT_NONE mappings corrupt paging accounting when they are remapped or removed.
Review effort: Balanced
Findings: 1
Open (1)
What changed in this PR
Ports memory-confinement protections to the Linux integration branch.
Changes:
- Refuses demand paging for foreign guests.
- Adds end-to-end
PROT_NONEsupport. - Persists protections across
mprotectandfork.
| File | Description |
|---|---|
userland/libc/src/peer.rs |
Adds the peer PROT_NONE ABI bit. |
userland/capsule_linux/src/linux/guest/region.rs |
Records and translates access permissions. |
userland/capsule_linux/src/linux/guest/region_prot.rs |
Updates stored region protections. |
userland/capsule_linux/src/linux/guest/mod.rs |
Registers and exports protection helpers. |
userland/capsule_linux/src/linux/guest/mem_reserve.rs |
Marks reservations inaccessible. |
userland/capsule_linux/src/linux/call/spawn/fork_copy.rs |
Preserves protection during fork. |
userland/capsule_linux/src/linux/call/mem/prot.rs |
Shares protection-mask handling. |
userland/capsule_linux/src/linux/call/mem/prot_span.rs |
Applies and records protection changes. |
src/process/foreign/peer_protect.rs |
Implements kernel-side PROT_NONE. |
src/process/foreign/peer_map.rs |
Reuses centralized permission conversion. |
src/process/foreign/peer_guard.rs |
Defines the kernel ABI bit. |
src/memory/paging/manager/faults/mod.rs |
Registers demand-refusal logic. |
src/memory/paging/manager/faults/demand.rs |
Applies refusal before demand allocation. |
src/memory/paging/manager/faults/demand_refuse.rs |
Refuses unauthorized guest faults. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| * at teardown, and absent for every access the guest makes. | ||
| */ | ||
| if prot & PROT_NONE != 0 { | ||
| return PagePermissions::READ; |
|
Reviewed at A note on what this review is worth: I wrote this branch, so the usual value of a second reader is missing. What follows is an adversarial pass over my own work, and it did turn up something real — but "the author re-read it" is not the same as review, and the two items below marked for a second opinion are exactly the ones I am least able to judge. Verdict: Request changes — on my own branch. The title claims a closed span stays closed. It does, through Critical1. Reachable in three calls on this branch as it stands:
Step 3 in detail. // Writable while the bytes go in; the old protection after.
if guest.map(at, span, true, false) < 0 { ... } // destination mapped writable, accessible
...
if guest.write(at, &bytes) < bytes.len() as i64 { ... }
if !write {
let _ = super::prot::mprotect(guest, at, span, PROT_READ);
}The destination is mapped writable and accessible, the bytes are copied in — This is #586's
#586 fixes it by passing the whole So the scope of this PR is wrong rather than its content. Either port Important2. I changed
The effect on this branch is that an The general lesson for the next port: Minor
For a second opinion — the two I cannot judge
Verified correct
What is still unprovenNothing here has been booted — no signing keys in the worktree, so a release build stops at |
|
Superseded. This work is integrated into the 0.9.2 release and ships in the current tree. Closing as part of the 0.9.2 consolidation. |

Two commits on top of
linux/syncat950541d8a, porting #586's memory confinement onto the branch that targetsmain.Why this, and why now
#593 is the branch pointed at
main, and it carries none of #586's memory work: nodemand_refuse.rs, noPROT_NONEbit inpeer_guard.rs, and afork_copythat maps each piece with write and exec only. Landing it as written ships the guard-page hole tomainand leaves #586 stranded on a base 127 commits behind, so whoever ports it afterwards resolvesfork_copy.rsagainst a third variant.Doing it here instead means the merge order stops mattering.
What was actually open on this branch
Scoping it turned up three holes rather than the one I expected, because
Regionhas diverged between the two branches —linux/syncgrewkept, #586 grewaccess.1. A PROT_NONE mmap was readable and writable.
map_anon::anonymoussendsprot == 0toGuest::reserve, which maps nothing, and the comment said what followed: "reserve it and let the first access fault a page in". The kernel demand-filled a zeroed frame on first touch, so a reservation read back as zeros and a musl pthread's guard page — the unopened bottom of a PROT_NONE stack reservation — took a real page and guarded nothing.2.
mprotect(PROT_NONE)on a backed span left it readable.protect_spanbuilt its bits fromPROT_WRITEandPROT_EXEConly, so PROT_NONE arrived as prot0, and the kernel'sperms_of(0)returnedREAD | USER. The guest kept reading a span it had just given up. This one is independent of the fork path and was not in my review of #593; I found it while scoping the port.3. Fork re-opened closed spans, and nothing recorded the protection anyway.
prot_ofdropped any notion of PROT_NONE, and the mprotect loop never wrote the new protection back to the region list — which is the list fork maps a child from.The change
Kernel (
4a50be4f3):demand_refuse.rsrefuses a not-present fault for a foreign guest. Every page a guest is meant to have is already mapped by its supervisor withMkPeerMap, so a fault elsewhere is a page nobody gave it. The fault path ends the thread and the supervisor is told, as before.PROT_NONEas a peer protection bit.perms_ofmaps it toREADwithoutUSER: present for the kernel, which copies it at fork and frees it at teardown, and absent for every access the guest makes.perms_ofmoves frompeer_map.rstopeer_protect.rsbeside the bit it reads.Userland (
f5d969af8):Regionrecordsaccessalongsidewriteandexec, andpeer_protturns the three into the bits a peer call carries.protect_spansends that and records it through a newset_prot, so the region list no longer holds the old protection.fork_copymaps each piece withspan.peer_prot().demand_refuse.rs,region_prot.rsand theperms_ofPROT_NONE arm are #586's, carried over unchanged. The rest is the same idea fitted to this branch'sRegion.What is verified, and what is not
Compile-checked clean, exit 0 in all four:
nonos_kernelagainstx86_64-nonos.jsoncapsule_linuxagainstx86_64-nonos-user.json, withRUSTFLAGS=-D warningscapsule_linux_proofsnonos_libcI also traced the whole chain and checked the bits agree across the boundary, since these are hand-synced in two places and have drifted here before:
It has not been booted. This was prepared in a worktree with no signing keys, so a release build stops at
build.rs:349and nothing can be run. That is the gap worth closing before this merges, and #586 already has the guests for it:guardpage(a musl pthread recursing into its guard page, which must end on SIGSEGV with status 139) andprotfork(a 2 MiB mapping forked read-write and forked closed, where the child's read must fault). Neither is on this branch. Running those two against this is the proof I could not produce.Worth a second opinion
Guest::reservepreviously marked the spanwrite: true; it is nowwrite: false, access: false. I could find nothing that readswriteon an unbacked region —fork_copyskips unbacked, andmprotectuses the prot it was passed — but that is an argument from absence.is_foreignnow runs on every not-present user fault, not only for guests. The kernel-half check returns before it, so a kernel-heap fault cannot re-enter the lock whileregistry::insertholds the write side across aVec::push. Worth a look from someone who knows that path better.escapeguest. Only the confinement part is here.