Conversation
A Linux program cannot run here, and making one run by putting a Linux personality in the kernel would mean putting Linux in the trusted computing base. This adds the opposite: a generic mechanism that lets an ordinary capsule host a binary the kernel does not understand. A foreign process is created with no capabilities at all. When it makes a syscall this kernel does not recognise, and NONOS numbers are four character tags so nothing legitimate lands there, the caller is parked and the frame is handed to the process that created it. That supervisor decides what the call means, builds the guest's address space through two peer calls, and answers. Six syscalls, one capability, and no knowledge of any other operating system in ring 0. Peer map and peer copy refuse any process the caller did not create, and copy through the guest's frames rather than its mapping, so a read-only code page can be filled without ever being mapped writable and executable at once. A guest dies with its supervisor. The first user is a Linux personality capsule holding the ABI, an ELF loader, and the calls a static binary makes. It is a userland program like any other, signed and capability-bound, and the kernel gains nothing from its presence.
The personality could host a binary that never opened a file, which is almost none of them. This adds the file layer: openat and open, read, write, close, lseek, stat, fstat, newfstatat and getdents64, over the store through the client the rest of the tree uses. Everything is opened under this capsule's own identity rather than the guest's. A hosted process holds no capabilities, so the store would refuse it, and the consequence is the useful one: a guest reaches exactly what this capsule is granted and nothing else. A descriptor now holds an open handle, a read offset, a pending write buffer and, for a directory, the listing taken when it was opened. Writes are held until close and then written as one file, because the store takes whole values and a program that writes a file expects it to appear whole or not at all. Paths are made absolute against a working directory and flattened here, since the store has neither dot nor dot-dot.
A static binary opens and stats before it reaches main, and asks three more questions on the way: whether its output is a terminal, which kernel it is on, and where it is. Answering those with ENOSYS stopped programs that the file layer alone would otherwise have carried. Added pread64, getcwd, access, readlink, ioctl, fcntl and uname. The answers are the true ones rather than the convenient ones. There is no terminal behind any descriptor, so ioctl is ENOTTY, which is what a libc is asking when it probes TCGETS and is what puts it into full buffering. The store holds no symbolic links, so readlink is EINVAL on a path that exists and ENOENT on one that does not. uname says Linux for the system name because that names the ABI this capsule implements, which is the question being asked, and the other fields say NONOS. A directory entry is reported with an unknown type rather than a guessed one, since the listing says what is there and not what each entry is, and a caller that cares will stat the name.
NONOS refuses to map a page writable and executable at once, and that is the right rule. It also means a loader cannot ask for the end state up front: a dynamic linker maps a library writable, applies relocations to it, and only then needs it executable. Without a way to make that second change, mprotect had to be a lie that returned success and did nothing, and the program faulted on its first call into the library. The kernel gains a seventh peer call. MkPeerProtect sets the protection of pages a guest already has, refusing any process the caller did not create and any page that is not mapped, since a caller asking for execute on a range it has not filled is not asking for what it thinks. It reuses the same permission conversion and the same W^X refusal as the mapping call, so the rule cannot be escaped through it. I said in the mechanism PR that the ABI could grow without the kernel changing again. That was too strong, and this is the exception: changing the protection of frames a process already holds is not something the six calls could express. It is still a generic primitive with no knowledge of any foreign format in it. With it, mprotect is real, and mmap can map a file privately: the pages are mapped writable, the bytes are read into them, and the requested protection is set afterwards. A request for write and execute together is refused with EPERM rather than quietly granted as one of the two.
A static binary was all this could host, which rules out most software that exists. A dynamic executable does not start at its own entry: the interpreter named in PT_INTERP is loaded beside it and started instead, and it maps the libraries and jumps to the program when it is done. The loader now reads the interpreter path, fetches it from the store, loads it at its own base, and reports where everything landed. A shared object is biased and an executable is not, decided from the ELF type in the file rather than from a flag a caller could get wrong. The auxiliary vector is the other half. An interpreter finds the program's headers through AT_PHDR, learns where it was itself placed from AT_BASE, and jumps to AT_ENTRY at the end. AT_RANDOM is not optional either: a C runtime takes its stack guard from those sixteen bytes before it runs anything, and they come from the system entropy source rather than a constant, which would make every guest's guard identical. The program itself now comes from this capsule's arguments and is read out of the store, so the personality runs what it is asked to run instead of being a wrapper around one embedded image. That image stays as the fallback, so the mechanism can still be proved on a machine with nothing in the store.
A toolkit is not single threaded, so nothing with a window was ever going to run without this. The kernel gains an eighth peer call. MkForeignThread makes a thread inside a guest: it shares the guest's address space, joins the same supervisor so its own unknown syscalls redirect there too, and carries a TLS base that the context switch loads from the control block, which it must, because a C runtime reads thread-local storage before it runs any program code. The futex needs no kernel support at all. A guest thread that traps is already parked inside its syscall with no reply, so a wait is the absence of a reply and a wake is the reply. The personality compares the word through a peer read, answers EAGAIN if it already changed, and otherwise leaves the caller parked with its address recorded. A wake replies to as many waiters on that address as were asked for. The serve loop therefore takes traps from any thread of the guest and no longer assumes one reply per trap. clone is written and refuses. musl resumes a cloned child at the instruction after its own syscall, so the child's entry is the caller's return address, and the kernel does not pass it yet: the foreign redirect was wired with a literal zero for the instruction pointer. Passing it means touching the SYSCALL entry assembly, which is not something to do untested, so clone answers ENOSYS and says why in the code rather than starting a thread somewhere that is not its own return point.
A runtime installs handlers before main and checks the return. Answering ENOSYS made programs abort at startup that would otherwise have run to completion, because most of them never raise anything. rt_sigaction, rt_sigprocmask and sigaltstack now succeed and record what was asked. Nothing is ever raised. Delivery means pushing a frame onto a guest thread's stack and redirecting it, and the trap mechanism hands out a register frame without any way to rewrite one, so it cannot be done from here yet. That limit is written in the file rather than hidden behind the success: a program that depends on SIGALRM will hang rather than misbehave quietly, which is the failure that can be diagnosed. SIGKILL and SIGSTOP are refused as uncatchable, as Linux refuses them.
net.sockets offers socket, connect, send, recv, close and a readiness poll, keyed by the caller's pid, which maps onto the Linux calls almost one to one. A guest's descriptor now holds a handle that service issued to this capsule, so a guest reaches only the sockets opened for it. read and write route by descriptor kind, so a program that treats a socket as a file, which most do, works without knowing the difference. Closing one closes the handle behind it. poll answers for every descriptor. A file or a console is always ready, which is what Linux reports too. A socket is asked one handle at a time, because that is the shape of the readiness call the service serves, and inventing a batched form here would mean a second protocol with nobody on the other end. The opcodes are transcribed rather than imported: the server is a binary and its protocol module is not a library. The file they came from is named beside them, since a number that changes there and not here is a wrong operation rather than a failed one.
supervised_asid returned an asid and dropped the lock that made it true, so peer map, unmap, copy and protect all ran against an asid the supervisor no longer held. It returns the guard with it now. The entry stub saves the callee-saved five plus a pad so a forked child resumes on the parent's register state. Nothing else on the path writes them to memory and a handler's prologue may already be using them, so it happens in the stub or not at all. Six slots keeps the frame 16-byte aligned and leaves existing offsets alone. libc gains wrappers for foreign exec, fork and resume, peer TLS and unmap, and the local signing and consent calls.
Zero capabilities was being treated as containment. brk, munmap and the wayland shm path took a guest number and acted on it, so a guest could unmap the personality or ask for a buffer no machine has. Limits come from one address plan in guest/layout.rs now. They had already drifted: only map.rs knew the mapping cursor's neighbour is EXEC_BASE and not the stack. Paths resolve to a Key only file::resolve can mint and the store wrappers take nothing else, so /linux is the root by construction. PT_INTERP and library mappings were loaded unproven, which left the attestation gate on the main image doing nothing useful. Both proven, ELF arithmetic checked. argv was read under MAX_PATH, so long arguments failed execve with nothing saying why. The resolver hands out 100.64/10 addresses and maps them back at connect, and the installer runs over the mixnet. resolve_host in net_sockets still does a clearnet lookup and is not fixed here. Package bytes are still unauthenticated, so place_entry refuses to vouch for them and the exec gate refuses what lands. tar::entries stopped at the first end-of-archive block, which in an apk is the end of the signature stream, so unpack had never written a file. Mirrored in python: one stream, sig+ctl+data, unterminated middle, garbage tail.
eKisNonos
force-pushed
the
linux/signals
branch
from
September 21, 2026 14:26
6babc8b to
5f1ce5f
Compare
Contributor
Author
|
Folded into #512. Both commits, |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two commits on top of what is already here. Neither can go to main:
userland/capsule_linuxandsrc/process/foreigndo not exist there.Peer boundary.
supervised_asidhanded back an asid and dropped the lock that made it true, so peer map, unmap, copy and protect all ran against an asid the supervisor no longer held. It returns the guard with it. The entry stub also saves the callee-saved five plus a pad so a forked child resumes on the parent's register state; nothing else on that path writes them to memory and a handler's prologue may already be using them, so it is the stub or nowhere. Six slots keeps the frame 16-byte aligned and every existing offset unchanged.Guest. Holding no capabilities was being read as containment.
brk,munmapand the wayland shm path took a guest supplied number and acted on it, so a guest could unmap the personality or ask for a buffer no machine has. Limits come from one address plan inguest/layout.rs. They had already drifted, and onlymap.rsknew the mapping cursor's neighbour isEXEC_BASErather than the stack.The guest also had the personality's whole VFS reach. Paths resolve to a
Keyonlyfile::resolvecan mint and the store wrappers take nothing else, so/linuxis the root by construction instead of by every caller remembering the prefix.PT_INTERPand library mappings were loaded unproven, which made the attestation gate on the main image pointless: prove the binary, then let it name an interpreter nobody looked at. Both proven, ELF arithmetic checked.argvwas read underMAX_PATH, so long arguments failedexecvewith no explanation.DNS: the resolver hands out 100.64/10 and maps back at connect, and the installer runs over the mixnet. One leak left, not fixed here,
resolve_hostinnet_socketsstill does a clearnet lookup before it dispatches on socket kind.Package bytes are still unauthenticated. APKINDEX signature unchecked,
C:checksum unread.place_entryrefuses to vouch for them and the exec gate refuses what lands. Checking them needs Alpine's RSA keys in the tree,verify_pkcs1v15already exists.tar::entriesstopped at the first end-of-archive block, which in an apk is the end of the signature stream, sounpackhad never written a file. Mirrored in python: one stream, sig plus control plus data, unterminated middle, garbage tail. The old walk on a real apk returns.SIGN.RSA.key.puband nothing else.Pairs with the local signing branch. That makes minting reachable; this is why that is safe.