Skip to content

linux: confine the guest and hold the supervised asid - #546

Closed
eKisNonos wants to merge 10 commits into
linux/signalsfrom
linux/peer-boundary
Closed

eKisNonos wants to merge 10 commits into
linux/signalsfrom
linux/peer-boundary

Conversation

@eKisNonos

Copy link
Copy Markdown
Contributor

Two commits on top of what is already here. Neither can go to main: userland/capsule_linux and src/process/foreign do not exist there.

Peer boundary. supervised_asid handed back an asid and dropped the lock that made it true, so peer map, unmap, copy and protect all ran against an asid the supervisor no longer held. It returns the guard with it. The entry stub also saves the callee-saved five plus a pad so a forked child resumes on the parent's register state; nothing else on that path writes them to memory and a handler's prologue may already be using them, so it is the stub or nowhere. Six slots keeps the frame 16-byte aligned and every existing offset unchanged.

Guest. Holding no capabilities was being read as containment. brk, munmap and the wayland shm path took a guest supplied number and acted on it, so a guest could unmap the personality or ask for a buffer no machine has. Limits come from one address plan in guest/layout.rs. They had already drifted, and only map.rs knew the mapping cursor's neighbour is EXEC_BASE rather than the stack.

The guest also had the personality's whole VFS reach. Paths resolve to a Key only file::resolve can mint and the store wrappers take nothing else, so /linux is the root by construction instead of by every caller remembering the prefix.

PT_INTERP and library mappings were loaded unproven, which made the attestation gate on the main image pointless: prove the binary, then let it name an interpreter nobody looked at. Both proven, ELF arithmetic checked. argv was read under MAX_PATH, so long arguments failed execve with no explanation.

DNS: the resolver hands out 100.64/10 and maps back at connect, and the installer runs over the mixnet. One leak left, not fixed here, resolve_host in net_sockets still does a clearnet lookup before it dispatches on socket kind.

Package bytes are still unauthenticated. APKINDEX signature unchecked, C: checksum unread. place_entry refuses to vouch for them and the exec gate refuses what lands. Checking them needs Alpine's RSA keys in the tree, verify_pkcs1v15 already exists.

tar::entries stopped at the first end-of-archive block, which in an apk is the end of the signature stream, so unpack had never written a file. Mirrored in python: one stream, sig plus control plus data, unterminated middle, garbage tail. The old walk on a real apk returns .SIGN.RSA.key.pub and nothing else.

Pairs with the local signing branch. That makes minting reachable; this is why that is safe.

A Linux program cannot run here, and making one run by putting a Linux
personality in the kernel would mean putting Linux in the trusted computing
base. This adds the opposite: a generic mechanism that lets an ordinary
capsule host a binary the kernel does not understand.

A foreign process is created with no capabilities at all. When it makes a
syscall this kernel does not recognise, and NONOS numbers are four
character tags so nothing legitimate lands there, the caller is parked and
the frame is handed to the process that created it. That supervisor decides
what the call means, builds the guest's address space through two peer
calls, and answers. Six syscalls, one capability, and no knowledge of any
other operating system in ring 0.

Peer map and peer copy refuse any process the caller did not create, and
copy through the guest's frames rather than its mapping, so a read-only
code page can be filled without ever being mapped writable and executable
at once. A guest dies with its supervisor.

The first user is a Linux personality capsule holding the ABI, an ELF
loader, and the calls a static binary makes. It is a userland program like
any other, signed and capability-bound, and the kernel gains nothing from
its presence.
The personality could host a binary that never opened a file, which is
almost none of them. This adds the file layer: openat and open, read,
write, close, lseek, stat, fstat, newfstatat and getdents64, over the
store through the client the rest of the tree uses.

Everything is opened under this capsule's own identity rather than the
guest's. A hosted process holds no capabilities, so the store would refuse
it, and the consequence is the useful one: a guest reaches exactly what
this capsule is granted and nothing else.

A descriptor now holds an open handle, a read offset, a pending write
buffer and, for a directory, the listing taken when it was opened. Writes
are held until close and then written as one file, because the store takes
whole values and a program that writes a file expects it to appear whole
or not at all. Paths are made absolute against a working directory and
flattened here, since the store has neither dot nor dot-dot.
A static binary opens and stats before it reaches main, and asks three
more questions on the way: whether its output is a terminal, which kernel
it is on, and where it is. Answering those with ENOSYS stopped programs
that the file layer alone would otherwise have carried.

Added pread64, getcwd, access, readlink, ioctl, fcntl and uname. The
answers are the true ones rather than the convenient ones. There is no
terminal behind any descriptor, so ioctl is ENOTTY, which is what a libc
is asking when it probes TCGETS and is what puts it into full buffering.
The store holds no symbolic links, so readlink is EINVAL on a path that
exists and ENOENT on one that does not. uname says Linux for the system
name because that names the ABI this capsule implements, which is the
question being asked, and the other fields say NONOS.

A directory entry is reported with an unknown type rather than a guessed
one, since the listing says what is there and not what each entry is, and
a caller that cares will stat the name.
NONOS refuses to map a page writable and executable at once, and that is
the right rule. It also means a loader cannot ask for the end state up
front: a dynamic linker maps a library writable, applies relocations to
it, and only then needs it executable. Without a way to make that second
change, mprotect had to be a lie that returned success and did nothing,
and the program faulted on its first call into the library.

The kernel gains a seventh peer call. MkPeerProtect sets the protection
of pages a guest already has, refusing any process the caller did not
create and any page that is not mapped, since a caller asking for execute
on a range it has not filled is not asking for what it thinks. It reuses
the same permission conversion and the same W^X refusal as the mapping
call, so the rule cannot be escaped through it.

I said in the mechanism PR that the ABI could grow without the kernel
changing again. That was too strong, and this is the exception: changing
the protection of frames a process already holds is not something the six
calls could express. It is still a generic primitive with no knowledge of
any foreign format in it.

With it, mprotect is real, and mmap can map a file privately: the pages
are mapped writable, the bytes are read into them, and the requested
protection is set afterwards. A request for write and execute together is
refused with EPERM rather than quietly granted as one of the two.
A static binary was all this could host, which rules out most software
that exists. A dynamic executable does not start at its own entry: the
interpreter named in PT_INTERP is loaded beside it and started instead,
and it maps the libraries and jumps to the program when it is done.

The loader now reads the interpreter path, fetches it from the store,
loads it at its own base, and reports where everything landed. A shared
object is biased and an executable is not, decided from the ELF type in
the file rather than from a flag a caller could get wrong.

The auxiliary vector is the other half. An interpreter finds the
program's headers through AT_PHDR, learns where it was itself placed from
AT_BASE, and jumps to AT_ENTRY at the end. AT_RANDOM is not optional
either: a C runtime takes its stack guard from those sixteen bytes before
it runs anything, and they come from the system entropy source rather
than a constant, which would make every guest's guard identical.

The program itself now comes from this capsule's arguments and is read
out of the store, so the personality runs what it is asked to run instead
of being a wrapper around one embedded image. That image stays as the
fallback, so the mechanism can still be proved on a machine with nothing
in the store.
A toolkit is not single threaded, so nothing with a window was ever going
to run without this.

The kernel gains an eighth peer call. MkForeignThread makes a thread
inside a guest: it shares the guest's address space, joins the same
supervisor so its own unknown syscalls redirect there too, and carries a
TLS base that the context switch loads from the control block, which it
must, because a C runtime reads thread-local storage before it runs any
program code.

The futex needs no kernel support at all. A guest thread that traps is
already parked inside its syscall with no reply, so a wait is the absence
of a reply and a wake is the reply. The personality compares the word
through a peer read, answers EAGAIN if it already changed, and otherwise
leaves the caller parked with its address recorded. A wake replies to as
many waiters on that address as were asked for. The serve loop therefore
takes traps from any thread of the guest and no longer assumes one reply
per trap.

clone is written and refuses. musl resumes a cloned child at the
instruction after its own syscall, so the child's entry is the caller's
return address, and the kernel does not pass it yet: the foreign redirect
was wired with a literal zero for the instruction pointer. Passing it
means touching the SYSCALL entry assembly, which is not something to do
untested, so clone answers ENOSYS and says why in the code rather than
starting a thread somewhere that is not its own return point.
A runtime installs handlers before main and checks the return. Answering
ENOSYS made programs abort at startup that would otherwise have run to
completion, because most of them never raise anything.

rt_sigaction, rt_sigprocmask and sigaltstack now succeed and record what
was asked. Nothing is ever raised. Delivery means pushing a frame onto a
guest thread's stack and redirecting it, and the trap mechanism hands out
a register frame without any way to rewrite one, so it cannot be done
from here yet.

That limit is written in the file rather than hidden behind the success:
a program that depends on SIGALRM will hang rather than misbehave
quietly, which is the failure that can be diagnosed. SIGKILL and SIGSTOP
are refused as uncatchable, as Linux refuses them.
net.sockets offers socket, connect, send, recv, close and a readiness
poll, keyed by the caller's pid, which maps onto the Linux calls almost
one to one. A guest's descriptor now holds a handle that service issued
to this capsule, so a guest reaches only the sockets opened for it.

read and write route by descriptor kind, so a program that treats a
socket as a file, which most do, works without knowing the difference.
Closing one closes the handle behind it.

poll answers for every descriptor. A file or a console is always ready,
which is what Linux reports too. A socket is asked one handle at a time,
because that is the shape of the readiness call the service serves, and
inventing a batched form here would mean a second protocol with nobody
on the other end.

The opcodes are transcribed rather than imported: the server is a binary
and its protocol module is not a library. The file they came from is
named beside them, since a number that changes there and not here is a
wrong operation rather than a failed one.
supervised_asid returned an asid and dropped the lock that made it
true, so peer map, unmap, copy and protect all ran against an asid the
supervisor no longer held. It returns the guard with it now.

The entry stub saves the callee-saved five plus a pad so a forked
child resumes on the parent's register state. Nothing else on the path
writes them to memory and a handler's prologue may already be using
them, so it happens in the stub or not at all. Six slots keeps the
frame 16-byte aligned and leaves existing offsets alone.

libc gains wrappers for foreign exec, fork and resume, peer TLS and
unmap, and the local signing and consent calls.
Zero capabilities was being treated as containment. brk, munmap and
the wayland shm path took a guest number and acted on it, so a guest
could unmap the personality or ask for a buffer no machine has. Limits
come from one address plan in guest/layout.rs now. They had already
drifted: only map.rs knew the mapping cursor's neighbour is EXEC_BASE
and not the stack.

Paths resolve to a Key only file::resolve can mint and the store
wrappers take nothing else, so /linux is the root by construction.

PT_INTERP and library mappings were loaded unproven, which left the
attestation gate on the main image doing nothing useful. Both proven,
ELF arithmetic checked. argv was read under MAX_PATH, so long
arguments failed execve with nothing saying why.

The resolver hands out 100.64/10 addresses and maps them back at
connect, and the installer runs over the mixnet. resolve_host in
net_sockets still does a clearnet lookup and is not fixed here.

Package bytes are still unauthenticated, so place_entry refuses to
vouch for them and the exec gate refuses what lands.

tar::entries stopped at the first end-of-archive block, which in an
apk is the end of the signature stream, so unpack had never written a
file. Mirrored in python: one stream, sig+ctl+data, unterminated
middle, garbage tail.
@eKisNonos

Copy link
Copy Markdown
Contributor Author

Folded into #512. Both commits, 3c3408a59 and 5f1ce5f37, are on linux/signals now that the chain has been rebased onto current main, so this branch has nothing left to add and its diff against the moved base is meaningless.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant