Skip to content

linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole - #582

Open
eKisNonos wants to merge 245 commits into
mainfrom
linux/go-guests
Open

eKisNonos wants to merge 245 commits into
mainfrom
linux/go-guests

Conversation

@eKisNonos

@eKisNonos eKisNonos commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Targets main. It also carries the 209 commits of #567 (linux/store-install), which come before it and hold the Linux personality these commits change, so it merges after #567 and shows those commits until #567 lands. What it adds is thirty-six commits.

Before this, a Linux-guest test image started no guest at all. Once it did, no guest could run a second thread: Go's threads called address zero, musl's pthread_create failed outright, and a thread's exit returned to it. A guest that did get threads reset the machine when it exited, left its threads running, and hung if one of them faulted. And nothing could wait: eventfd2 was not served, so every Go program with a timer died at its first one, while epoll_wait and a futex with a timeout returned at once or never.

What Linux guests can do now

  • Go programs start their runtime threads and run to a clean exit.
  • musl programs create pthreads, lock mutexes and join every thread.
  • A thread that faults ends the whole process, as on Linux; nothing is left running or waiting.
  • A fork from any thread gives the child that thread's registers and thread pointer.
  • Go's timers and poller work: time.Sleep, a timer woken early through Go's eventfd, and a pipe read through the poller to end of file.
  • musl programs wait as on Linux: timed condition variables, blocking eventfd reads, epoll_wait, poll and select timeouts, non-blocking pipes, several readers or a writer waiting on one pipe, edge-triggered epoll, and timerfd one-shot, periodic and absolute.
  • The descriptor ioctls, the scheduler calls, epoll_create, epoll_pwait2 and tgkill are served as Linux serves them to an unprivileged process on one CPU.

What changed

The boot guest starts

  • mk bf9d9dbf3: the Linux-guest test image (NONOS_LINUX_GUESTS=1) builds microkernel-desktop-gui with nonos-stark-attest, without first-boot setup. Under the setup profile every app, the personality included, waits for setup to exit, and setup waits for keys, so the unattended image never started its guest: 0 [APP-LINUX] or [LINUX] lines in 180 s. Every other build of nonos-mk-desktop-gui-prod is unchanged.

Threads start, run and end

  • foreign 3fc1ac19d: a clone child starts on a copy of the calling thread's registers, with rax 0, the new stack, and the parent's thread pointer unless a TLS is named. MkForeignThread takes the calling thread as a fifth argument; it must be supervised by the caller and in the same thread group as the target, so no guest receives another's registers. Zero keeps the fresh start. Go's child calls r12 and musl's calls r9, so both called address 0 before. A TLS is taken only when CLONE_SETTLS is set. abi/syscalls.toml records the new argument.
  • linux addec49af: clone writes the new tid where CLONE_PARENT_SETTID and CLONE_CHILD_SETTID ask. musl's thread-list lock stores its owner's tid, and a thread that never learned its tid locked the list as 0, which the lock reads as free.
  • linux 1e6b8a673: mprotect commits a PROT_NONE reservation it is asked to open. musl reserves every pthread stack with PROT_NONE and opens the usable part with mprotect; MkPeerProtect reprotects only pages that exist, so pthread_create failed. A backed region is reprotected as before; a reservation is backed with the asked protection, keeping any page the guest already touched, and recorded as backed so fork copies it. A span outside every region is refused with ENOMEM, as on Linux.
  • linux 10947aa96: a thread's exit ends the thread and never returns to it. The word it named through CLONE_CHILD_CLEARTID or set_tid_address is zeroed and one waiter woken, as Linux does. musl names its thread-list lock there and pthread_join waits for that lock, so every join hung before. Go's exitThread falls into INT3 when exit returns, which the fault report below would have turned into the end of the process.

A process ends whole

  • exit 5ab0a0912: a thread group's page tables stay until its last member leaves the process table. Release freed them when the leader was finalized, under a Go thread still running on them, and the machine reset: CR3 was the guest's, FS held the Go thread's TLS, CR2 was the IDT entry for vector 8. The tables now pass to a process still running on them, whose capability token is rebound to the ASID it now owns.
  • kill 874f3ecfc: MkKill admits the caller the foreign registry names as the target's supervisor. A thread's parent is its leader, not the supervisor, so every thread kill at a guest's exit returned EPERM and the threads ran on. A refused kill now prints [LINUX] kill refused: pid <n> outlives its process, errno <e>.
  • foreign 2311afc39: when a guest thread ends on a signal, and only when it ends itself so a supervisor's own kill cannot loop back, the kernel leaves a one-way notice for its supervisor, collected on its next MkForeignWait as a frame numbered FOREIGN_NR_DIED. The kernel reports; what a dead thread means is the personality's policy.
  • linux 959b57f98: the personality ends the process on that notice, as Linux ends a thread group on an unhandled fault, and reports the signal as 128+signo.

Fork

  • foreign ffc1b542b: a fork copies the calling thread's parked frame, and the kernel carries that thread's own thread pointer to the child. Before, every fork took the leader's registers and the personality's single fs_base, which is only the last thread to set one.

Waiting, as Linux waits

  • linux 69f3b48d4: a pipe has Linux's ends. The family notes, each time it lends its pipes to the guest being answered, which ends are still open anywhere in the family. From that note: a read end is readable while it holds bytes and hung up once no write end is left; a write end is writable while there is room and in error once no read end is left. An empty pipe with no write end reads as end of file, and a write no reader can take is refused with EPIPE. poll, epoll and select report hang-up and error unasked. Before, every pipe was readable and writable to all three.
  • linux 3e48edbaf: a descriptor keeps O_NONBLOCK. F_SETFL sets it, F_GETFL reports it with the access mode, pipe2 takes it, dup and fork carry it, and a non-blocking read of an empty pipe answers EAGAIN. Before, F_SETFL was dropped, which is how Go's runtime prepares every pipe for its poller.
  • linux bfb07892f: eventfd2 and eventfd are served. The counter is kept with the pipes as the family's, so dup, fork and every thread reach one. EFD_SEMAPHORE, EFD_NONBLOCK and EFD_CLOEXEC as Linux; all ones and a short buffer EINVAL. Go's runtime makes one at its first timer and threw when it could not.
  • linux 06a266ac5: epoll_wait and epoll_pwait wait for their timeout (an int of ms; negative waits until ready). A blocking eventfd read or write and a write to a full pipe wait too. A waiting call is left parked in its trap and tried again after every answer and at its deadline, so a write from another thread ends the wait at once. A wait on a socket or timer is looked at every 10 ms. A pipe write of up to PIPE_BUF goes in whole or waits. maxevents of zero or less is EINVAL.
  • linux 32d2a1a06: EPOLLET entries are reported as readiness rises and re-armed by an EAGAIN; EPOLLONESHOT reports once. epoll_ctl refuses as Linux: EBADF, EPERM for a regular file or directory, EINVAL, EEXIST, ENOENT. Go registers every pipe and socket with EPOLLET and EPOLLOUT, so a level-triggered report of an always-writable end would answer its poller at once, every time.
  • linux aa0087d75: a futex wait ends with ETIMEDOUT at its timeout. FUTEX_WAIT_BITSET (absolute, MONOTONIC or REALTIME), FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE are served. The value is compared as 32 bits, so musl's sign-extended -1 matches. musl's timed waits and condition variables, and Go's sysmon, use these.
  • linux b1ead15c5: two proof guests, gopoll and cwait (below).

More waiting, timers, descriptors and threads

  • linux b716a34af: any number of threads can wait to read one pipe. A read parked in one slot per process, so a second reader of an empty pipe took the slot and the first was never answered. Pipe reads now wait with the other parked calls; the family settles them after it reaps, so a writer that left with its process reads as end of file at once. A zero-byte read answers zero.
  • linux 439ed66c6: poll, ppoll, select and pselect6 wait for their timeout (poll's int of ms, ppoll's and pselect6's timespec, select's timeval; null or negative waits until ready; bad fields EINVAL). poll(NULL, 0, ms) sleeps. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out. A negative poll descriptor is never ready.
  • linux 85b3d7724: a closed descriptor leaves every epoll interest list, as on Linux, so its number can be added again; before, that add was EEXIST and the old entry reported the new file under the old token. dup2 closes an open target first (a file's buffered bytes were lost, a socket's handle kept) and refuses a number past the table.
  • linux b9efc7f96: timerfd has Linux's semantics and is kept with the family like the eventfd counter, so dup and fork share it. A periodic timer keeps firing and a read returns how many times it fired; the clock is Linux's set (else EINVAL); TFD_NONBLOCK, TFD_CLOEXEC and TFD_TIMER_ABSTIME are honoured; the old setting is written when asked; timerfd_gettime is served; a blocking read waits for the firing; a wait watching a timer is woken when it fires.
  • linux 262c06280: FIONREAD (a pipe's bytes, a file's remainder), FIONBIO, FIOCLEX and FIONCLEX. Any other request is ENOTTY, as for a descriptor with no terminal or device behind it.
  • linux 8f778075b: a dup'd or inherited epoll descriptor carries the interest list it had; before, it started empty.
  • linux ff778c704: sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam, sched_get_priority_max and _min, sched_setaffinity, epoll_create, epoll_pwait2. Every thread is SCHED_OTHER at zero; SCHED_FIFO and SCHED_RR are EINVAL outside 1 to 99 and EPERM inside it, the range checked first as Linux checks it; a mask without the one CPU is EINVAL; pid arguments are mapped like kill's.
  • linux 9039b893d: tgkill's thread group and thread are mapped into the family's pids. gettid and getpid answer with the family's numbers, so tgkill(getpid(), gettid(), sig), which Go's runtime and glibc's pthread_kill use, was ESRCH.
  • linux 84e10f538, 7c3967f18, 468f5aa97, 2f0db3b25, a045bec06, 66b55faf8, afd96ef58, 24cafe103: the guest gopreempt, and seven more cwait parts, one per change above; cwait now runs every part and names each one that fails.

Cleanups

  • linux 289272374: delete call/clone.rs, a copy left behind when clone moved to call/spawn/; no module included it.
  • mm ce395f7ff: the [PF] demand fill line is printed only after a fill, not before the handler refuses the null page or the kernel half.

Evidence

q35 under TCG, one vCPU. Each guest is the boot guest of its own store.

Threads, faults and exit (the first thirteen commits)

guest proves line calls served unserved triple faults
gohello Go runtime threads, GC, clean exit [GO] hello PASS 271 0 0
goconc 8 goroutines on runtime threads [GO] conc PASS: 8 goroutines summed 3199960000 236 0 0
cthreads 8 musl pthreads, a mutex, 8 joins [C] cthreads PASS: 8 pthreads joined, summed 3199960000 87 0 0
threadfault a worker faults while main waits [LINUX] guest thread 51 ended on a signal; ending the process, then [LINUX] guest exited 7 0 0

threadfault's worker runs on a stack mapped read-write outright through clone(), so it does not depend on the pthread path. Its SURVIVED line, printed only if main outlives the faulting thread, never appears.

Waiting (the next seven commits)

Both guests pass on host Linux, which is what they are measured against. On one build carrying all twenty commits:

guest proves line calls served unserved triple faults
gopoll Go's timer and poller: a 50 ms sleep, a 10 ms timer set while the poller waits on a 3 s one, a pipe read through the poller [GO] poll PASS: 3 parts, 3454ms in all: slept 69 ms, the early timer ended in 18 ms, the pipe read to end of file in 339 ms 399 0 0
cwait musl waiting in 9 parts [C] cwait PASS: 9 parts in 951 ms: cond_timedwait ETIMEDOUT after 102 ms, a blocking eventfd read woken after 147 ms, epoll_wait timed out after 100 ms and woken after 153 ms, POLLHUP on a closed pipe, a full-pipe write waited 148 ms, edge-triggered counts as Linux 196 0 0

The same build ran the first four guests again: gohello 283 calls, goconc 253, cthreads 87, threadfault 7 with its death line; 0 unserved in each.

More waiting, timers, descriptors and threads (the next sixteen commits)

On one build carrying all thirty-six commits:

guest proves line calls served unserved triple faults
cwait 16 parts, the 9 above and 7 more [C] cwait PASS: 16 parts in 2313 ms; SCHED_FIFO at priority 1 errno 1 and a CPU-1-only mask errno 22; a 50 ms periodic timer fired 3 times in 180 ms, a one-shot read waited 103 ms, an absolute timer woke epoll after 105 ms 345 0 0
gopoll as above [GO] poll PASS: 3 parts, 3437ms in all 386 0 0
gopreempt a goroutine spinning with no call beside main, on one P no line in 360 s: this set does not reach a running thread, the gap named below 0

The same build ran the regression set: gohello 282 calls, goconc 248, cthreads 87, threadfault 7 with its death line; 0 unserved in each.

Each fix was also run with its change removed, on the same guest:

change removed guest result without it result with it
the supervisor clause in MkKill goconc 3 kill refused ... errno 1 lines; guest threads fault after guest exited 0 and 0
the kernel's death notice threadfault SURVIVED printed, no death line death line, no SURVIVED
the clear-tid write and wake cthreads no PASS, no FAIL, no exit in 241 s after the guest started PASS
the mprotect commit (the pthread build of threadfault, before the fix) threadfault [C] threadfault FAIL: no worker pthread_create succeeds (cthreads)
a parked wait tried again when something changes (only at its deadline instead) gopoll a 10ms timer set under a 3s one took 2823ms, FAIL 18 ms, PASS
the same cwait hangs in its blocking eventfd read; nothing after part 3 in 420 s PASS
eventfd2 served gopoll [LINUX] unserved nr=290, then fatal error: runtime: eventfd failed PASS
the futex timeout cwait hangs in cond_timedwait; no line in 420 s ETIMEDOUT after 102 ms
EPOLLET (reported level-triggered instead) cwait FAIL: edge-triggered counts (222, 1): the second look counts 2, not 1 PASS
the same gopoll PASS, with 1282 calls served: the poller answered at once while the pipe was open 399 calls
a pipe's hang-up cwait FAIL: end of file after the write end closed (0, 0): poll sees nothing POLLHUP
O_NONBLOCK kept by fcntl and pipe2 cwait hangs in the non-blocking pipe read; nothing after part 6 in 420 s EAGAIN at once
the personality before the next sixteen commits cwait eight parts pass, then it hangs in pipe-readers: the first of two blocked readers is never answered PASS
poll, ppoll and select waiting (answering at once instead), a closed descriptor leaving epoll, periodic timers, FIONREAD, sched_getscheduler, the tgkill mapping, all six removed on one build cwait FAIL: 6 parts failed, 10 passed: poll and ppoll returned after 2 ms, the re-add was EEXIST, the periodic timer counted 1, FIONREAD failed, [LINUX] unserved nr=145, tgkill ESRCH 16 parts pass
the epoll list carried through fork cwait FAIL: epoll list through fork (13, 256): the child finds nothing ready and exits 1 PASS

Each of these ran on its own build with only that change removed, except the six that fail different cwait parts, which were removed together. gopoll passes with the hang-up removed, as its writer closes before its reader waits, and with O_NONBLOCK dropped, as Go's pipe reads then park instead of polling.

The store settled in 12.8 to 84.2 s across the 22 guest boots for the first thirteen commits, in 38.0 to 104.8 s across the 16 for the next seven, and in 17.2 to 117.7 s across the 10 for the sixteen after; the personality waits up to 300 s.

Checks

  • x86_64 kernel crate: 0 errors; 43 warnings, the same set before and after.
  • check_stubs, check_allows, check_dark_features: 0 new sites. check_syscall_abi: 111 published syscalls reach a handler.
  • check_unreachable: 1 new site, has_children, which linux/store-install reports too; this branch has 1257 sites against its 1258.
  • Each of the first 13 commits builds on its own: the kernel (microkernel-core) and the personality were checked at every one, with 0 errors. The next twenty-three change only the personality and the guests; the personality was checked at each with 0 errors and 0 warnings, and each guest passes on host Linux.

Not done here

  • A PROT_NONE reservation is not enforced: the kernel demand-fills any guest page on first touch, so a guard page does not stop a stack overflow.
  • mprotect on a backed region does not update the recorded protection, so a fork after it gives the child the protection the region was mapped with.
  • A fork does not copy a reserved region the guest has touched, since only backed regions are copied.
  • The leader's own exit (not exit_group) still ends the whole process; on Linux the others would run on.
  • A caught signal reaches a guest thread only when it next makes a call. A thread parked in a wait sees it when the wait ends, and a thread running user code never does, so Go cannot preempt a goroutine that spins without a call (gopreempt, above). The kernel change that stops a running thread for its supervisor follows this set.
  • A caught signal sent with kill to a child process is queued in the sender, not the child, so the child's handler never runs; an uncaught one ends the child as it should.
  • A write to a pipe with no reader is EPIPE, but SIGPIPE is not raised.
  • A wait on a socket is looked at again every 10 ms; nothing tells the family sooner that a socket became ready.
  • A dup'd or inherited epoll descriptor gets a copy of the interest list; on Linux both share one, so a change made on one side after the dup or fork is not seen on the other.
  • TFD_TIMER_CANCEL_ON_SET is accepted and has no effect.

eKisNonos and others added 30 commits September 20, 2026 12:00
Reads the signed marketplace index, lists what the running image can
install, and installs through the queue init already owns.

The install buttons did nothing: the painter and the hit test each
computed their own rectangles, so a click never matched a row. Both
derive from one geometry module now and the click path reaches the
same install::ask as Enter.

Search filters as typed, the list scrolls with a real scrollbar, and
selection survives a refresh.

init gains an install queue and a wake path so a store request is
serviced on arrival rather than on the next supervisor spin. The
surface registry can attach frames to a surface a capsule owns.

Scrollbar arithmetic mirrored in python: empty, shorter than the
viewport, thumb at both ends, single row overflow.
Release encode, decode and signing move into marketplace_abi, so the
tool that writes the index and the capsule that reads it share a codec
instead of having one each.

install_ready checks arch and readiness against the running image
rather than the index, and adds a seventh gate for whether a release
carries a zk trailer for its own measurement. The other six are
signatures over the artifact.

The catalogue generator reads what is on disk and writes JSON; the CLI
encodes the binary the market ingests. Signing is a separate step with
the operator seed, which is not in any build rule.
IconId::Store points table.rs at assets/icons/store.a8, which only
existed on the app-store branch, so every build of this branch failed
to read it. Add the mask and its SVG source here so the table stands
on its own.
text::line returns the drawn width, so the bare match evaluated to i32
where the function body expects (), failing the capsule build.
nonos-data/marketplace/index.bin has no make rule, so naming it as a
hard prerequisite failed every build on a checkout without it (CI:
No rule to make target). Wrapping it in $(wildcard) keeps the rebuild
on a newer catalogue where it exists and drops the prerequisite where
it does not.
The branch's manifest predated the switch to in-process Ed25519 and
dropped the dependency while verify/crypto.rs imports it, so the
capsule failed with an unresolved import. Restore main's manifest and
add only the app_skeleton dependency boot_index.rs needs.
Twenty-five syscalls were published as caps = ["valid_token"] while the
cap table demands a hardware, dev-root or time capability. Twenty-two take
any one of Admin or a hardware capability, now published with caps_any; the
dev-root and time calls need one capability, published with caps.
The cap table gates MDRO with MDRQ and MDRC on can_enrol_dev_root. The old
syscall caps check cannot resolve that predicate and wants valid_token, so
this fails it until abi/caps-check-fail-closed lands; the fixed check passes.
Every lane was pinned to -accel hvf -cpu host, which only macOS has, so
no Linux host could boot an image. KVM when /dev/kvm opens read-write,
hvf on macOS, TCG otherwise, the rule the boot matrix already uses; the
display and audio backends follow the host too.
The trailer's magic picked the verifier for every root, so a Pedersen
trailer was checked against the vendor root too. A local root's leaf is a
commitment to a secret this kernel holds; the vendor root's is not, so there
the trailer may no longer choose the weaker proof.
MkLocalSign let a LocalSign holder prove any capability it held, including
LocalSign itself, so one signer could hand out the right to sign. A local
proof now names nothing beyond AMBIENT_CAPS. crypto_proofs checks a minted
trailer and its refusals against the kernel's own verifier files.
An Alpine package index is signed that way, and the capsule offered SHA-1
only without the prefix, so the index could not be checked at all. The
rsa crate rebuilds the whole padded block and compares it. Nothing here
signs, so offering SHA-1 verification mints nothing new with it.
resolve.rs names super::root, which the crate never mounted, so it did not
build and nothing noticed because no workflow ran it. It mounts root.rs,
follows the rename of absolute to visible, adds the /linux confinement and
Phdr::file_range tests, and joins the proof-crate matrix.
A signature over a package covers the compressed bytes of one member, so a
verifier needs to know where each one starts and ends. members() refuses a
file with any byte that belongs to no member, where gunzip ends the stream
and ignores the rest.
The index's .SIGN.RSA entry is checked against Alpine's x86_64 keys by the
crypto service, the package's control member against the index's C: SHA-1,
and its data against the control member's datahash. Only a Verified value
reaches the store, so unauthenticated bytes are refused, not kept unvouched.
Without the operator-key rotation: NOX_OPERATOR_V1 keeps main's key, now
read from .keys/marketplace_operator_ed25519.pub. The capsule embedded an
index nothing built, so tools/nonos-market-index writes one, signed and
verified when the operator seed is present and empty otherwise. nonos-mk
moves to 78dae45 for the zk_trailer_hash field; the icon table is 49 long.
A release naming x86_64-linux counted as having its attestation, so any
release could claim the exemption for itself. It now needs the linux.
namespace too, which is where the store routes it: to the installer that
authenticates the bytes before the machine mints their proof.
Conflicts were both sides adding: the install and app-store capsules sit
together in Cargo.toml, mk and userspace, and init/mod.rs names the
install queue once. The init loop and install queue are taken as the PR
wrote them; the commit after this replaces its wake path.
The scheduler takes every ready process's priority lock from the timer
interrupt. wake.rs and boost_init_for_drain took init's with interrupts on
from syscall context, so a tick inside either spun forever on one CPU. The
install queue now raises through the guarded setter the window queue uses.
MkAppInstall took a package name and the store's own readiness flag. It now
takes a listing and release; init asks the market for readiness and the
release's package hash, and the installer refuses bytes of any other BLAKE3.
The index keeps each record's D: and p: lines, so dependencies come too.
…each

CryptoMachineKey derives the machine key for any label a Crypto holder
names, so a key the kernel keeps for itself needs a label no syscall can
ask for. Kernel labels start with a zero byte, and the syscall refuses any
label that does.
The local signing identity was random each boot, so consent was too. It is
now derived from the machine key, and first-boot setup, which alone holds
EnrolDevRoot, grants the local root as a named step and keeps a token only
this machine can make; later boots restore it. The desktop profile now
includes setup and the market, and builds every capsule it embeds.
MkAppLaunch queues a run for init, which spawns the personality to start
the program the installer recorded outside /linux. Whether it may start is
the exec gate's answer. The store drops its console-code enrolment, which
it never held the capability for, and gains an o key to open.
A guest thread's exit was answered with zero, so the thread ran on past it.
musl loops on exit, and its second call parked the thread for good, never
reaped. Go's exitThread falls into INT3 when exit returns, which the fault
report now turns into the end of the whole process.

The exit is now Linux's. The thread is left unanswered and killed, so it never
returns. The word it named for clearing, through CLONE_CHILD_CLEARTID or
set_tid_address, is zeroed and one waiter woken. musl names its thread-list
lock there, and pthread_join waits for that lock to be released, so every join
hung before. set_tid_address now records the word and still answers with the
tid. A refused kill is named in the log.
cthreads creates 8 musl pthreads, sums known ranges under a mutex, joins them
all and checks the total. pthread_create exercises the reserved stack committed
by mprotect and the tid written by clone, the mutex the futex, and each join the
clear-tid word zeroed and woken at the thread's exit.
A pipe reported readable and writable to poll, epoll and select whatever it
held, a read of an empty pipe whose writers had all gone answered EAGAIN, and
a write that no reader could ever take was kept. A program that waits for a
pipe to become readable before reading it either spun or never saw end of
file.

The family now notes, each time it lends its pipes to the guest being
answered, which ends of each pipe are still open anywhere in the family,
since a fork leaves them in more than one process. From that note, as
Linux's pipe_poll: a read end is readable while it holds bytes and hung up
once no write end is left; a write end is writable while there is room and
in error once no read end is left. An empty pipe with no write end reads as
end of file, and a write with no read end is refused with EPIPE. poll, epoll
and select report hang-up and error whether they were asked for or not.
fcntl accepted F_SETFL and dropped it, and F_GETFL answered O_RDWR for every
descriptor. pipe2 ignored O_NONBLOCK. So a read of an empty pipe parked its
thread even when the program had asked never to wait, which is how Go's
runtime sets every pipe it gets before handing it to its poller.

A descriptor now keeps O_NONBLOCK: F_SETFL sets or clears it, pipe2 sets it
on both ends, dup and fork carry it. F_GETFL answers the access mode (a pipe's
read end read-only, its write end write-only, as pipe2 makes them) with
O_NONBLOCK when it is set. A non-blocking read of an empty pipe is answered
EAGAIN at once, or end of file once no write end is left. The other status
flags F_SETFL takes are accepted and still have no effect.
eventfd2 was not served, and Go's runtime makes one for its poller at the
first timer and throws when it cannot, so every Go program that sleeps or
sets a timer died there.

eventfd2 and eventfd now make a counter. The counter is the object and
the descriptor only names it, so it is kept with the pipes as the family's,
lent to the guest being answered: dup, fork and every thread reach the
same one, as on Linux. A read takes the whole count, or one under
EFD_SEMAPHORE; a write adds to it; a starting value is an unsigned int; a
write of all ones and a buffer under eight bytes are refused with EINVAL.
Where Linux would wait (a read at zero, a write that would pass the most a
counter holds), the call answers EAGAIN for now. poll, epoll and select see
it readable while the count is above zero and writable while one more fits.
EFD_CLOEXEC and EFD_NONBLOCK set the descriptor's flags; any other flag is
refused with EINVAL.
epoll_wait and epoll_pwait answered at once whatever their timeout, so a
program waiting for a descriptor or a deadline spun, and one waiting with
no timeout spun for good. Go's runtime waits there for its next timer, and
wakes that wait from another thread by writing its eventfd. A read of an
empty eventfd and a write to a full pipe answered EAGAIN where Linux waits.

A call that has to wait is now left parked in its trap and tried again
after every answer the family gives and at its deadline, so a write from
another thread or process completes the wait at once. epoll_wait takes its
timeout as an int of milliseconds: zero reports what is ready now, a
negative one waits until something is, and a deadline that passes answers
zero. A maxevents of zero or less, or past Linux's bound, is refused with
EINVAL. A read or write of an eventfd, and a write to a full pipe, wait the
same way unless the descriptor is non-blocking. A pipe write of up to
PIPE_BUF bytes goes in whole or waits, as Linux's does. A wait on a socket
or a timer is looked at again every 10 ms, since nothing the family answers
changes them. The loop wakes for the nearest deadline. A killed thread's
parked call is dropped, so it cannot take what a live one waits for.
Every epoll entry was reported for as long as its descriptor was ready,
whatever it asked for. Go registers every pipe and socket with EPOLLET and
EPOLLOUT, and a write end with room is always writable, so once epoll_wait
waits, Go's poller would be answered at once on every call and spin.
epoll_ctl accepted a regular file, adding twice, and changing or dropping
an entry that was not there.

An EPOLLET entry is now reported when its readiness rises: bits already
seen at the last look are left out until they fall, or until a read,
write, send, receive or accept on the descriptor answers EAGAIN, or a
connect EINPROGRESS, which is where a program using EPOLLET stops and
waits. An EPOLLONESHOT entry reports once and then nothing until it is
modified. What was seen is kept only once the events are delivered.
epoll_ctl refuses as Linux does: EBADF for a closed descriptor, EPERM for a
regular file or directory (Go's os.Open then reads it blocking), EINVAL for
the list itself, EEXIST for adding twice, ENOENT for changing or dropping
what is not there. A change looks at the entry afresh.
A futex wait ignored its timeout and waited until it was woken, and
FUTEX_WAIT_BITSET, FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE
answered ENOSYS. musl's pthread_cond_timedwait, sem_timedwait and
pthread_mutex_timedlock wait with a timeout, musl hands a condition
variable's waiters to its mutex with FUTEX_REQUEUE, and Go's sysmon sleeps
on a timed futex. A value compared against the word was also taken as 64
bits, so musl's sign-extended int of -1 never matched.

A timed wait now ends with ETIMEDOUT at its deadline, unless something wakes
it first, and at once if the deadline has already passed. FUTEX_WAIT's
timeout is relative; FUTEX_WAIT_BITSET's is absolute on CLOCK_MONOTONIC, or
on CLOCK_REALTIME with FUTEX_CLOCK_REALTIME. A timespec with negative
seconds or nanoseconds past a second is EINVAL. The bitset forms are served
with the bitset taken as matching every waiter, which at worst is a
spurious wake, and a zero bitset is EINVAL. FUTEX_REQUEUE wakes up to val
waiters and moves up to val2 more to the second word; FUTEX_CMP_REQUEUE
first checks the word against val3. The value is compared as 32 bits. The
loop wakes for the nearest futex deadline, and a thread being ended has all
its waits dropped.
gopoll is Go's network poller in three parts, each printed as it passes: a
50 ms sleep that needs epoll_pwait to wait its timeout, a 10 ms timer set
while the poller already waits on a 3 s one, which needs Go's eventfd write
from another thread to end that wait, and an os.Pipe read through the
poller to end of file, which needs O_NONBLOCK, edge-triggered readiness and
a hang-up once the write end closes.

cwait is the same for musl in nine parts: cond_timedwait's futex timeout, a
condition variable broadcast, eventfd counts and semaphore reads, a blocking
eventfd read woken by another thread, epoll_wait's timeout and its wake from
another thread, a non-blocking pipe with fcntl and end of file, a write to a
full pipe waiting for a reader, and edge-triggered epoll. Both pass on host
Linux, which is what they measure the personality against.
A read parked on an empty pipe was kept in one slot per process, so a
second thread reading an empty pipe took the slot and the first was never
answered. A read of zero bytes from an empty pipe with a writer answered
EAGAIN rather than zero.

A read of an empty pipe now waits with every other parked call, tried again
after each answer until it gets bytes, or end of file once no write end is
left anywhere in the family. The slot and its settling are gone. The family
settles waits after it reaps, so a write end that left with its process
reads as end of file at once. A read of zero bytes answers zero.
The four answered at once whatever their timeout, so a program waiting on
descriptors spun, and poll(NULL, 0, ms), the idiom for a short sleep, did
not sleep. select also narrowed the caller's sets even when nothing was
ready, and copied a whole 1024-bit set whatever nfds was. poll counted an
entry with a negative descriptor, which Linux ignores and programs use to
switch an entry off, as closed and so ready.

Each now waits with the other parked calls until something it watches is
ready or its timeout passes: poll's int of milliseconds, negative for none;
ppoll's and pselect6's timespec and select's timeval, null for none, with
negative or out-of-range fields refused with EINVAL. A timeout of zero
only looks. select writes its sets back only when something is ready,
copies only the longs nfds covers, always empties the exception set, and
empties all three when its time runs out, as Linux does. A negative poll
descriptor is never ready. A wait watching a socket or a timer is looked
at again every 10 ms, now for poll and select as for epoll.
Two parts added to cwait. Two threads block reading one empty pipe and both
must be answered once bytes arrive. poll, ppoll and select each wait out a
100 ms timeout with nothing ready, poll(NULL, 0, 50) sleeps, and a select
with no timeout is woken by another thread's write, with its set narrowed
to the ready descriptor. Both pass on host Linux.
…aces

A closed descriptor stayed in every epoll interest list. Linux drops it, so
a program that closes without EPOLL_CTL_DEL and gets the same number back
from its next open adds it again; here that add was refused with EEXIST,
and the old entry reported the new file under the old token. dup2 onto an
open descriptor overwrote it without closing it, so a file's buffered bytes
were lost and a socket's handle kept, and a target number past the table
grew the table to reach it.

close now drops the descriptor from the guest's interest lists. dup2 closes
an open target first, as Linux does, and refuses a number past the table
with EBADF.
A twelfth part: a watched pipe end is closed without EPOLL_CTL_DEL, a new
pipe takes its number, and adding it again succeeds with no stale event.
dup2 to a number past the table is EBADF. Passes on host Linux.
A timerfd fired once and forgot its interval, so a periodic timer stopped
after its first read, and a read always answered one. timerfd_create
ignored its clock and its TFD_NONBLOCK and TFD_CLOEXEC flags,
timerfd_settime ignored TFD_TIMER_ABSTIME and never wrote the old setting,
timerfd_gettime was not served, and a read of a timer that had not fired
answered EAGAIN even to a blocking descriptor. The expiry lived in the
descriptor, so dup and fork gave a timer that never fired.

A timer is now an object the family keeps, lent with the pipes and eventfd
counters, so every descriptor onto it sees one timer. It is made on a
clock Linux names (anything else is EINVAL) with its flags; set relative or
absolute on its own clock, one-shot or periodic, with the old setting
written when asked; read for how many times it has fired since the last
read, stepping a periodic timer past them; read blocking until it fires
unless the descriptor is non-blocking; and reported readable once it has
fired. A wait watching a timer is woken when it fires rather than on a
10 ms look.
A thirteenth part: an unarmed non-blocking timer reads EAGAIN, a 50 ms
periodic timer counts its firings over 180 ms and reports its interval, a
one-shot read blocks until it fires, and an absolute time 100 ms ahead
wakes an epoll wait. Passes on host Linux.
gopreempt spins a goroutine with no call in it beside main, on one P. Go
moves such a goroutine off the CPU only by sending SIGURG to its thread
while it runs, so main waking from a 20 ms sleep and a garbage collection,
which stops the world, both depend on that signal landing. On host Linux
main runs again after 25 ms and the collection ends by 45 ms, with the
tgkill deliveries visible under strace.
ioctl answered ENOTTY to every request, so a program asking how many bytes
a pipe or file held, setting a descriptor non-blocking the BSD way, or
marking it close-on-exec without fcntl was refused.

FIONREAD answers what a pipe's read end holds, or what is left of a file
past its offset; FIONBIO sets or clears O_NONBLOCK from the int it is
given; FIOCLEX and FIONCLEX set and clear close-on-exec. A socket's
FIONREAD, and every device request, is still ENOTTY: there is no terminal
or device behind a descriptor here, which is also how isatty says no.
A duplicated or inherited epoll descriptor started with an empty interest
list, so a child that waited on the epoll its parent set up before fork
waited on nothing. The list is now copied with the descriptor. Linux shares
one list between them; the copy holds what was registered at the dup or
fork, and a change made afterwards on one side is not seen on the other.
A fourteenth part: FIONREAD counts three bytes in a pipe, FIONBIO makes its
read end answer EAGAIN once drained, FIOCLEX sets close-on-exec, and a
forked child finds a ready entry in the epoll list its parent built.
Passes on host Linux.
sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam,
sched_get_priority_max and _min and sched_setaffinity were not served, nor
epoll_create and epoll_pwait2, so a program that sets a thread's policy or
pins it, or a libc that reaches for the older or newer epoll form, died
there.

They are answered as Linux answers an unprivileged process on the one CPU
the guest is shown: every thread is SCHED_OTHER at priority zero; asking
for SCHED_OTHER, BATCH or IDLE at zero is accepted and changes nothing,
since scheduling is the kernel's; asking for FIFO or RR is EINVAL outside
the priority range 1 to 99 and EPERM inside it, the range checked first as
Linux checks it, since no guest holds the privilege Linux asks for; the
priority ranges are Linux's. An affinity mask that includes the one CPU is
accepted and one that leaves it out is EINVAL. A pid argument is mapped
from the guest's namespace like kill's, and one outside the family is
ESRCH. epoll_create checks its size is positive and makes a list;
epoll_pwait2 waits like epoll_pwait with a timespec timeout.
A fifteenth part: the policy is SCHED_OTHER, SCHED_FIFO at priority zero is
EINVAL, SCHED_OTHER is accepted, the real-time range is 1 to 99, CPU 0 can
be pinned, epoll_create refuses a size of zero, and epoll_pwait2 waits its
50 ms timespec. SCHED_FIFO at priority one and a CPU-1-only mask depend on
privilege and CPU count, so they are printed, not checked. Passes on host
Linux.
gettid and getpid answer with the family's own numbers, and kill and tkill
map theirs back, but tgkill's were passed through as they came. A thread
signalling itself with tgkill(getpid(), gettid(), sig), which is how Go's
runtime preempts a goroutine and how glibc's pthread_kill and raise reach a
thread, named numbers no kernel thread has and was answered ESRCH.

Both of tgkill's pid arguments are now mapped, and one outside the family
is ESRCH, as for kill.
A sixteenth part: a caught SIGUSR1 sent with tgkill(getpid(), gettid())
runs its handler, and a thread the family does not have is ESRCH. Passes
on host Linux.
cwait stopped at its first failing part, so a run with more than one
change removed showed only the first. Every part now runs, each failure is
printed, and the last line counts them. A hang still stops it at the part
it hangs in.
@eKisNonos eKisNonos changed the title linux: run Go and musl threads in a guest, and end the guest whole linux: run Go and musl threads in a guest, let them wait as on Linux, and end the guest whole Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants