Conversation
Reads the signed marketplace index, lists what the running image can install, and installs through the queue init already owns. The install buttons did nothing: the painter and the hit test each computed their own rectangles, so a click never matched a row. Both derive from one geometry module now and the click path reaches the same install::ask as Enter. Search filters as typed, the list scrolls with a real scrollbar, and selection survives a refresh. init gains an install queue and a wake path so a store request is serviced on arrival rather than on the next supervisor spin. The surface registry can attach frames to a surface a capsule owns. Scrollbar arithmetic mirrored in python: empty, shorter than the viewport, thumb at both ends, single row overflow.
Release encode, decode and signing move into marketplace_abi, so the tool that writes the index and the capsule that reads it share a codec instead of having one each. install_ready checks arch and readiness against the running image rather than the index, and adds a seventh gate for whether a release carries a zk trailer for its own measurement. The other six are signatures over the artifact. The catalogue generator reads what is on disk and writes JSON; the CLI encodes the binary the market ingests. Signing is a separate step with the operator seed, which is not in any build rule.
IconId::Store points table.rs at assets/icons/store.a8, which only existed on the app-store branch, so every build of this branch failed to read it. Add the mask and its SVG source here so the table stands on its own.
text::line returns the drawn width, so the bare match evaluated to i32 where the function body expects (), failing the capsule build.
nonos-data/marketplace/index.bin has no make rule, so naming it as a hard prerequisite failed every build on a checkout without it (CI: No rule to make target). Wrapping it in $(wildcard) keeps the rebuild on a newer catalogue where it exists and drops the prerequisite where it does not.
The branch's manifest predated the switch to in-process Ed25519 and dropped the dependency while verify/crypto.rs imports it, so the capsule failed with an unresolved import. Restore main's manifest and add only the app_skeleton dependency boot_index.rs needs.
Twenty-five syscalls were published as caps = ["valid_token"] while the cap table demands a hardware, dev-root or time capability. Twenty-two take any one of Admin or a hardware capability, now published with caps_any; the dev-root and time calls need one capability, published with caps.
The cap table gates MDRO with MDRQ and MDRC on can_enrol_dev_root. The old syscall caps check cannot resolve that predicate and wants valid_token, so this fails it until abi/caps-check-fail-closed lands; the fixed check passes.
Every lane was pinned to -accel hvf -cpu host, which only macOS has, so no Linux host could boot an image. KVM when /dev/kvm opens read-write, hvf on macOS, TCG otherwise, the rule the boot matrix already uses; the display and audio backends follow the host too.
The trailer's magic picked the verifier for every root, so a Pedersen trailer was checked against the vendor root too. A local root's leaf is a commitment to a secret this kernel holds; the vendor root's is not, so there the trailer may no longer choose the weaker proof.
MkLocalSign let a LocalSign holder prove any capability it held, including LocalSign itself, so one signer could hand out the right to sign. A local proof now names nothing beyond AMBIENT_CAPS. crypto_proofs checks a minted trailer and its refusals against the kernel's own verifier files.
An Alpine package index is signed that way, and the capsule offered SHA-1 only without the prefix, so the index could not be checked at all. The rsa crate rebuilds the whole padded block and compares it. Nothing here signs, so offering SHA-1 verification mints nothing new with it.
resolve.rs names super::root, which the crate never mounted, so it did not build and nothing noticed because no workflow ran it. It mounts root.rs, follows the rename of absolute to visible, adds the /linux confinement and Phdr::file_range tests, and joins the proof-crate matrix.
A signature over a package covers the compressed bytes of one member, so a verifier needs to know where each one starts and ends. members() refuses a file with any byte that belongs to no member, where gunzip ends the stream and ignores the rest.
The index's .SIGN.RSA entry is checked against Alpine's x86_64 keys by the crypto service, the package's control member against the index's C: SHA-1, and its data against the control member's datahash. Only a Verified value reaches the store, so unauthenticated bytes are refused, not kept unvouched.
Without the operator-key rotation: NOX_OPERATOR_V1 keeps main's key, now read from .keys/marketplace_operator_ed25519.pub. The capsule embedded an index nothing built, so tools/nonos-market-index writes one, signed and verified when the operator seed is present and empty otherwise. nonos-mk moves to 78dae45 for the zk_trailer_hash field; the icon table is 49 long.
A release naming x86_64-linux counted as having its attestation, so any release could claim the exemption for itself. It now needs the linux. namespace too, which is where the store routes it: to the installer that authenticates the bytes before the machine mints their proof.
Conflicts were both sides adding: the install and app-store capsules sit together in Cargo.toml, mk and userspace, and init/mod.rs names the install queue once. The init loop and install queue are taken as the PR wrote them; the commit after this replaces its wake path.
The scheduler takes every ready process's priority lock from the timer interrupt. wake.rs and boost_init_for_drain took init's with interrupts on from syscall context, so a tick inside either spun forever on one CPU. The install queue now raises through the guarded setter the window queue uses.
MkAppInstall took a package name and the store's own readiness flag. It now takes a listing and release; init asks the market for readiness and the release's package hash, and the installer refuses bytes of any other BLAKE3. The index keeps each record's D: and p: lines, so dependencies come too.
…each CryptoMachineKey derives the machine key for any label a Crypto holder names, so a key the kernel keeps for itself needs a label no syscall can ask for. Kernel labels start with a zero byte, and the syscall refuses any label that does.
The local signing identity was random each boot, so consent was too. It is now derived from the machine key, and first-boot setup, which alone holds EnrolDevRoot, grants the local root as a named step and keeps a token only this machine can make; later boots restore it. The desktop profile now includes setup and the market, and builds every capsule it embeds.
MkAppLaunch queues a run for init, which spawns the personality to start the program the installer recorded outside /linux. Whether it may start is the exec gate's answer. The store drops its console-code enrolment, which it never held the capability for, and gains an o key to open.
A guest thread's exit was answered with zero, so the thread ran on past it. musl loops on exit, and its second call parked the thread for good, never reaped. Go's exitThread falls into INT3 when exit returns, which the fault report now turns into the end of the whole process. The exit is now Linux's. The thread is left unanswered and killed, so it never returns. The word it named for clearing, through CLONE_CHILD_CLEARTID or set_tid_address, is zeroed and one waiter woken. musl names its thread-list lock there, and pthread_join waits for that lock to be released, so every join hung before. set_tid_address now records the word and still answers with the tid. A refused kill is named in the log.
cthreads creates 8 musl pthreads, sums known ranges under a mutex, joins them all and checks the total. pthread_create exercises the reserved stack committed by mprotect and the tid written by clone, the mutex the futex, and each join the clear-tid word zeroed and woken at the thread's exit.
A pipe reported readable and writable to poll, epoll and select whatever it held, a read of an empty pipe whose writers had all gone answered EAGAIN, and a write that no reader could ever take was kept. A program that waits for a pipe to become readable before reading it either spun or never saw end of file. The family now notes, each time it lends its pipes to the guest being answered, which ends of each pipe are still open anywhere in the family, since a fork leaves them in more than one process. From that note, as Linux's pipe_poll: a read end is readable while it holds bytes and hung up once no write end is left; a write end is writable while there is room and in error once no read end is left. An empty pipe with no write end reads as end of file, and a write with no read end is refused with EPIPE. poll, epoll and select report hang-up and error whether they were asked for or not.
fcntl accepted F_SETFL and dropped it, and F_GETFL answered O_RDWR for every descriptor. pipe2 ignored O_NONBLOCK. So a read of an empty pipe parked its thread even when the program had asked never to wait, which is how Go's runtime sets every pipe it gets before handing it to its poller. A descriptor now keeps O_NONBLOCK: F_SETFL sets or clears it, pipe2 sets it on both ends, dup and fork carry it. F_GETFL answers the access mode (a pipe's read end read-only, its write end write-only, as pipe2 makes them) with O_NONBLOCK when it is set. A non-blocking read of an empty pipe is answered EAGAIN at once, or end of file once no write end is left. The other status flags F_SETFL takes are accepted and still have no effect.
eventfd2 was not served, and Go's runtime makes one for its poller at the first timer and throws when it cannot, so every Go program that sleeps or sets a timer died there. eventfd2 and eventfd now make a counter. The counter is the object and the descriptor only names it, so it is kept with the pipes as the family's, lent to the guest being answered: dup, fork and every thread reach the same one, as on Linux. A read takes the whole count, or one under EFD_SEMAPHORE; a write adds to it; a starting value is an unsigned int; a write of all ones and a buffer under eight bytes are refused with EINVAL. Where Linux would wait (a read at zero, a write that would pass the most a counter holds), the call answers EAGAIN for now. poll, epoll and select see it readable while the count is above zero and writable while one more fits. EFD_CLOEXEC and EFD_NONBLOCK set the descriptor's flags; any other flag is refused with EINVAL.
epoll_wait and epoll_pwait answered at once whatever their timeout, so a program waiting for a descriptor or a deadline spun, and one waiting with no timeout spun for good. Go's runtime waits there for its next timer, and wakes that wait from another thread by writing its eventfd. A read of an empty eventfd and a write to a full pipe answered EAGAIN where Linux waits. A call that has to wait is now left parked in its trap and tried again after every answer the family gives and at its deadline, so a write from another thread or process completes the wait at once. epoll_wait takes its timeout as an int of milliseconds: zero reports what is ready now, a negative one waits until something is, and a deadline that passes answers zero. A maxevents of zero or less, or past Linux's bound, is refused with EINVAL. A read or write of an eventfd, and a write to a full pipe, wait the same way unless the descriptor is non-blocking. A pipe write of up to PIPE_BUF bytes goes in whole or waits, as Linux's does. A wait on a socket or a timer is looked at again every 10 ms, since nothing the family answers changes them. The loop wakes for the nearest deadline. A killed thread's parked call is dropped, so it cannot take what a live one waits for.
Every epoll entry was reported for as long as its descriptor was ready, whatever it asked for. Go registers every pipe and socket with EPOLLET and EPOLLOUT, and a write end with room is always writable, so once epoll_wait waits, Go's poller would be answered at once on every call and spin. epoll_ctl accepted a regular file, adding twice, and changing or dropping an entry that was not there. An EPOLLET entry is now reported when its readiness rises: bits already seen at the last look are left out until they fall, or until a read, write, send, receive or accept on the descriptor answers EAGAIN, or a connect EINPROGRESS, which is where a program using EPOLLET stops and waits. An EPOLLONESHOT entry reports once and then nothing until it is modified. What was seen is kept only once the events are delivered. epoll_ctl refuses as Linux does: EBADF for a closed descriptor, EPERM for a regular file or directory (Go's os.Open then reads it blocking), EINVAL for the list itself, EEXIST for adding twice, ENOENT for changing or dropping what is not there. A change looks at the entry afresh.
A futex wait ignored its timeout and waited until it was woken, and FUTEX_WAIT_BITSET, FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE answered ENOSYS. musl's pthread_cond_timedwait, sem_timedwait and pthread_mutex_timedlock wait with a timeout, musl hands a condition variable's waiters to its mutex with FUTEX_REQUEUE, and Go's sysmon sleeps on a timed futex. A value compared against the word was also taken as 64 bits, so musl's sign-extended int of -1 never matched. A timed wait now ends with ETIMEDOUT at its deadline, unless something wakes it first, and at once if the deadline has already passed. FUTEX_WAIT's timeout is relative; FUTEX_WAIT_BITSET's is absolute on CLOCK_MONOTONIC, or on CLOCK_REALTIME with FUTEX_CLOCK_REALTIME. A timespec with negative seconds or nanoseconds past a second is EINVAL. The bitset forms are served with the bitset taken as matching every waiter, which at worst is a spurious wake, and a zero bitset is EINVAL. FUTEX_REQUEUE wakes up to val waiters and moves up to val2 more to the second word; FUTEX_CMP_REQUEUE first checks the word against val3. The value is compared as 32 bits. The loop wakes for the nearest futex deadline, and a thread being ended has all its waits dropped.
gopoll is Go's network poller in three parts, each printed as it passes: a 50 ms sleep that needs epoll_pwait to wait its timeout, a 10 ms timer set while the poller already waits on a 3 s one, which needs Go's eventfd write from another thread to end that wait, and an os.Pipe read through the poller to end of file, which needs O_NONBLOCK, edge-triggered readiness and a hang-up once the write end closes. cwait is the same for musl in nine parts: cond_timedwait's futex timeout, a condition variable broadcast, eventfd counts and semaphore reads, a blocking eventfd read woken by another thread, epoll_wait's timeout and its wake from another thread, a non-blocking pipe with fcntl and end of file, a write to a full pipe waiting for a reader, and edge-triggered epoll. Both pass on host Linux, which is what they measure the personality against.
A read parked on an empty pipe was kept in one slot per process, so a second thread reading an empty pipe took the slot and the first was never answered. A read of zero bytes from an empty pipe with a writer answered EAGAIN rather than zero. A read of an empty pipe now waits with every other parked call, tried again after each answer until it gets bytes, or end of file once no write end is left anywhere in the family. The slot and its settling are gone. The family settles waits after it reaps, so a write end that left with its process reads as end of file at once. A read of zero bytes answers zero.
The four answered at once whatever their timeout, so a program waiting on descriptors spun, and poll(NULL, 0, ms), the idiom for a short sleep, did not sleep. select also narrowed the caller's sets even when nothing was ready, and copied a whole 1024-bit set whatever nfds was. poll counted an entry with a negative descriptor, which Linux ignores and programs use to switch an entry off, as closed and so ready. Each now waits with the other parked calls until something it watches is ready or its timeout passes: poll's int of milliseconds, negative for none; ppoll's and pselect6's timespec and select's timeval, null for none, with negative or out-of-range fields refused with EINVAL. A timeout of zero only looks. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out, as Linux does. A negative poll descriptor is never ready. A wait watching a socket or a timer is looked at again every 10 ms, now for poll and select as for epoll.
Two parts added to cwait. Two threads block reading one empty pipe and both must be answered once bytes arrive. poll, ppoll and select each wait out a 100 ms timeout with nothing ready, poll(NULL, 0, 50) sleeps, and a select with no timeout is woken by another thread's write, with its set narrowed to the ready descriptor. Both pass on host Linux.
…aces A closed descriptor stayed in every epoll interest list. Linux drops it, so a program that closes without EPOLL_CTL_DEL and gets the same number back from its next open adds it again; here that add was refused with EEXIST, and the old entry reported the new file under the old token. dup2 onto an open descriptor overwrote it without closing it, so a file's buffered bytes were lost and a socket's handle kept, and a target number past the table grew the table to reach it. close now drops the descriptor from the guest's interest lists. dup2 closes an open target first, as Linux does, and refuses a number past the table with EBADF.
A twelfth part: a watched pipe end is closed without EPOLL_CTL_DEL, a new pipe takes its number, and adding it again succeeds with no stale event. dup2 to a number past the table is EBADF. Passes on host Linux.
A timerfd fired once and forgot its interval, so a periodic timer stopped after its first read, and a read always answered one. timerfd_create ignored its clock and its TFD_NONBLOCK and TFD_CLOEXEC flags, timerfd_settime ignored TFD_TIMER_ABSTIME and never wrote the old setting, timerfd_gettime was not served, and a read of a timer that had not fired answered EAGAIN even to a blocking descriptor. The expiry lived in the descriptor, so dup and fork gave a timer that never fired. A timer is now an object the family keeps, lent with the pipes and eventfd counters, so every descriptor onto it sees one timer. It is made on a clock Linux names (anything else is EINVAL) with its flags; set relative or absolute on its own clock, one-shot or periodic, with the old setting written when asked; read for how many times it has fired since the last read, stepping a periodic timer past them; read blocking until it fires unless the descriptor is non-blocking; and reported readable once it has fired. A wait watching a timer is woken when it fires rather than on a 10 ms look.
A thirteenth part: an unarmed non-blocking timer reads EAGAIN, a 50 ms periodic timer counts its firings over 180 ms and reports its interval, a one-shot read blocks until it fires, and an absolute time 100 ms ahead wakes an epoll wait. Passes on host Linux.
gopreempt spins a goroutine with no call in it beside main, on one P. Go moves such a goroutine off the CPU only by sending SIGURG to its thread while it runs, so main waking from a 20 ms sleep and a garbage collection, which stops the world, both depend on that signal landing. On host Linux main runs again after 25 ms and the collection ends by 45 ms, with the tgkill deliveries visible under strace.
ioctl answered ENOTTY to every request, so a program asking how many bytes a pipe or file held, setting a descriptor non-blocking the BSD way, or marking it close-on-exec without fcntl was refused. FIONREAD answers what a pipe's read end holds, or what is left of a file past its offset; FIONBIO sets or clears O_NONBLOCK from the int it is given; FIOCLEX and FIONCLEX set and clear close-on-exec. A socket's FIONREAD, and every device request, is still ENOTTY: there is no terminal or device behind a descriptor here, which is also how isatty says no.
A duplicated or inherited epoll descriptor started with an empty interest list, so a child that waited on the epoll its parent set up before fork waited on nothing. The list is now copied with the descriptor. Linux shares one list between them; the copy holds what was registered at the dup or fork, and a change made afterwards on one side is not seen on the other.
A fourteenth part: FIONREAD counts three bytes in a pipe, FIONBIO makes its read end answer EAGAIN once drained, FIOCLEX sets close-on-exec, and a forked child finds a ready entry in the epoll list its parent built. Passes on host Linux.
sched_getscheduler, sched_setscheduler, sched_getparam, sched_setparam, sched_get_priority_max and _min and sched_setaffinity were not served, nor epoll_create and epoll_pwait2, so a program that sets a thread's policy or pins it, or a libc that reaches for the older or newer epoll form, died there. They are answered as Linux answers an unprivileged process on the one CPU the guest is shown: every thread is SCHED_OTHER at priority zero; asking for SCHED_OTHER, BATCH or IDLE at zero is accepted and changes nothing, since scheduling is the kernel's; asking for FIFO or RR is EINVAL outside the priority range 1 to 99 and EPERM inside it, the range checked first as Linux checks it, since no guest holds the privilege Linux asks for; the priority ranges are Linux's. An affinity mask that includes the one CPU is accepted and one that leaves it out is EINVAL. A pid argument is mapped from the guest's namespace like kill's, and one outside the family is ESRCH. epoll_create checks its size is positive and makes a list; epoll_pwait2 waits like epoll_pwait with a timespec timeout.
A fifteenth part: the policy is SCHED_OTHER, SCHED_FIFO at priority zero is EINVAL, SCHED_OTHER is accepted, the real-time range is 1 to 99, CPU 0 can be pinned, epoll_create refuses a size of zero, and epoll_pwait2 waits its 50 ms timespec. SCHED_FIFO at priority one and a CPU-1-only mask depend on privilege and CPU count, so they are printed, not checked. Passes on host Linux.
gettid and getpid answer with the family's own numbers, and kill and tkill map theirs back, but tgkill's were passed through as they came. A thread signalling itself with tgkill(getpid(), gettid(), sig), which is how Go's runtime preempts a goroutine and how glibc's pthread_kill and raise reach a thread, named numbers no kernel thread has and was answered ESRCH. Both of tgkill's pid arguments are now mapped, and one outside the family is ESRCH, as for kill.
A sixteenth part: a caught SIGUSR1 sent with tgkill(getpid(), gettid()) runs its handler, and a thread the family does not have is ESRCH. Passes on host Linux.
cwait stopped at its first failing part, so a run with more than one change removed showed only the first. Every part now runs, each failure is printed, and the last line counts them. A hang still stops it at the part it hangs in.
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Targets
main. It also carries the 209 commits of #567 (linux/store-install), which come before it and hold the Linux personality these commits change, so it merges after #567 and shows those commits until #567 lands. What it adds is thirty-six commits.Before this, a Linux-guest test image started no guest at all. Once it did, no guest could run a second thread: Go's threads called address zero, musl's
pthread_createfailed outright, and a thread's exit returned to it. A guest that did get threads reset the machine when it exited, left its threads running, and hung if one of them faulted. And nothing could wait:eventfd2was not served, so every Go program with a timer died at its first one, whileepoll_waitand a futex with a timeout returned at once or never.What Linux guests can do now
time.Sleep, a timer woken early through Go's eventfd, and a pipe read through the poller to end of file.epoll_wait,pollandselecttimeouts, non-blocking pipes, several readers or a writer waiting on one pipe, edge-triggered epoll, and timerfd one-shot, periodic and absolute.epoll_create,epoll_pwait2andtgkillare served as Linux serves them to an unprivileged process on one CPU.What changed
The boot guest starts
bf9d9dbf3: the Linux-guest test image (NONOS_LINUX_GUESTS=1) buildsmicrokernel-desktop-guiwithnonos-stark-attest, without first-boot setup. Under the setup profile every app, the personality included, waits for setup to exit, and setup waits for keys, so the unattended image never started its guest: 0[APP-LINUX]or[LINUX]lines in 180 s. Every other build ofnonos-mk-desktop-gui-prodis unchanged.Threads start, run and end
3fc1ac19d: a clone child starts on a copy of the calling thread's registers, with rax 0, the new stack, and the parent's thread pointer unless a TLS is named.MkForeignThreadtakes the calling thread as a fifth argument; it must be supervised by the caller and in the same thread group as the target, so no guest receives another's registers. Zero keeps the fresh start. Go's child calls r12 and musl's calls r9, so both called address 0 before. A TLS is taken only whenCLONE_SETTLSis set.abi/syscalls.tomlrecords the new argument.addec49af: clone writes the new tid whereCLONE_PARENT_SETTIDandCLONE_CHILD_SETTIDask. musl's thread-list lock stores its owner's tid, and a thread that never learned its tid locked the list as 0, which the lock reads as free.1e6b8a673:mprotectcommits a PROT_NONE reservation it is asked to open. musl reserves every pthread stack with PROT_NONE and opens the usable part withmprotect;MkPeerProtectreprotects only pages that exist, sopthread_createfailed. A backed region is reprotected as before; a reservation is backed with the asked protection, keeping any page the guest already touched, and recorded as backed so fork copies it. A span outside every region is refused with ENOMEM, as on Linux.10947aa96: a thread'sexitends the thread and never returns to it. The word it named throughCLONE_CHILD_CLEARTIDorset_tid_addressis zeroed and one waiter woken, as Linux does. musl names its thread-list lock there andpthread_joinwaits for that lock, so every join hung before. Go'sexitThreadfalls into INT3 whenexitreturns, which the fault report below would have turned into the end of the process.A process ends whole
5ab0a0912: a thread group's page tables stay until its last member leaves the process table. Release freed them when the leader was finalized, under a Go thread still running on them, and the machine reset: CR3 was the guest's, FS held the Go thread's TLS, CR2 was the IDT entry for vector 8. The tables now pass to a process still running on them, whose capability token is rebound to the ASID it now owns.874f3ecfc:MkKilladmits the caller the foreign registry names as the target's supervisor. A thread's parent is its leader, not the supervisor, so every thread kill at a guest's exit returned EPERM and the threads ran on. A refused kill now prints[LINUX] kill refused: pid <n> outlives its process, errno <e>.2311afc39: when a guest thread ends on a signal, and only when it ends itself so a supervisor's own kill cannot loop back, the kernel leaves a one-way notice for its supervisor, collected on its nextMkForeignWaitas a frame numberedFOREIGN_NR_DIED. The kernel reports; what a dead thread means is the personality's policy.959b57f98: the personality ends the process on that notice, as Linux ends a thread group on an unhandled fault, and reports the signal as 128+signo.Fork
ffc1b542b: a fork copies the calling thread's parked frame, and the kernel carries that thread's own thread pointer to the child. Before, every fork took the leader's registers and the personality's singlefs_base, which is only the last thread to set one.Waiting, as Linux waits
69f3b48d4: a pipe has Linux's ends. The family notes, each time it lends its pipes to the guest being answered, which ends are still open anywhere in the family. From that note: a read end is readable while it holds bytes and hung up once no write end is left; a write end is writable while there is room and in error once no read end is left. An empty pipe with no write end reads as end of file, and a write no reader can take is refused with EPIPE. poll, epoll and select report hang-up and error unasked. Before, every pipe was readable and writable to all three.3e48edbaf: a descriptor keeps O_NONBLOCK. F_SETFL sets it, F_GETFL reports it with the access mode, pipe2 takes it, dup and fork carry it, and a non-blocking read of an empty pipe answers EAGAIN. Before, F_SETFL was dropped, which is how Go's runtime prepares every pipe for its poller.bfb07892f: eventfd2 and eventfd are served. The counter is kept with the pipes as the family's, so dup, fork and every thread reach one. EFD_SEMAPHORE, EFD_NONBLOCK and EFD_CLOEXEC as Linux; all ones and a short buffer EINVAL. Go's runtime makes one at its first timer and threw when it could not.06a266ac5: epoll_wait and epoll_pwait wait for their timeout (an int of ms; negative waits until ready). A blocking eventfd read or write and a write to a full pipe wait too. A waiting call is left parked in its trap and tried again after every answer and at its deadline, so a write from another thread ends the wait at once. A wait on a socket or timer is looked at every 10 ms. A pipe write of up to PIPE_BUF goes in whole or waits. maxevents of zero or less is EINVAL.32d2a1a06: EPOLLET entries are reported as readiness rises and re-armed by an EAGAIN; EPOLLONESHOT reports once. epoll_ctl refuses as Linux: EBADF, EPERM for a regular file or directory, EINVAL, EEXIST, ENOENT. Go registers every pipe and socket with EPOLLET and EPOLLOUT, so a level-triggered report of an always-writable end would answer its poller at once, every time.aa0087d75: a futex wait ends with ETIMEDOUT at its timeout. FUTEX_WAIT_BITSET (absolute, MONOTONIC or REALTIME), FUTEX_WAKE_BITSET, FUTEX_REQUEUE and FUTEX_CMP_REQUEUE are served. The value is compared as 32 bits, so musl's sign-extended -1 matches. musl's timed waits and condition variables, and Go's sysmon, use these.b1ead15c5: two proof guests, gopoll and cwait (below).More waiting, timers, descriptors and threads
b716a34af: any number of threads can wait to read one pipe. A read parked in one slot per process, so a second reader of an empty pipe took the slot and the first was never answered. Pipe reads now wait with the other parked calls; the family settles them after it reaps, so a writer that left with its process reads as end of file at once. A zero-byte read answers zero.439ed66c6:poll,ppoll,selectandpselect6wait for their timeout (poll's int of ms, ppoll's and pselect6's timespec, select's timeval; null or negative waits until ready; bad fields EINVAL).poll(NULL, 0, ms)sleeps. select writes its sets back only when something is ready, copies only the longs nfds covers, always empties the exception set, and empties all three when its time runs out. A negative poll descriptor is never ready.85b3d7724: a closed descriptor leaves every epoll interest list, as on Linux, so its number can be added again; before, that add was EEXIST and the old entry reported the new file under the old token.dup2closes an open target first (a file's buffered bytes were lost, a socket's handle kept) and refuses a number past the table.b9efc7f96: timerfd has Linux's semantics and is kept with the family like the eventfd counter, so dup and fork share it. A periodic timer keeps firing and a read returns how many times it fired; the clock is Linux's set (else EINVAL); TFD_NONBLOCK, TFD_CLOEXEC and TFD_TIMER_ABSTIME are honoured; the old setting is written when asked;timerfd_gettimeis served; a blocking read waits for the firing; a wait watching a timer is woken when it fires.262c06280: FIONREAD (a pipe's bytes, a file's remainder), FIONBIO, FIOCLEX and FIONCLEX. Any other request is ENOTTY, as for a descriptor with no terminal or device behind it.8f778075b: a dup'd or inherited epoll descriptor carries the interest list it had; before, it started empty.ff778c704:sched_getscheduler,sched_setscheduler,sched_getparam,sched_setparam,sched_get_priority_maxand_min,sched_setaffinity,epoll_create,epoll_pwait2. Every thread is SCHED_OTHER at zero; SCHED_FIFO and SCHED_RR are EINVAL outside 1 to 99 and EPERM inside it, the range checked first as Linux checks it; a mask without the one CPU is EINVAL; pid arguments are mapped like kill's.9039b893d: tgkill's thread group and thread are mapped into the family's pids. gettid and getpid answer with the family's numbers, sotgkill(getpid(), gettid(), sig), which Go's runtime and glibc'spthread_killuse, was ESRCH.84e10f538,7c3967f18,468f5aa97,2f0db3b25,a045bec06,66b55faf8,afd96ef58,24cafe103: the guest gopreempt, and seven more cwait parts, one per change above; cwait now runs every part and names each one that fails.Cleanups
289272374: deletecall/clone.rs, a copy left behind when clone moved tocall/spawn/; no module included it.ce395f7ff: the[PF] demand fillline is printed only after a fill, not before the handler refuses the null page or the kernel half.Evidence
q35 under TCG, one vCPU. Each guest is the boot guest of its own store.
Threads, faults and exit (the first thirteen commits)
[GO] hello PASS[GO] conc PASS: 8 goroutines summed 3199960000[C] cthreads PASS: 8 pthreads joined, summed 3199960000[LINUX] guest thread 51 ended on a signal; ending the process, then[LINUX] guest exitedthreadfault's worker runs on a stack mapped read-write outright through
clone(), so it does not depend on the pthread path. ItsSURVIVEDline, printed only if main outlives the faulting thread, never appears.Waiting (the next seven commits)
Both guests pass on host Linux, which is what they are measured against. On one build carrying all twenty commits:
[GO] poll PASS: 3 parts, 3454ms in all: slept 69 ms, the early timer ended in 18 ms, the pipe read to end of file in 339 ms[C] cwait PASS: 9 parts in 951 ms: cond_timedwait ETIMEDOUT after 102 ms, a blocking eventfd read woken after 147 ms, epoll_wait timed out after 100 ms and woken after 153 ms, POLLHUP on a closed pipe, a full-pipe write waited 148 ms, edge-triggered counts as LinuxThe same build ran the first four guests again: gohello 283 calls, goconc 253, cthreads 87, threadfault 7 with its death line; 0 unserved in each.
More waiting, timers, descriptors and threads (the next sixteen commits)
On one build carrying all thirty-six commits:
[C] cwait PASS: 16 parts in 2313 ms; SCHED_FIFO at priority 1 errno 1 and a CPU-1-only mask errno 22; a 50 ms periodic timer fired 3 times in 180 ms, a one-shot read waited 103 ms, an absolute timer woke epoll after 105 ms[GO] poll PASS: 3 parts, 3437ms in allThe same build ran the regression set: gohello 282 calls, goconc 248, cthreads 87, threadfault 7 with its death line; 0 unserved in each.
Each fix was also run with its change removed, on the same guest:
MkKillkill refused ... errno 1lines; guest threads fault afterguest exitedSURVIVEDprinted, no death lineSURVIVEDmprotectcommit (the pthread build of threadfault, before the fix)[C] threadfault FAIL: no workerpthread_createsucceeds (cthreads)a 10ms timer set under a 3s one took 2823ms, FAILeventfd2served[LINUX] unserved nr=290, thenfatal error: runtime: eventfd failedFAIL: edge-triggered counts (222, 1): the second look counts 2, not 1FAIL: end of file after the write end closed (0, 0): poll sees nothingFAIL: 6 parts failed, 10 passed: poll and ppoll returned after 2 ms, the re-add was EEXIST, the periodic timer counted 1, FIONREAD failed,[LINUX] unserved nr=145, tgkill ESRCHFAIL: epoll list through fork (13, 256): the child finds nothing ready and exits 1Each of these ran on its own build with only that change removed, except the six that fail different cwait parts, which were removed together. gopoll passes with the hang-up removed, as its writer closes before its reader waits, and with O_NONBLOCK dropped, as Go's pipe reads then park instead of polling.
The store settled in 12.8 to 84.2 s across the 22 guest boots for the first thirteen commits, in 38.0 to 104.8 s across the 16 for the next seven, and in 17.2 to 117.7 s across the 10 for the sixteen after; the personality waits up to 300 s.
Checks
check_stubs,check_allows,check_dark_features: 0 new sites.check_syscall_abi: 111 published syscalls reach a handler.check_unreachable: 1 new site,has_children, whichlinux/store-installreports too; this branch has 1257 sites against its 1258.microkernel-core) and the personality were checked at every one, with 0 errors. The next twenty-three change only the personality and the guests; the personality was checked at each with 0 errors and 0 warnings, and each guest passes on host Linux.Not done here
mprotecton a backed region does not update the recorded protection, so a fork after it gives the child the protection the region was mapped with.exit(notexit_group) still ends the whole process; on Linux the others would run on.killto a child process is queued in the sender, not the child, so the child's handler never runs; an uncaught one ends the child as it should.