From cf2a8d9114d1a0368b59f5598ebc2451489c7617 Mon Sep 17 00:00:00 2001 From: Bruno Herrera Date: Sun, 20 Sep 2026 23:08:41 -0300 Subject: [PATCH 1/6] docs: propose LKMM validation for printk payload accesses Record a pending exception for investigating the printk ringbuffer port after Loom rejects its speculative payload reads. Separate the state protocol model from the C access and compiler contracts without claiming that an atomic control validates ordinary C payload accesses. Keep the proposal inactive until accepted. The existing stop criteria continue to apply to production ports. Assisted-by: LLM [Codex] --- AGENTS.md | 36 ++++++++++++++++++++++++++++++++++++ 1 file changed, 36 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index 1db99ef3c42faa..c9c65a7e61ab9b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -308,3 +308,39 @@ Project decision: commits in this tree use kernel style. If agents keep making the same mistake, or a decision changes, propose an edit here. Do not add workarounds in code. + +### Proposed: printk LKMM access validation (pending acceptance) + +This proposal is not an active exception to the rules above. The printk +ringbuffer reads ordinary payload memory speculatively and validates the +descriptor afterward. A direct Rust translation fails Loom's data-race +checks, including when descriptor operations are sequentially consistent. +The in-tree `rust/kernel/sync/atomic.rs` already distinguishes LKMM from +the userspace Rust memory model, but does not supply a contract for these +bulk copies. + +For this port, investigate a split validation approach: + +- Keep all ringbuffer decisions, reservation, recycling and descriptor + validation in Rust. Keep headers and C callers unchanged. +- Access intentionally racing C-owned memory only through a minimal C + boundary. Each helper must preserve an existing C access operation and + its barrier placement, with a documented U1 contract. This is not + permission to retain the C algorithm behind a wrapper. +- Never form Rust references to concurrently recycled payloads. Copies + returned to Rust remain untrusted until descriptor validation succeeds; + lengths, offsets and bit validity must be checked before use. +- Use Loom for protocols it can faithfully model. Preserve the failing + direct-translation diagnostic. An atomic payload model is a control, + not evidence that C's ordinary payload accesses are race-free. +- Use LKMM/herd7 litmus tests for the C boundary's publication and reuse + ordering. Derive them from the source's named LMM barrier pairs and + document which properties and accesses the model does not cover. +- Require a compiler/FFI argument for the boundary as well as layout, + differential, existing KUnit and C/Rust build validation. Passing herd7 + alone does not establish Rust soundness or compiler correctness. + +If accepted, criterion 2 would allow this explicit combination of Loom +and LKMM evidence for printk. Failure to establish the C/Rust boundary +contract still stops the port; acceptance is permission to investigate, +not certification of the implementation. From 51f59c06d8fc4f8d254a205743d3acbcbfed2636 Mon Sep 17 00:00:00 2001 From: Bruno Herrera Date: Sun, 20 Sep 2026 23:33:14 -0300 Subject: [PATCH 2/6] docs: clarify acceptance and constraints of printk proposal Define merging PR #24 as acceptance of the printk-specific criterion-2 amendment. Preserve the single-forwarding-call FFI restriction and state that the existing stop criterion remains tripped until acceptance. Cite the local diagnostic commit and distinguish its minimal race model from complete ringbuffer certification. Assisted-by: LLM [Codex] --- AGENTS.md | 29 +++++++++++++++++++++++------ 1 file changed, 23 insertions(+), 6 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index c9c65a7e61ab9b..8a408b5b032f42 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -311,14 +311,25 @@ here. Do not add workarounds in code. ### Proposed: printk LKMM access validation (pending acceptance) -This proposal is not an active exception to the rules above. The printk -ringbuffer reads ordinary payload memory speculatively and validates the -descriptor afterward. A direct Rust translation fails Loom's data-race +Merging PR #24 into `linux-rust` accepts and activates this scoped amendment +to criterion 2. Until then, it is not an active exception: printk port +reports must state that criterion 2 is tripped and the port is stopped. +This proposal belongs here because it changes a project decision. + +The printk ringbuffer reads ordinary payload memory speculatively and +validates the descriptor afterward. A direct Rust translation fails Loom's data-race checks, including when descriptor operations are sequentially consistent. The in-tree `rust/kernel/sync/atomic.rs` already distinguishes LKMM from the userspace Rust memory model, but does not supply a contract for these bulk copies. +Evidence: local, unpushed commit `818c24e` in `misttech/linux-rust`, on +`codex/gpt-6/feature/printk-port`, contains the reproducer at +`harness/printk_ringbuffer/loom/` and the record at +`units/printk_ringbuffer.md`. The speculative-copy test fails; the atomic +payload control passes. This is a minimal feasibility model, not a model +of the complete ringbuffer. The commit is not yet available on GitHub. + For this port, investigate a split validation approach: - Keep all ringbuffer decisions, reservation, recycling and descriptor @@ -327,6 +338,11 @@ For this port, investigate a split validation approach: boundary. Each helper must preserve an existing C access operation and its barrier placement, with a documented U1 contract. This is not permission to retain the C algorithm behind a wrapper. + The existing FFI rule remains binding: each helper is a single forwarding + call to an existing C function, inline or macro, with its prototype above + the definition. No helper loops, branches or new access algorithms are + authorized. If this boundary requires more, stop and propose a separate + amendment to the FFI rule. - Never form Rust references to concurrently recycled payloads. Copies returned to Rust remain untrusted until descriptor validation succeeds; lengths, offsets and bit validity must be checked before use. @@ -340,7 +356,8 @@ For this port, investigate a split validation approach: differential, existing KUnit and C/Rust build validation. Passing herd7 alone does not establish Rust soundness or compiler correctness. -If accepted, criterion 2 would allow this explicit combination of Loom -and LKMM evidence for printk. Failure to establish the C/Rust boundary -contract still stops the port; acceptance is permission to investigate, +Upon acceptance by merging PR #24, criterion 2 allows this explicit +combination of Loom and LKMM evidence for printk only. All other ports +remain subject to the original criterion. Failure to establish the C/Rust +boundary contract still stops the port; acceptance is permission to investigate, not certification of the implementation. From 688cad45bb7b540589ef29f77707455b3edc64df Mon Sep 17 00:00:00 2001 From: Bruno Herrera Date: Mon, 21 Sep 2026 09:20:49 -0300 Subject: [PATCH 3/6] docs: record authorization for printk access investigation Record the explicit authorization to implement printk using the scoped Loom and LKMM validation approach while PR #24 is under development. Keep certification separate from permission to investigate, and link the now-published diagnostic and validation evidence. Assisted-by: LLM [Codex] --- AGENTS.md | 16 ++++++++-------- 1 file changed, 8 insertions(+), 8 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 8a408b5b032f42..7f68db9ca72bd1 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -309,12 +309,11 @@ Project decision: commits in this tree use kernel style. If agents keep making the same mistake, or a decision changes, propose an edit here. Do not add workarounds in code. -### Proposed: printk LKMM access validation (pending acceptance) +### Printk LKMM access validation -Merging PR #24 into `linux-rust` accepts and activates this scoped amendment -to criterion 2. Until then, it is not an active exception: printk port -reports must state that criterion 2 is tripped and the port is stopped. -This proposal belongs here because it changes a project decision. +The user approved using this scoped amendment while implementing printk +in PR #24. This authorizes investigation and implementation with the +validation below; it does not certify the C/Rust boundary or the port. The printk ringbuffer reads ordinary payload memory speculatively and validates the descriptor afterward. A direct Rust translation fails Loom's data-race @@ -323,12 +322,13 @@ The in-tree `rust/kernel/sync/atomic.rs` already distinguishes LKMM from the userspace Rust memory model, but does not supply a contract for these bulk copies. -Evidence: local, unpushed commit `818c24e` in `misttech/linux-rust`, on +Evidence: commit `818c24e` in `misttech/linux-rust`, on `codex/gpt-6/feature/printk-port`, contains the reproducer at `harness/printk_ringbuffer/loom/` and the record at `units/printk_ringbuffer.md`. The speculative-copy test fails; the atomic payload control passes. This is a minimal feasibility model, not a model -of the complete ringbuffer. The commit is not yet available on GitHub. +of the complete ringbuffer. The diagnostic and subsequent validation +commit `3edea95` are published on that branch. For this port, investigate a split validation approach: @@ -356,7 +356,7 @@ For this port, investigate a split validation approach: differential, existing KUnit and C/Rust build validation. Passing herd7 alone does not establish Rust soundness or compiler correctness. -Upon acceptance by merging PR #24, criterion 2 allows this explicit +For the authorized printk work, criterion 2 allows this explicit combination of Loom and LKMM evidence for printk only. All other ports remain subject to the original criterion. Failure to establish the C/Rust boundary contract still stops the port; acceptance is permission to investigate, From 47c3d708e5e5c15199b12f17819ee6b56917f5b6 Mon Sep 17 00:00:00 2001 From: Bruno Herrera Date: Mon, 21 Sep 2026 09:20:49 -0300 Subject: [PATCH 4/6] printk: implement the Rust ring buffer replacement Replace the ring buffer algorithm under CONFIG_RUST_KERNEL while retaining the C header and caller ABI. Generate layout assertions from the active C configuration, including execution-context metadata. Keep reservation and recycling decisions in Rust. Forward shared data accesses, atomics and barriers through C primitives without retaining the C algorithm behind the boundary. This is an implementation checkpoint for the printk series. Sequential differential traces, two metadata layouts, C/Rust SMP KUnit runs and three LKMM barrier projections pass. The full compiler-boundary argument, source-comment audit and unsafe classification remain under review. Assisted-by: LLM [Codex] --- Kbuild | 2 +- kernel/port-layout.c | 50 + kernel/printk/Makefile | 2 +- kernel/printk/printk_ringbuffer.rs | 1508 +++++++++++++++++++++++++ kernel/printk/printk_ringbuffer_ffi.c | 182 +++ main.rs | 4 + rust/Makefile | 1 + 7 files changed, 1747 insertions(+), 2 deletions(-) create mode 100644 kernel/printk/printk_ringbuffer.rs create mode 100644 kernel/printk/printk_ringbuffer_ffi.c diff --git a/Kbuild b/Kbuild index e0ea130b11e4a6..f157d8aecdc0ec 100644 --- a/Kbuild +++ b/Kbuild @@ -50,7 +50,7 @@ $(port-layout-header): kernel/port-layout.s FORCE define filechk_port_layout_rs echo "// SPDX-License-Identifier: GPL-2.0"; \ echo "// Generated from the active C kernel configuration."; \ - awk '/^#define PORT_LOCKREF_/ { \ + awk '/^#define PORT_/ { \ type = ($$2 == "PORT_LOCKREF_DEAD_VAL" ? "i32" : "usize"); \ print "pub(crate) const " $$2 ": " type " = " $$3 ";"; \ if ($$2 == "PORT_LOCKREF_ALIGN") align = $$3; \ diff --git a/kernel/port-layout.c b/kernel/port-layout.c index f39fc5cdca4f89..15eb83c2ef4e1e 100644 --- a/kernel/port-layout.c +++ b/kernel/port-layout.c @@ -6,6 +6,7 @@ #define COMPILE_OFFSETS #include #include +#include "printk/printk_ringbuffer.h" int main(void) { @@ -19,6 +20,55 @@ int main(void) OFFSET(PORT_LOCKREF_LOCK_COUNT, lockref, lock_count); #else DEFINE(PORT_LOCKREF_LOCK_COUNT, 0); +#endif +#ifdef CONFIG_PRINTK + DEFINE(PORT_PRB_DATA_BLK_LPOS_SIZE, sizeof(struct prb_data_blk_lpos)); + DEFINE(PORT_PRB_DATA_BLK_LPOS_ALIGN, __alignof__(struct prb_data_blk_lpos)); + OFFSET(PORT_PRB_DATA_BLK_LPOS_BEGIN, prb_data_blk_lpos, begin); + OFFSET(PORT_PRB_DATA_BLK_LPOS_NEXT, prb_data_blk_lpos, next); + DEFINE(PORT_PRB_DESC_SIZE, sizeof(struct prb_desc)); + DEFINE(PORT_PRB_DESC_ALIGN, __alignof__(struct prb_desc)); + OFFSET(PORT_PRB_DESC_STATE_VAR, prb_desc, state_var); + OFFSET(PORT_PRB_DESC_TEXT_BLK_LPOS, prb_desc, text_blk_lpos); + DEFINE(PORT_PRB_DATA_RING_SIZE, sizeof(struct prb_data_ring)); + DEFINE(PORT_PRB_DATA_RING_ALIGN, __alignof__(struct prb_data_ring)); + OFFSET(PORT_PRB_DATA_RING_SIZE_BITS, prb_data_ring, size_bits); + OFFSET(PORT_PRB_DATA_RING_DATA, prb_data_ring, data); + OFFSET(PORT_PRB_DATA_RING_HEAD_LPOS, prb_data_ring, head_lpos); + OFFSET(PORT_PRB_DATA_RING_TAIL_LPOS, prb_data_ring, tail_lpos); + DEFINE(PORT_PRB_DESC_RING_SIZE, sizeof(struct prb_desc_ring)); + DEFINE(PORT_PRB_DESC_RING_ALIGN, __alignof__(struct prb_desc_ring)); + OFFSET(PORT_PRB_DESC_RING_COUNT_BITS, prb_desc_ring, count_bits); + OFFSET(PORT_PRB_DESC_RING_DESCS, prb_desc_ring, descs); + OFFSET(PORT_PRB_DESC_RING_INFOS, prb_desc_ring, infos); + OFFSET(PORT_PRB_DESC_RING_HEAD_ID, prb_desc_ring, head_id); + OFFSET(PORT_PRB_DESC_RING_TAIL_ID, prb_desc_ring, tail_id); + OFFSET(PORT_PRB_DESC_RING_LAST_FINALIZED_SEQ, prb_desc_ring, last_finalized_seq); + DEFINE(PORT_PRINTK_RINGBUFFER_SIZE, sizeof(struct printk_ringbuffer)); + DEFINE(PORT_PRINTK_RINGBUFFER_ALIGN, __alignof__(struct printk_ringbuffer)); + OFFSET(PORT_PRINTK_RINGBUFFER_DESC_RING, printk_ringbuffer, desc_ring); + OFFSET(PORT_PRINTK_RINGBUFFER_TEXT_DATA_RING, printk_ringbuffer, text_data_ring); + OFFSET(PORT_PRINTK_RINGBUFFER_FAIL, printk_ringbuffer, fail); + DEFINE(PORT_PRB_RESERVED_ENTRY_SIZE, sizeof(struct prb_reserved_entry)); + DEFINE(PORT_PRB_RESERVED_ENTRY_ALIGN, __alignof__(struct prb_reserved_entry)); + OFFSET(PORT_PRB_RESERVED_ENTRY_RB, prb_reserved_entry, rb); + OFFSET(PORT_PRB_RESERVED_ENTRY_IRQFLAGS, prb_reserved_entry, irqflags); + OFFSET(PORT_PRB_RESERVED_ENTRY_ID, prb_reserved_entry, id); + OFFSET(PORT_PRB_RESERVED_ENTRY_TEXT_SPACE, prb_reserved_entry, text_space); + DEFINE(PORT_PRINTK_RECORD_SIZE, sizeof(struct printk_record)); + DEFINE(PORT_PRINTK_RECORD_ALIGN, __alignof__(struct printk_record)); + OFFSET(PORT_PRINTK_RECORD_INFO, printk_record, info); + OFFSET(PORT_PRINTK_RECORD_TEXT_BUF, printk_record, text_buf); + OFFSET(PORT_PRINTK_RECORD_TEXT_BUF_SIZE, printk_record, text_buf_size); + DEFINE(PORT_PRINTK_INFO_SIZE, sizeof(struct printk_info)); + DEFINE(PORT_PRINTK_INFO_ALIGN, __alignof__(struct printk_info)); + OFFSET(PORT_PRINTK_INFO_SEQ, printk_info, seq); + OFFSET(PORT_PRINTK_INFO_TS_NSEC, printk_info, ts_nsec); + OFFSET(PORT_PRINTK_INFO_TEXT_LEN, printk_info, text_len); + OFFSET(PORT_PRINTK_INFO_FACILITY, printk_info, facility); + OFFSET(PORT_PRINTK_INFO_CALLER_ID, printk_info, caller_id); + OFFSET(PORT_PRINTK_INFO_DEV_INFO, printk_info, dev_info); + DEFINE(PORT_PRINTK_INFO_DEV_INFO_SIZE, sizeof(struct dev_printk_info)); #endif return 0; } diff --git a/kernel/printk/Makefile b/kernel/printk/Makefile index f8004ac3983da2..79d183ebe2d1a1 100644 --- a/kernel/printk/Makefile +++ b/kernel/printk/Makefile @@ -5,7 +5,7 @@ obj-$(CONFIG_A11Y_BRAILLE_CONSOLE) += braille.o obj-$(CONFIG_PRINTK_INDEX) += index.o obj-$(CONFIG_PRINTK) += printk_support.o -printk_support-y := printk_ringbuffer.o +printk_support-y := $(if $(CONFIG_RUST_KERNEL),printk_ringbuffer_ffi.o,printk_ringbuffer.o) printk_support-$(CONFIG_SYSCTL) += sysctl.o obj-$(CONFIG_PRINTK_RINGBUFFER_KUNIT_TEST) += printk_ringbuffer_kunit_test.o diff --git a/kernel/printk/printk_ringbuffer.rs b/kernel/printk/printk_ringbuffer.rs new file mode 100644 index 00000000000000..457c8c2c7dbf35 --- /dev/null +++ b/kernel/printk/printk_ringbuffer.rs @@ -0,0 +1,1508 @@ +// SPDX-License-Identifier: GPL-2.0-only +//! In-place printk ring buffer replacement. +//! +//! Shared C storage is accessed through the LKMM forwarding boundary. +//! Descriptor validation governs whether speculative snapshots may be used. + +// Descriptor IDs are unsigned long; sequence numbers are always u64. +// Keep the widening conversions explicit for 32-bit kernels. +#![allow(clippy::unnecessary_cast)] + +use core::ffi::{c_char, c_long, c_uint, c_ulong}; +use core::ffi::{c_int, c_void}; +use core::mem::{size_of, MaybeUninit}; +use core::ptr::{addr_of, addr_of_mut, null_mut}; + +unsafe extern "C" { + fn c_prb_load(p: *const c_long) -> c_long; + fn c_prb_load_acquire(p: *const c_long) -> c_long; + fn c_prb_store(p: *mut c_long, value: c_long); + fn c_prb_cas(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; + fn c_prb_cas_relaxed(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; + fn c_prb_cas_release(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; + fn c_prb_inc(p: *mut c_long); + fn c_prb_rmb(); + fn c_prb_irq_save(flags: *mut c_ulong); + fn c_prb_irq_restore(flags: c_ulong); + fn c_prb_copy(dst: *mut c_void, src: *const c_void, n: usize) -> *mut c_void; + fn c_prb_clear(dst: *mut c_void, value: c_int, n: usize) -> *mut c_void; + fn c_prb_find(src: *const c_void, value: c_int, n: usize) -> *mut c_void; + fn c_prb_panic_cpu() -> bool; + fn c_prb_warn_0(condition: bool) -> bool; + fn c_prb_warn_1(condition: bool) -> bool; + fn c_prb_warn_2(condition: bool) -> bool; + fn c_prb_warn_3(condition: bool) -> bool; + fn c_prb_warn_4(condition: bool) -> bool; + fn c_prb_warn_5(condition: bool) -> bool; + fn c_prb_warn_6(condition: bool) -> bool; + fn c_prb_warn_7(condition: bool) -> bool; + fn c_prb_warn_8(condition: bool) -> bool; + fn c_prb_warn_9(condition: bool) -> bool; + fn c_prb_warn_10(condition: bool) -> bool; + fn c_prb_warn_11(condition: bool) -> bool; + fn c_prb_warn_len_zero(length: u16); + fn c_prb_warn_len_max(length: u16, max: c_uint); + static debug_non_panic_cpus: bool; + static legacy_allow_panic_sync: bool; +} + +const RESERVED: i32 = 0; +const EINVAL: c_int = -22; +const ENOENT: c_int = -2; + +/// # Safety +/// ring's arrays are valid; desc is private writable storage. +unsafe fn finalized_seq(ring: *mut DescRing, id: c_ulong, seq: u64, desc: &mut Desc) -> c_int { + let mut actual = 0; + // SAFETY: (U1) snapshot access and state validation use the C boundary. + let s = unsafe { desc_read(ring, id, desc, &mut actual, null_mut()) }; + if s == MISS || s == RESERVED || s == COMMITTED || actual != seq { + return EINVAL; + } + if s == REUSABLE + || (desc.text_blk_lpos.begin == FAILED_LPOS && desc.text_blk_lpos.next == FAILED_LPOS) + { + return ENOENT; + } + 0 +} + +/// # Safety +/// text points to a C ring data range; its contents are speculative. +unsafe fn count_lines(text: *const c_char, size: c_uint) -> c_uint { + // SAFETY: (U1) memchr performs the original C speculative text access. + unsafe { + let mut remaining = size as usize; + let mut next = text; + let mut count = 1; + while remaining != 0 { + let found = c_prb_find(next.cast(), 10, remaining).cast::(); + if found.is_null() { + break; + } + count += 1; + let consumed = found.offset_from(next) as usize + 1; + remaining -= consumed; + next = found.add(1); + } + count + } +} + +/// # Safety +/// ring is valid, lpos is a descriptor snapshot, output buffers are caller-owned. +unsafe fn copy_data( + ring: *const DataRing, + lpos: &DataBlkLpos, + len: u16, + buf: *mut c_char, + size: c_uint, + lines: *mut c_uint, +) -> bool { + if (buf.is_null() || size == 0) && lines.is_null() { + return true; + } + // SAFETY: (U1) the C access layer copies only the range validated by get_data. + unsafe { + let mut available = 0; + let data = get_data(ring, lpos, &mut available); + if data.is_null() || available < c_uint::from(len) { + return false; + } + if !lines.is_null() { + *lines = count_lines(data, c_uint::from(len)); + } + if !buf.is_null() && size != 0 { + c_prb_copy( + buf.cast(), + data.cast(), + size.min(c_uint::from(len)) as usize, + ); + } // LMM(copy_data:A) + true + } +} + +/// # Safety +/// rb is initialized; optional outputs are writable and do not alias shared data. +unsafe fn read_record(rb: *mut Ringbuffer, seq: u64, r: *mut Record, lines: *mut c_uint) -> c_int { + // SAFETY: (U1) speculative C copies are bracketed by descriptor validation. + unsafe { + let ring = addr_of_mut!((*rb).desc_ring); + let info = to_info(ring, seq); + let d = to_desc(ring, seq); + let id = load(addr_of!((*d).state_var)) & ID_MASK; + let mut desc = empty_desc(); + let err = finalized_seq(ring, id, seq, &mut desc); + if err != 0 || r.is_null() { + return err; + } + if !(*r).info.is_null() { + c_prb_copy((*r).info.cast(), info.cast(), size_of::()); + } + if !copy_data( + addr_of!((*rb).text_data_ring), + &desc.text_blk_lpos, + read(addr_of!((*info).text_len)), + (*r).text_buf, + (*r).text_buf_size, + lines, + ) { + return ENOENT; + } + finalized_seq(ring, id, seq, &mut desc) + } +} + +/// Find the sequence number at the current tail, including a data-less record. +/// # Safety +/// rb and all its arrays must remain alive and initialized. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_first_seq(rb: *mut Ringbuffer) -> u64 { + // SAFETY: (U1) the original C tail/descriptor barriers are preserved. + unsafe { + let ring = addr_of_mut!((*rb).desc_ring); + loop { + let id = load(addr_of!((*ring).tail_id)); + let mut seq = 0; + let mut desc = empty_desc(); + let s = desc_read(ring, id, &mut desc, &mut seq, null_mut()); + if s == FINALIZED || s == REUSABLE { + return seq; + } + // Guarantee the last state load from desc_read() is before + // reloading @tail_id in order to see a new tail in the case + // that the descriptor has been recycled. This pairs with + // desc_reserve:D. + // + // Memory barrier involvement: + // + // If prb_first_seq:B reads from desc_reserve:F, then + // prb_first_seq:A reads from desc_push_tail:B. + // + // Relies on: + // + // MB from desc_push_tail:B to desc_reserve:F + // matching + // RMB from prb_first_seq:B to prb_first_seq:A + c_prb_rmb(); // LMM(prb_first_seq:C) + } + } +} + +/// Get the sequence number that the next reservation will receive. +/// # Safety +/// rb and all its arrays must remain alive and initialized. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_next_reserve_seq(rb: *mut Ringbuffer) -> u64 { + // SAFETY: (U1) sequence snapshots use the C atomic publication protocol. + unsafe { + let ring = addr_of_mut!((*rb).desc_ring); + loop { + let seq = last_finalized(rb); + let head = load(addr_of!((*ring).head_id)); + let d = to_desc(ring, seq); + let mut id = load(addr_of!((*d).state_var)) & ID_MASK; + let mut desc = empty_desc(); + if finalized_seq(ring, id, seq, &mut desc) == EINVAL { + if seq != 0 { + continue; + } + let initial = desc_count(ring).wrapping_add(1).wrapping_neg() & ID_MASK; + if head == initial { + return 0; + } + id = initial.wrapping_add(1); + } + return seq + .wrapping_add(head.wrapping_sub(id) as u64) + .wrapping_add(1); + } + } +} + +/// # Safety +/// rb is valid and seq and optional outputs belong to the reader. +unsafe fn read_valid( + rb: *mut Ringbuffer, + seq: &mut u64, + r: *mut Record, + lines: *mut c_uint, +) -> bool { + // SAFETY: (U1) record snapshots and panic-state accesses use C operations. + unsafe { + loop { + let err = read_record(rb, *seq, r, lines); + if err == 0 { + return true; + } + let tail = prb_first_seq(rb); + if *seq < tail { + *seq = tail; + } else if err == ENOENT + || (c_prb_panic_cpu() + && (read(addr_of!(debug_non_panic_cpus).cast::()) == 0 + || read(addr_of!(legacy_allow_panic_sync).cast::()) != 0) + && seq.wrapping_add(1) < prb_next_reserve_seq(rb)) + { + *seq = seq.wrapping_add(1); + } else { + return false; + } + } + } +} + +/// Read a record or the next available record. +/// # Safety +/// rb is initialized; r and its optional buffers are writable reader-owned storage. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_read_valid(rb: *mut Ringbuffer, mut seq: u64, r: *mut Record) -> bool { + // SAFETY: (U1) the caller satisfies the C reader output-buffer contract. + unsafe { read_valid(rb, &mut seq, r, null_mut()) } +} + +/// Read record metadata and optionally count lines. +/// # Safety +/// rb is initialized; info and lines, when non-null, are writable private storage. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_read_valid_info( + rb: *mut Ringbuffer, + mut seq: u64, + info: *mut PrintkInfo, + lines: *mut c_uint, +) -> bool { + let mut r = Record { + info, + text_buf: null_mut(), + text_buf_size: 0, + }; + // SAFETY: (U1) the local record forwards the caller's valid optional outputs. + unsafe { read_valid(rb, &mut seq, &mut r, lines) } +} + +/// Get the oldest available record's sequence number. +/// # Safety +/// rb and its arrays must remain alive and initialized. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_first_valid_seq(rb: *mut Ringbuffer) -> u64 { + let mut seq = 0; + // SAFETY: (U1) no output buffers are requested. + if unsafe { read_valid(rb, &mut seq, null_mut(), null_mut()) } { + seq + } else { + 0 + } +} + +/// Get the sequence number immediately after the last available record. +/// # Safety +/// rb and its arrays must remain alive and initialized. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_next_seq(rb: *mut Ringbuffer) -> u64 { + // SAFETY: (U1) validated sequence reads follow C publication ordering. + unsafe { + let mut seq = last_finalized(rb); + if seq != 0 { + seq = seq.wrapping_add(1); + } + while read_valid(rb, &mut seq, null_mut(), null_mut()) { + seq = seq.wrapping_add(1); + } + seq + } +} + +/// Reserve space for a new record. +/// # Safety +/// rb is initialized. e and r are private writer storage; the caller must commit +/// every successful reservation in the same context before re-enabling IRQs. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_reserve( + e: *mut ReservedEntry, + rb: *mut Ringbuffer, + r: *mut Record, +) -> bool { + // SAFETY: (U1) C operations establish reservation ownership and save IRQ state. + unsafe { + let ring = addr_of_mut!((*rb).desc_ring); + let data = addr_of_mut!((*rb).text_data_ring); + if data_check_size(data, (*r).text_buf_size) { + c_prb_irq_save(addr_of_mut!((*e).irqflags)); + let mut id = 0; + if desc_reserve(rb, &mut id) { + let d = to_desc(ring, id as u64); + let info = to_info(ring, id as u64); + let old_seq = read(addr_of!((*info).seq)); + c_prb_clear(info.cast(), 0, size_of::()); + (*e).rb = rb; + (*e).id = id; + let index = id & (desc_count(ring) - 1); + let seq = if old_seq == 0 && index != 0 { + index as u64 + } else { + old_seq.wrapping_add(desc_count(ring) as u64) + }; + write(addr_of_mut!((*info).seq), seq); + if seq != 0 { + desc_make_final(rb, id.wrapping_sub(1) & ID_MASK); + } + (*r).text_buf = + data_alloc(rb, (*r).text_buf_size, addr_of_mut!((*d).text_blk_lpos), id); + if (*r).text_buf_size == 0 || !(*r).text_buf.is_null() { + (*r).info = info; + let pos = DataBlkLpos { + begin: read(addr_of!((*d).text_blk_lpos.begin)), + next: read(addr_of!((*d).text_blk_lpos.next)), + }; + (*e).text_space = space_used(data, &pos); + return true; + } + prb_commit(e); + } else { + c_prb_inc(addr_of_mut!((*rb).fail)); + c_prb_irq_restore((*e).irqflags); + } + } + c_prb_clear(r.cast(), 0, size_of::()); + false + } +} + +/// Extend the newest committed record owned by caller_id. +/// # Safety +/// Same reservation and IRQ obligations as prb_reserve(). +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_reserve_in_last( + e: *mut ReservedEntry, + rb: *mut Ringbuffer, + r: *mut Record, + caller_id: u32, + max: c_uint, +) -> bool { + // SAFETY: (U1) the C compare-exchange obtains writer ownership before updates. + unsafe { + c_prb_irq_save(addr_of_mut!((*e).irqflags)); + let ring = addr_of_mut!((*rb).desc_ring); + let data = addr_of_mut!((*rb).text_data_ring); + let mut id = 0; + let d = desc_reopen_last(ring, caller_id, &mut id); + if d.is_null() { + c_prb_irq_restore((*e).irqflags); + } else { + let info = to_info(ring, id as u64); + (*e).rb = rb; + (*e).id = id; + let success = 'reserve: { + if caller_id != read(addr_of!((*info).caller_id)) { + break 'reserve false; + } + let pos = DataBlkLpos { + begin: read(addr_of!((*d).text_blk_lpos.begin)), + next: read(addr_of!((*d).text_blk_lpos.next)), + }; + let mut len = read(addr_of!((*info).text_len)); + if dataless(&pos) { + if c_prb_warn_9(len != 0) { + c_prb_warn_len_zero(len); + write(addr_of_mut!((*info).text_len), 0u16); + } + if !data_check_size(data, (*r).text_buf_size) || (*r).text_buf_size > max { + break 'reserve false; + } + (*r).text_buf = + data_alloc(rb, (*r).text_buf_size, addr_of_mut!((*d).text_blk_lpos), id); + } else { + let mut size = 0; + if get_data(data, &pos, &mut size).is_null() { + break 'reserve false; + } + if c_prb_warn_10(c_uint::from(len) > size) { + c_prb_warn_len_max(len, size); + len = size as u16; + write(addr_of_mut!((*info).text_len), len); + } + (*r).text_buf_size = (*r).text_buf_size.wrapping_add(c_uint::from(len)); + if !data_check_size(data, (*r).text_buf_size) || (*r).text_buf_size > max { + break 'reserve false; + } + (*r).text_buf = + data_realloc(rb, (*r).text_buf_size, addr_of_mut!((*d).text_blk_lpos), id); + } + if (*r).text_buf_size != 0 && (*r).text_buf.is_null() { + break 'reserve false; + } + (*r).info = info; + let pos = DataBlkLpos { + begin: read(addr_of!((*d).text_blk_lpos.begin)), + next: read(addr_of!((*d).text_blk_lpos.next)), + }; + (*e).text_space = space_used(data, &pos); + true + }; + if success { + return true; + } + prb_commit(e); + } + c_prb_clear(r.cast(), 0, size_of::()); + false + } +} + +/// Initialize caller-provided descriptor, metadata and text arrays. +/// # Safety +/// All arrays have the power-of-two capacities supplied and valid C alignment. +/// rb and its arrays are exclusively owned until initialization returns. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_init( + rb: *mut Ringbuffer, + text: *mut c_char, + textbits: c_uint, + descs: *mut Desc, + descbits: c_uint, + infos: *mut PrintkInfo, +) { + // SAFETY: (U7) the caller supplies exclusively owned C storage for initialization. + unsafe { + let count = 1usize << descbits; + let id = (count as c_ulong).wrapping_add(1).wrapping_neg() & ID_MASK; + c_prb_clear(descs.cast(), 0, count * size_of::()); + c_prb_clear(infos.cast(), 0, count * size_of::()); + (*rb).desc_ring.count_bits = descbits; + (*rb).desc_ring.descs = descs; + (*rb).desc_ring.infos = infos; + c_prb_store(addr_of_mut!((*rb).desc_ring.head_id), id as c_long); + c_prb_store(addr_of_mut!((*rb).desc_ring.tail_id), id as c_long); + c_prb_store(addr_of_mut!((*rb).desc_ring.last_finalized_seq), 0); + (*rb).text_data_ring.size_bits = textbits; + (*rb).text_data_ring.data = text; + let lpos = (1 as c_ulong).wrapping_shl(textbits).wrapping_neg(); + c_prb_store(addr_of_mut!((*rb).text_data_ring.head_lpos), lpos as c_long); + c_prb_store(addr_of_mut!((*rb).text_data_ring.tail_lpos), lpos as c_long); + c_prb_store(addr_of_mut!((*rb).fail), 0); + let last = descs.add(count - 1); + c_prb_store(addr_of_mut!((*last).state_var), sv(id, REUSABLE) as c_long); + (*last).text_blk_lpos.begin = FAILED_LPOS; + (*last).text_blk_lpos.next = FAILED_LPOS; + (*infos).seq = (count as u64).wrapping_neg(); + (*infos.add(count - 1)).seq = 0; + } +} + +/// Query total text-ring space consumed by a successful reservation. +/// # Safety +/// e belongs to the caller and contains a successful reservation. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_record_text_space(e: *mut ReservedEntry) -> c_uint { + // SAFETY: (U3) the writer owns this initialized reservation handle. + unsafe { (*e).text_space } +} +const COMMITTED: i32 = 1; +const FINALIZED: i32 = 2; +const REUSABLE: i32 = 3; +const MISS: i32 = -1; +const FLAGS_SHIFT: u32 = c_ulong::BITS - 2; +const ID_MASK: c_ulong = c_ulong::MAX >> 2; +const FAILED_LPOS: c_ulong = 1; +fn dataless(lpos: &DataBlkLpos) -> bool { + lpos.begin & lpos.next & 1 != 0 +} + +/// # Safety +/// ring is an initialized data ring. +unsafe fn wrapped(ring: *const DataRing, begin: c_ulong, next: c_ulong) -> bool { + // SAFETY: (U3) size_bits is immutable for the lifetime of the ring. + unsafe { begin >> (*ring).size_bits != next.wrapping_sub(1) >> (*ring).size_bits } +} + +/// # Safety +/// ring is an initialized data ring. +unsafe fn next_lpos(ring: *const DataRing, begin: c_ulong, size: c_uint) -> c_ulong { + let next = begin.wrapping_add(c_ulong::from(size)); + // SAFETY: (U3) only the ring's immutable geometry is inspected. + unsafe { + if !wrapped(ring, begin, next) { + next + } else { + (next & !(data_size(ring) - 1)).wrapping_add(c_ulong::from(size)) + } + } +} + +/// # Safety +/// rb is alive and id is held reserved; lpos is its writable block location. +unsafe fn data_alloc( + rb: *mut Ringbuffer, + size: c_uint, + lpos: *mut DataBlkLpos, + id: c_ulong, +) -> *mut c_char { + // SAFETY: (U1) payload writes/copies and state changes use the C access layer. + unsafe { + if size == 0 { + write(addr_of_mut!((*lpos).begin), EMPTY_LINE_LPOS); + write(addr_of_mut!((*lpos).next), EMPTY_LINE_LPOS); + return null_mut(); + } + let ring = addr_of_mut!((*rb).text_data_ring); + let size = to_blk_size(size); + let mut begin = load(addr_of!((*ring).head_lpos)); + let next = loop { + let next = next_lpos(ring, begin, size); + if c_prb_warn_2(next.wrapping_sub(begin) > data_size(ring)) + || !data_push_tail(rb, next.wrapping_sub(data_size(ring))) + { + write(addr_of_mut!((*lpos).begin), FAILED_LPOS); + write(addr_of_mut!((*lpos).next), FAILED_LPOS); + return null_mut(); + } + if cas(addr_of_mut!((*ring).head_lpos), &mut begin, next) { + break next; + } + // 1. Guarantee any descriptor states that have transitioned + // to reusable are stored before modifying the newly + // allocated data area. A full memory barrier is needed + // since other CPUs may have made the descriptor states + // reusable. See data_push_tail:A about why the reusable + // states are visible. This pairs with desc_read:D. + // + // 2. Guarantee any updated tail lpos is stored before + // modifying the newly allocated data area. Another CPU may + // be in data_make_reusable() and is reading a block ID + // from this area. data_make_reusable() can handle reading + // a garbage block ID value, but then it must be able to + // load a new tail lpos. A full memory barrier is needed + // since other CPUs may have updated the tail lpos. This + // pairs with data_push_tail:B. + }; // LMM(data_alloc:A) + let mut block = to_block(ring, begin); + write(block, id); // LMM(data_alloc:B) + if wrapped(ring, begin, next) { + block = to_block(ring, 0); + write(block, id); + } + write(addr_of_mut!((*lpos).begin), begin); + write(addr_of_mut!((*lpos).next), next); + block.add(1).cast() + } +} + +/// # Safety +/// rb is alive; the caller has reopened id and owns the block. +unsafe fn data_realloc( + rb: *mut Ringbuffer, + size: c_uint, + lpos: *mut DataBlkLpos, + id: c_ulong, +) -> *mut c_char { + // SAFETY: (U1) C accessors preserve the ring protocol for recycled memory. + unsafe { + let ring = addr_of_mut!((*rb).text_data_ring); + let begin = read(addr_of!((*lpos).begin)); + let end = read(addr_of!((*lpos).next)); + let mut head = load(addr_of!((*ring).head_lpos)); + if head != end { + return null_mut(); + } + let was_wrapped = wrapped(ring, begin, end); + let next = next_lpos(ring, begin, to_blk_size(size)); + if !need_more_space(ring, head, next) { + return to_block(ring, if was_wrapped { 0 } else { begin }) + .add(1) + .cast(); + } + if c_prb_warn_3(next.wrapping_sub(begin) > data_size(ring)) + || !data_push_tail(rb, next.wrapping_sub(data_size(ring))) + { + return null_mut(); + } + if !cas(addr_of_mut!((*ring).head_lpos), &mut head, next) { + return null_mut(); + } + let mut block = to_block(ring, begin); + if wrapped(ring, begin, next) { + let old = block; + block = to_block(ring, 0); + write(block, id); + if !was_wrapped { + c_prb_copy( + block.add(1).cast(), + old.add(1).cast(), + end.wrapping_sub(begin) as usize - size_of::(), + ); + } + } + write(addr_of_mut!((*lpos).next), next); + block.add(1).cast() + } +} + +/// # Safety +/// ring is alive, lpos is a validated local snapshot or reserved block. +unsafe fn space_used(ring: *const DataRing, lpos: &DataBlkLpos) -> c_uint { + if dataless(lpos) { + return 0; + } + // SAFETY: (U3) the geometry is immutable and the snapshot is validated. + unsafe { + let size = data_size(ring); + let begin = lpos.begin & (size - 1); + let end = lpos.next & (size - 1); + if !wrapped(ring, lpos.begin, lpos.next) { + end.wrapping_sub(begin) as c_uint + } else { + end.wrapping_add(size).wrapping_sub(begin) as c_uint + } + } +} + +/// # Safety +/// ring is initialized. lpos is a snapshot; returned memory remains speculative. +unsafe fn get_data(ring: *const DataRing, lpos: &DataBlkLpos, size: &mut c_uint) -> *const c_char { + // SAFETY: (U1) warnings use the C boundary; block geometry is checked before use. + unsafe { + if dataless(lpos) { + if lpos.begin == EMPTY_LINE_LPOS && lpos.next == EMPTY_LINE_LPOS { + *size = 0; + return c"".to_bytes_with_nul().as_ptr().cast(); + } + return null_mut(); + } + let block; + if !wrapped(ring, lpos.begin, lpos.next) { + block = to_block(ring, lpos.begin); + *size = lpos.next.wrapping_sub(lpos.begin) as c_uint; + } else if !wrapped(ring, lpos.begin.wrapping_add(data_size(ring)), lpos.next) { + block = to_block(ring, 0); + *size = (lpos.next & (data_size(ring) - 1)) as c_uint; + } else { + c_prb_warn_4(true); + return null_mut(); + } + let mask = size_of::() as c_ulong - 1; + if c_prb_warn_5(lpos.begin & mask != 0) || c_prb_warn_6(lpos.next & mask != 0) { + return null_mut(); + } + if c_prb_warn_7(*size as usize <= size_of::()) { + return null_mut(); + } + *size -= size_of::() as c_uint; + if c_prb_warn_8(!data_check_size(ring, *size)) { + return null_mut(); + } + block.add(1).cast() + } +} + +/// # Safety +/// ring is alive; caller requests ownership of its last committed descriptor. +unsafe fn desc_reopen_last(ring: *mut DescRing, caller: u32, id_out: &mut c_ulong) -> *mut Desc { + // SAFETY: (U1) the C compare-exchange acquires the reserved descriptor state. + unsafe { + let id = load(addr_of!((*ring).head_id)); + let mut desc = empty_desc(); + let mut cid = 0; + if desc_read(ring, id, &mut desc, null_mut(), &mut cid) != COMMITTED || cid != caller { + return null_mut(); + } + let d = to_desc(ring, id as u64); + let mut old = sv(id, COMMITTED); + if !cas(addr_of_mut!((*d).state_var), &mut old, sv(id, RESERVED)) { + return null_mut(); + } + *id_out = id; + d + } +} + +/// # Safety +/// rb is an initialized ringbuffer that remains alive. +unsafe fn last_finalized(rb: *mut Ringbuffer) -> u64 { + // SAFETY: (U1) acquire load pairs with the C release publication protocol. + unsafe { + let seq = c_prb_load_acquire(addr_of!((*rb).desc_ring.last_finalized_seq)) as c_ulong; + #[cfg(target_pointer_width = "64")] + { + seq as u64 + } + #[cfg(target_pointer_width = "32")] + { + let first = prb_first_seq(rb); + first.wrapping_sub((first as u32).wrapping_sub(seq as u32) as i32 as u64) + } + } +} + +/// # Safety +/// rb remains alive and contains initialized arrays. +unsafe fn update_last_finalized(rb: *mut Ringbuffer) { + // SAFETY: (U1) release compare-exchange publishes only validated sequence numbers. + unsafe { + let mut old_seq = last_finalized(rb); + loop { + let mut finalized = old_seq; + let mut next = finalized.wrapping_add(1); + while read_valid(rb, &mut next, null_mut(), null_mut()) { + finalized = next; + next = next.wrapping_add(1); + } + if finalized == old_seq { + return; + } + let mut old = old_seq as c_long; + if c_prb_cas_release( + addr_of_mut!((*rb).desc_ring.last_finalized_seq), + &mut old, + finalized as c_long, + ) { + return; + } + #[cfg(target_pointer_width = "64")] + { + old_seq = old as u64; + } + #[cfg(target_pointer_width = "32")] + { + let first = prb_first_seq(rb); + old_seq = first.wrapping_sub((first as u32).wrapping_sub(old as u32) as i32 as u64); + } + } + } +} + +/// # Safety +/// rb is alive; id refers to a descriptor that may be finalized. +unsafe fn desc_make_final(rb: *mut Ringbuffer, id: c_ulong) { + // SAFETY: (U1) the C compare-exchange respects descriptor ID and state. + unsafe { + let d = to_desc(addr_of_mut!((*rb).desc_ring), id as u64); + let mut old = sv(id, COMMITTED) as c_long; + if c_prb_cas_relaxed( + addr_of_mut!((*d).state_var), + &mut old, + sv(id, FINALIZED) as c_long, + ) { + update_last_finalized(rb); + } + } +} + +/// # Safety +/// e owns a successfully reserved entry with saved IRQ flags. +unsafe fn commit(e: *mut ReservedEntry, s: i32) { + // SAFETY: (U1) C atomic commit publishes the writer's data before restoring IRQs. + unsafe { + let d = to_desc(addr_of_mut!((*(*e).rb).desc_ring), (*e).id as u64); + let mut old = sv((*e).id, RESERVED); + if !cas(addr_of_mut!((*d).state_var), &mut old, sv((*e).id, s)) { + c_prb_warn_11(true); + } + c_prb_irq_restore((*e).irqflags); + } +} + +/// Commit a reserved record, leaving it extendable while it remains newest. +/// # Safety +/// e must be a successfully reserved entry belonging to this execution context. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_commit(e: *mut ReservedEntry) { + // SAFETY: (U1) the caller holds the reservation through the C commit boundary. + unsafe { + let rb = (*e).rb; + commit(e, COMMITTED); + if load(addr_of!((*rb).desc_ring.head_id)) != (*e).id { + desc_make_final(rb, (*e).id); + } + } +} + +/// Commit and finalize a reserved record. +/// # Safety +/// e must be a successfully reserved entry belonging to this execution context. +#[unsafe(no_mangle)] +pub unsafe extern "C" fn prb_final_commit(e: *mut ReservedEntry) { + // SAFETY: (U1) the caller owns the reservation being finalized. + unsafe { + commit(e, FINALIZED); + update_last_finalized((*e).rb); + } +} +const EMPTY_LINE_LPOS: c_ulong = 3; + +// Snapshots of shared C data must consist only of initialized integer bytes. +// This helper never creates a reference to the C allocation. +/// # Safety +/// Both ranges must be valid for T and T must admit all source bit patterns. +/// The C memcpy access must satisfy the documented LKMM snapshot contract. +unsafe fn read(src: *const T) -> T { + let mut out = MaybeUninit::::uninit(); + // SAFETY: (U1) the caller provides a valid C snapshot source and local output. + unsafe { c_prb_copy(out.as_mut_ptr().cast(), src.cast(), size_of::()) }; + // SAFETY: (U3) the private snapshot is initialized by memcpy above; the + // caller guarantees that every source bit pattern is a valid T. + unsafe { out.as_ptr().read() } +} + +/// # Safety +/// dst is writable C storage for T; the caller owns the reservation or init. +unsafe fn write(dst: *mut T, value: T) { + // SAFETY: (U1) memcpy forwards the reservation owner's scalar write. + unsafe { c_prb_copy(dst.cast(), addr_of!(value).cast(), size_of::()) }; +} + +fn sv(id: c_ulong, state: i32) -> c_ulong { + id | ((state as c_ulong) << FLAGS_SHIFT) +} +fn state(id: c_ulong, value: c_ulong) -> i32 { + if value & ID_MASK != id { + MISS + } else { + (value >> FLAGS_SHIFT) as i32 + } +} + +/// # Safety +/// p points to an initialized atomic_long_t governed by the C ring protocol. +unsafe fn load(p: *const c_long) -> c_ulong { + // SAFETY: (U1) atomic_long_read preserves the kernel's LKMM access. + unsafe { c_prb_load(p) as c_ulong } +} + +/// # Safety +/// p is a C atomic_long_t; old is a private snapshot for the compare-exchange. +unsafe fn cas(p: *mut c_long, old: &mut c_ulong, new: c_ulong) -> bool { + let mut expected = *old as c_long; + // SAFETY: (U1) C updates only the supplied atomic and local expected value. + let ok = unsafe { c_prb_cas(p, &mut expected, new as c_long) }; + *old = expected as c_ulong; + ok +} + +/// # Safety +/// ring is initialized and its immutable configuration remains alive. +unsafe fn data_size(ring: *const DataRing) -> c_ulong { + // SAFETY: (U3) size_bits is fixed throughout the ring's active lifetime. + 1 << unsafe { (*ring).size_bits } +} + +/// # Safety +/// ring is initialized and its immutable configuration remains alive. +unsafe fn desc_count(ring: *const DescRing) -> c_ulong { + // SAFETY: (U3) count_bits is fixed throughout the ring's active lifetime. + 1 << unsafe { (*ring).count_bits } +} + +/// # Safety +/// ring's descriptor array is alive and contains 2^count_bits entries. +unsafe fn to_desc(ring: *const DescRing, n: u64) -> *mut Desc { + // SAFETY: (U3) masking confines the index to the initialized C array. + unsafe { + (*ring) + .descs + .add((n as c_ulong & (desc_count(ring) - 1)) as usize) + } +} + +/// # Safety +/// ring's metadata array is alive and contains 2^count_bits entries. +unsafe fn to_info(ring: *const DescRing, n: u64) -> *mut PrintkInfo { + // SAFETY: (U3) masking confines the index to the initialized C array. + unsafe { + (*ring) + .infos + .add((n as c_ulong & (desc_count(ring) - 1)) as usize) + } +} + +/// # Safety +/// ring's text array is alive and contains 2^size_bits bytes. +unsafe fn to_block(ring: *const DataRing, lpos: c_ulong) -> *mut c_ulong { + // SAFETY: (U3) the caller's valid lpos and the mask select the C data block. + unsafe { + (*ring) + .data + .add((lpos & (data_size(ring) - 1)) as usize) + .cast() + } +} + +fn to_blk_size(size: c_uint) -> c_uint { + let word = size_of::() as c_uint; + size.wrapping_add(word).wrapping_add(word - 1) & !(word - 1) +} + +/// # Safety +/// ring is a valid initialized data ring. +unsafe fn data_check_size(ring: *const DataRing, size: c_uint) -> bool { + // SAFETY: (U3) size_bits is immutable. + size == 0 || c_ulong::from(to_blk_size(size)) <= unsafe { data_size(ring) } / 2 +} + +/// # Safety +/// ring is a valid initialized data ring. +unsafe fn need_more_space(ring: *const DataRing, current: c_ulong, target: c_ulong) -> bool { + // SAFETY: (U3) size_bits is immutable. + target.wrapping_sub(current).wrapping_sub(1) < unsafe { data_size(ring) } +} + +/// # Safety +/// ring and its arrays remain alive. Output pointers, when non-null, are private. +unsafe fn desc_read( + ring: *mut DescRing, + id: c_ulong, + out: *mut Desc, + seq: *mut u64, + caller: *mut u32, +) -> i32 { + // SAFETY: (U1) all shared snapshots and barriers use the C LKMM boundary; + // outputs belong to this reader. No shared payload reference is formed. + unsafe { + let desc = to_desc(ring, id as u64); + let info = to_info(ring, id as u64); + let mut value = load(addr_of!((*desc).state_var)); // LMM(desc_read:A) + let mut s = state(id, value); + if s != MISS && s != RESERVED { + // Guarantee the state is loaded before copying the descriptor + // content. This avoids copying obsolete descriptor content that might + // not apply to the descriptor state. This pairs with _prb_commit:B. + // + // Memory barrier involvement: + // + // If desc_read:A reads from _prb_commit:B, then desc_read:C reads + // from _prb_commit:A. + // + // Relies on: + // + // WMB from _prb_commit:A to _prb_commit:B + // matching + // RMB from desc_read:A to desc_read:C + c_prb_rmb(); // LMM(desc_read:B) + if !out.is_null() { + c_prb_copy( + addr_of_mut!((*out).text_blk_lpos).cast(), + addr_of!((*desc).text_blk_lpos).cast(), + size_of::(), + ); + // Copy the descriptor data. The data is not valid until the + // state has been re-checked. A memcpy() for all of @desc + // cannot be used because of the atomic_t @state_var field. + } // LMM(desc_read:C) + if !seq.is_null() { + *seq = read(addr_of!((*info).seq)); + } + if !caller.is_null() { + *caller = read(addr_of!((*info).caller_id)); + } + // 1. Guarantee the descriptor content is loaded before re-checking + // the state. This avoids reading an obsolete descriptor state + // that may not apply to the copied content. This pairs with + // desc_reserve:F. + // + // Memory barrier involvement: + // + // If desc_read:C reads from desc_reserve:G, then desc_read:E + // reads from desc_reserve:F. + // + // Relies on: + // + // WMB from desc_reserve:F to desc_reserve:G + // matching + // RMB from desc_read:C to desc_read:E + // + // 2. Guarantee the record data is loaded before re-checking the + // state. This avoids reading an obsolete descriptor state that may + // not apply to the copied data. This pairs with data_alloc:A and + // data_realloc:A. + // + // Memory barrier involvement: + // + // If copy_data:A reads from data_alloc:B, then desc_read:E + // reads from desc_make_reusable:A. + // + // Relies on: + // + // MB from desc_make_reusable:A to data_alloc:B + // matching + // RMB from desc_read:C to desc_read:E + // + // Note: desc_make_reusable:A and data_alloc:B can be different + // CPUs. However, the data_alloc:B CPU (which performs the + // full memory barrier) must have previously seen + // desc_make_reusable:A. + c_prb_rmb(); // LMM(desc_read:D) + // The data has been copied. Return the current descriptor state, + // which may have changed since the load above. + value = load(addr_of!((*desc).state_var)); // LMM(desc_read:E) + s = state(id, value); + } + if !out.is_null() { + (*out).state_var = value as c_long; + } + s + } +} + +/// # Safety +/// ring and its descriptors remain alive throughout the transition. +unsafe fn desc_make_reusable(ring: *mut DescRing, id: c_ulong) { + // SAFETY: (U1) the C cmpxchg conditionally invalidates the requested ID. + unsafe { + let d = to_desc(ring, id as u64); + let mut old = sv(id, FINALIZED) as c_long; + c_prb_cas_relaxed( + addr_of_mut!((*d).state_var), + &mut old, + sv(id, REUSABLE) as c_long, + ); + } +} + +fn empty_desc() -> Desc { + Desc { + state_var: 0, + text_blk_lpos: DataBlkLpos { begin: 0, next: 0 }, + } +} + +/// # Safety +/// rb and arrays are initialized; out is private writable storage. +unsafe fn data_make_reusable( + rb: *mut Ringbuffer, + mut begin: c_ulong, + end: c_ulong, + out: &mut c_ulong, +) -> bool { + // SAFETY: (U1) C accessors read racing payload and preserve descriptor ordering. + unsafe { + let data = addr_of_mut!((*rb).text_data_ring); + let ring = addr_of_mut!((*rb).desc_ring); + while need_more_space(data, begin, end) { + // Load the block ID from the data block. This is a data race + // against a writer that may have newly reserved this data + // area. If the loaded value matches a valid descriptor ID, + // the blk_lpos of that descriptor will be checked to make + // sure it points back to this data block. If the check fails, + // the data area has been recycled by another writer. + let id = read(to_block(data, begin)); // LMM(data_make_reusable:A) + let mut desc = empty_desc(); + match desc_read(ring, id, &mut desc, null_mut(), null_mut()) { + FINALIZED => { + if desc.text_blk_lpos.begin != begin { + return false; + } + desc_make_reusable(ring, id); + } + REUSABLE => { + if desc.text_blk_lpos.begin != begin { + return false; + } + } + _ => return false, + } + begin = desc.text_blk_lpos.next; + } + *out = begin; + true + } +} + +/// # Safety +/// rb and arrays remain valid throughout concurrent recycling. +unsafe fn data_push_tail(rb: *mut Ringbuffer, lpos: c_ulong) -> bool { + if lpos & 1 != 0 { + return true; + } + // SAFETY: (U1) the C atomics and barriers preserve the original tail protocol. + unsafe { + let ring = addr_of_mut!((*rb).text_data_ring); + // Any descriptor states that have transitioned to reusable due to the + // data tail being pushed to this loaded value will be visible to this + // CPU. This pairs with data_push_tail:D. + // + // Memory barrier involvement: + // + // If data_push_tail:A reads from data_push_tail:D, then this CPU can + // see desc_make_reusable:A. + // + // Relies on: + // + // MB from desc_make_reusable:A to data_push_tail:D + // matches + // READFROM from data_push_tail:D to data_push_tail:A + // thus + // READFROM from desc_make_reusable:A to this CPU + let mut tail = load(addr_of!((*ring).tail_lpos)); // LMM(data_push_tail:A) + while need_more_space(ring, tail, lpos) { + let mut next = 0; + if !data_make_reusable(rb, tail, lpos, &mut next) { + // 1. Guarantee the block ID loaded in + // data_make_reusable() is performed before + // reloading the tail lpos. The failed + // data_make_reusable() may be due to a newly + // recycled data area causing the tail lpos to + // have been previously pushed. This pairs with + // data_alloc:A and data_realloc:A. + // + // Memory barrier involvement: + // + // If data_make_reusable:A reads from data_alloc:B, + // then data_push_tail:C reads from + // data_push_tail:D. + // + // Relies on: + // + // MB from data_push_tail:D to data_alloc:B + // matching + // RMB from data_make_reusable:A to + // data_push_tail:C + // + // Note: data_push_tail:D and data_alloc:B can be + // different CPUs. However, the data_alloc:B + // CPU (which performs the full memory + // barrier) must have previously seen + // data_push_tail:D. + // + // 2. Guarantee the descriptor state loaded in + // data_make_reusable() is performed before + // reloading the tail lpos. The failed + // data_make_reusable() may be due to a newly + // recycled descriptor causing the tail lpos to + // have been previously pushed. This pairs with + // desc_reserve:D. + // + // Memory barrier involvement: + // + // If data_make_reusable:B reads from + // desc_reserve:F, then data_push_tail:C reads + // from data_push_tail:D. + // + // Relies on: + // + // MB from data_push_tail:D to desc_reserve:F + // matching + // RMB from data_make_reusable:B to + // data_push_tail:C + // + // Note: data_push_tail:D and desc_reserve:F can + // be different CPUs. However, the + // desc_reserve:F CPU (which performs the + // full memory barrier) must have previously + // seen data_push_tail:D. + c_prb_rmb(); // LMM(data_push_tail:B) + let new_tail = load(addr_of!((*ring).tail_lpos)); // LMM(data_push_tail:C) + if new_tail == tail { + return false; + } + tail = new_tail; + continue; + } + if cas(addr_of_mut!((*ring).tail_lpos), &mut tail, next) { + break; + } + // Guarantee any descriptor states that have transitioned to + // reusable are stored before pushing the tail lpos. A full + // memory barrier is needed since other CPUs may have made + // the descriptor states reusable. This pairs with + // data_push_tail:A. + } // LMM(data_push_tail:D) + true + } +} + +/// # Safety +/// rb and arrays remain valid throughout concurrent recycling. +unsafe fn desc_push_tail(rb: *mut Ringbuffer, tail: c_ulong) -> bool { + // SAFETY: (U1) all shared state transitions use the original C atomics. + unsafe { + let ring = addr_of_mut!((*rb).desc_ring); + let mut desc = empty_desc(); + match desc_read(ring, tail, &mut desc, null_mut(), null_mut()) { + MISS => { + return (desc.state_var as c_ulong & ID_MASK) + != (tail.wrapping_sub(desc_count(ring)) & ID_MASK) + } + RESERVED | COMMITTED => return false, + FINALIZED => desc_make_reusable(ring, tail), + _ => {} + } + if !data_push_tail(rb, desc.text_blk_lpos.next) { + return false; + } + let next = tail.wrapping_add(1) & ID_MASK; + // Check the next descriptor after @tail_id before pushing the tail + // to it because the tail must always be in a finalized or reusable + // state. The implementation of prb_first_seq() relies on this. + // + // A successful read implies that the next descriptor is less than or + // equal to @head_id so there is no risk of pushing the tail past the + // head. + let s = desc_read(ring, next, &mut desc, null_mut(), null_mut()); // LMM(desc_push_tail:A) + if s == FINALIZED || s == REUSABLE { + let mut expected = tail; + // Guarantee any descriptor states that have transitioned to + // reusable are stored before pushing the tail ID. This allows + // verifying the recycled descriptor state. A full memory + // barrier is needed since other CPUs may have made the + // descriptor states reusable. This pairs with desc_reserve:D. + cas(addr_of_mut!((*ring).tail_id), &mut expected, next); // LMM(desc_push_tail:B) + } else { + // Guarantee the last state load from desc_read() is before + // reloading @tail_id in order to see a new tail ID in the + // case that the descriptor has been recycled. This pairs + // with desc_reserve:D. + // + // Memory barrier involvement: + // + // If desc_push_tail:A reads from desc_reserve:F, then + // desc_push_tail:D reads from desc_push_tail:B. + // + // Relies on: + // + // MB from desc_push_tail:B to desc_reserve:F + // matching + // RMB from desc_push_tail:A to desc_push_tail:D + // + // Note: desc_push_tail:B and desc_reserve:F can be different + // CPUs. However, the desc_reserve:F CPU (which performs + // the full memory barrier) must have previously seen + // desc_push_tail:B. + c_prb_rmb(); // LMM(desc_push_tail:C) + if load(addr_of!((*ring).tail_id)) == tail { + return false; + } + } + true + } +} + +/// # Safety +/// rb and arrays remain alive; id_out is private storage for the reservation. +unsafe fn desc_reserve(rb: *mut Ringbuffer, id_out: &mut c_ulong) -> bool { + // SAFETY: (U1) C atomics/barriers implement the reservation ownership protocol. + unsafe { + let ring = addr_of_mut!((*rb).desc_ring); + let mut head = load(addr_of!((*ring).head_id)); + let (id, previous) = loop { + let id = head.wrapping_add(1) & ID_MASK; + let previous = id.wrapping_sub(desc_count(ring)) & ID_MASK; + // Guarantee the head ID is read before reading the tail ID. + // Since the tail ID is updated before the head ID, this + // guarantees that @id_prev_wrap is never ahead of the tail + // ID. This pairs with desc_reserve:D. + // + // Memory barrier involvement: + // + // If desc_reserve:A reads from desc_reserve:D, then + // desc_reserve:C reads from desc_push_tail:B. + // + // Relies on: + // + // MB from desc_push_tail:B to desc_reserve:D + // matching + // RMB from desc_reserve:A to desc_reserve:C + // + // Note: desc_push_tail:B and desc_reserve:D can be different + // CPUs. However, the desc_reserve:D CPU (which performs + // the full memory barrier) must have previously seen + // desc_push_tail:B. + c_prb_rmb(); // LMM(desc_reserve:B) + if previous == load(addr_of!((*ring).tail_id)) && !desc_push_tail(rb, previous) { + return false; + } + if cas(addr_of_mut!((*ring).head_id), &mut head, id) { + break (id, previous); + } + // 1. Guarantee the tail ID is read before validating the + // recycled descriptor state. A read memory barrier is + // sufficient for this. This pairs with desc_push_tail:B. + // + // Memory barrier involvement: + // + // If desc_reserve:C reads from desc_push_tail:B, then + // desc_reserve:E reads from desc_make_reusable:A. + // + // Relies on: + // + // MB from desc_make_reusable:A to desc_push_tail:B + // matching + // RMB from desc_reserve:C to desc_reserve:E + // + // Note: desc_make_reusable:A and desc_push_tail:B can be + // different CPUs. However, the desc_push_tail:B CPU + // (which performs the full memory barrier) must have + // previously seen desc_make_reusable:A. + // + // 2. Guarantee the tail ID is stored before storing the head + // ID. This pairs with desc_reserve:B. + // + // 3. Guarantee any data ring tail changes are stored before + // recycling the descriptor. Data ring tail changes can + // happen via desc_push_tail()->data_push_tail(). A full + // memory barrier is needed since another CPU may have + // pushed the data ring tails. This pairs with + // data_push_tail:B. + // + // 4. Guarantee a new tail ID is stored before recycling the + // descriptor. A full memory barrier is needed since + // another CPU may have pushed the tail ID. This pairs + // with desc_push_tail:C and this also pairs with + // prb_first_seq:C. + // + // 5. Guarantee the head ID is stored before trying to + // finalize the previous descriptor. This pairs with + // _prb_commit:B. + }; // LMM(desc_reserve:D) + let desc = to_desc(ring, id as u64); + // If the descriptor has been recycled, verify the old state val. + // See "ABA Issues" about why this verification is performed. + let mut old = load(addr_of!((*desc).state_var)); // LMM(desc_reserve:E) + if old != 0 && state(previous, old) != REUSABLE { + c_prb_warn_0(true); + return false; + } + if !cas(addr_of_mut!((*desc).state_var), &mut old, sv(id, RESERVED)) { + c_prb_warn_1(true); + return false; + // Assign the descriptor a new ID and set its state to reserved. + // See "ABA Issues" about why cmpxchg() instead of set() is used. + // + // Guarantee the new descriptor ID and state is stored before making + // any other changes. A write memory barrier is sufficient for this. + // This pairs with desc_read:D. + } // LMM(desc_reserve:F) + *id_out = id; + true + } +} +use crate::port_layout::*; + +/// Logical position and extent of a C ring data block. +#[repr(C)] +pub struct DataBlkLpos { + begin: c_ulong, + next: c_ulong, +} + +kr::static_assert_layout!(DataBlkLpos, size = PORT_PRB_DATA_BLK_LPOS_SIZE, align = PORT_PRB_DATA_BLK_LPOS_ALIGN, + begin @ PORT_PRB_DATA_BLK_LPOS_BEGIN, + next @ PORT_PRB_DATA_BLK_LPOS_NEXT, +); + +/// C descriptor containing publication state and text positions. +#[repr(C)] +pub struct Desc { + state_var: c_long, + text_blk_lpos: DataBlkLpos, +} + +kr::static_assert_layout!(Desc, size = PORT_PRB_DESC_SIZE, align = PORT_PRB_DESC_ALIGN, + state_var @ PORT_PRB_DESC_STATE_VAR, + text_blk_lpos @ PORT_PRB_DESC_TEXT_BLK_LPOS, +); + +/// C text ring geometry and atomic positions. +#[repr(C)] +pub struct DataRing { + size_bits: c_uint, + data: *mut c_char, + head_lpos: c_long, + tail_lpos: c_long, +} + +kr::static_assert_layout!(DataRing, size = PORT_PRB_DATA_RING_SIZE, align = PORT_PRB_DATA_RING_ALIGN, + size_bits @ PORT_PRB_DATA_RING_SIZE_BITS, + data @ PORT_PRB_DATA_RING_DATA, + head_lpos @ PORT_PRB_DATA_RING_HEAD_LPOS, + tail_lpos @ PORT_PRB_DATA_RING_TAIL_LPOS, +); + +/// C descriptor ring geometry and atomic sequence state. +#[repr(C)] +pub struct DescRing { + count_bits: c_uint, + descs: *mut Desc, + infos: *mut PrintkInfo, + head_id: c_long, + tail_id: c_long, + last_finalized_seq: c_long, +} + +kr::static_assert_layout!(DescRing, size = PORT_PRB_DESC_RING_SIZE, align = PORT_PRB_DESC_RING_ALIGN, + count_bits @ PORT_PRB_DESC_RING_COUNT_BITS, + descs @ PORT_PRB_DESC_RING_DESCS, + infos @ PORT_PRB_DESC_RING_INFOS, + head_id @ PORT_PRB_DESC_RING_HEAD_ID, + tail_id @ PORT_PRB_DESC_RING_TAIL_ID, + last_finalized_seq @ PORT_PRB_DESC_RING_LAST_FINALIZED_SEQ, +); + +/// C printk ring buffer shared by readers and writers. +#[repr(C)] +pub struct Ringbuffer { + desc_ring: DescRing, + text_data_ring: DataRing, + fail: c_long, +} + +kr::static_assert_layout!(Ringbuffer, size = PORT_PRINTK_RINGBUFFER_SIZE, align = PORT_PRINTK_RINGBUFFER_ALIGN, + desc_ring @ PORT_PRINTK_RINGBUFFER_DESC_RING, + text_data_ring @ PORT_PRINTK_RINGBUFFER_TEXT_DATA_RING, + fail @ PORT_PRINTK_RINGBUFFER_FAIL, +); + +/// Writer-owned reservation returned by the C ABI. +#[repr(C)] +pub struct ReservedEntry { + rb: *mut Ringbuffer, + irqflags: c_ulong, + id: c_ulong, + text_space: c_uint, +} + +kr::static_assert_layout!(ReservedEntry, size = PORT_PRB_RESERVED_ENTRY_SIZE, align = PORT_PRB_RESERVED_ENTRY_ALIGN, + rb @ PORT_PRB_RESERVED_ENTRY_RB, + irqflags @ PORT_PRB_RESERVED_ENTRY_IRQFLAGS, + id @ PORT_PRB_RESERVED_ENTRY_ID, + text_space @ PORT_PRB_RESERVED_ENTRY_TEXT_SPACE, +); + +/// Caller-provided metadata and text buffers. +#[repr(C)] +pub struct Record { + info: *mut PrintkInfo, + text_buf: *mut c_char, + text_buf_size: c_uint, +} + +kr::static_assert_layout!(Record, size = PORT_PRINTK_RECORD_SIZE, align = PORT_PRINTK_RECORD_ALIGN, + info @ PORT_PRINTK_RECORD_INFO, + text_buf @ PORT_PRINTK_RECORD_TEXT_BUF, + text_buf_size @ PORT_PRINTK_RECORD_TEXT_BUF_SIZE, +); + +// The record metadata is owned by printk's C writers. Its bitfields and +// configuration-dependent execution context stay opaque to the ring buffer. +// Scalar accessors below use the C-generated offsets. +/// C record metadata, including opaque execution-context fields. +#[repr(C)] +pub struct PrintkInfo { + seq: u64, + ts_nsec: u64, + text_len: u16, + facility: u8, + flags_level: u8, + caller_id: u32, + execution_context: [u8; PORT_PRINTK_INFO_DEV_INFO - PORT_PRINTK_INFO_CALLER_ID - 4], + dev_info: [u8; PORT_PRINTK_INFO_DEV_INFO_SIZE], +} + +kr::static_assert_layout!(PrintkInfo, + size = PORT_PRINTK_INFO_SIZE, align = PORT_PRINTK_INFO_ALIGN, + seq @ PORT_PRINTK_INFO_SEQ, ts_nsec @ PORT_PRINTK_INFO_TS_NSEC, + text_len @ PORT_PRINTK_INFO_TEXT_LEN, facility @ PORT_PRINTK_INFO_FACILITY, + flags_level @ (PORT_PRINTK_INFO_FACILITY + 1), + caller_id @ PORT_PRINTK_INFO_CALLER_ID, + execution_context @ (PORT_PRINTK_INFO_CALLER_ID + 4), + dev_info @ PORT_PRINTK_INFO_DEV_INFO, +); diff --git a/kernel/printk/printk_ringbuffer_ffi.c b/kernel/printk/printk_ringbuffer_ffi.c new file mode 100644 index 00000000000000..92e5808bea8278 --- /dev/null +++ b/kernel/printk/printk_ringbuffer_ffi.c @@ -0,0 +1,182 @@ +// SPDX-License-Identifier: GPL-2.0-only +/* LKMM primitives used by the Rust printk ring buffer. */ +#include +#include +#include +#include +#include "internal.h" +#include "printk_ringbuffer.h" +#include + +EXPORT_SYMBOL_IF_KUNIT(prb_reserve); +EXPORT_SYMBOL_IF_KUNIT(prb_commit); +EXPORT_SYMBOL_IF_KUNIT(prb_read_valid); +EXPORT_SYMBOL_IF_KUNIT(prb_init); + +long c_prb_load(const atomic_long_t *p); +long c_prb_load(const atomic_long_t *p) +{ + return atomic_long_read(p); +} + +long c_prb_load_acquire(const atomic_long_t *p); +long c_prb_load_acquire(const atomic_long_t *p) +{ + return atomic_long_read_acquire(p); +} + +void c_prb_store(atomic_long_t *p, long value); +void c_prb_store(atomic_long_t *p, long value) +{ + atomic_long_set(p, value); +} + +bool c_prb_cas(atomic_long_t *p, long *old, long value); +bool c_prb_cas(atomic_long_t *p, long *old, long value) +{ + return atomic_long_try_cmpxchg(p, old, value); +} + +bool c_prb_cas_relaxed(atomic_long_t *p, long *old, long value); +bool c_prb_cas_relaxed(atomic_long_t *p, long *old, long value) +{ + return atomic_long_try_cmpxchg_relaxed(p, old, value); +} + +bool c_prb_cas_release(atomic_long_t *p, long *old, long value); +bool c_prb_cas_release(atomic_long_t *p, long *old, long value) +{ + return atomic_long_try_cmpxchg_release(p, old, value); +} + +void c_prb_inc(atomic_long_t *p); +void c_prb_inc(atomic_long_t *p) +{ + atomic_long_inc(p); +} + +void c_prb_rmb(void); +void c_prb_rmb(void) +{ + smp_rmb(); +} + +void c_prb_irq_save(unsigned long *flags); +void c_prb_irq_save(unsigned long *flags) +{ + local_irq_save(*flags); +} + +void c_prb_irq_restore(unsigned long flags); +void c_prb_irq_restore(unsigned long flags) +{ + local_irq_restore(flags); +} + +void * c_prb_copy(void *dst, const void *src, size_t n); +void * c_prb_copy(void *dst, const void *src, size_t n) +{ + return memcpy(dst, src, n); +} + +void * c_prb_clear(void *dst, int value, size_t n); +void * c_prb_clear(void *dst, int value, size_t n) +{ + return memset(dst, value, n); +} + +void * c_prb_find(const void *src, int value, size_t n); +void * c_prb_find(const void *src, int value, size_t n) +{ + return memchr(src, value, n); +} + +bool c_prb_panic_cpu(void); +bool c_prb_panic_cpu(void) +{ + return panic_on_this_cpu(); +} + +bool c_prb_warn_0(bool condition); +bool c_prb_warn_0(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_1(bool condition); +bool c_prb_warn_1(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_2(bool condition); +bool c_prb_warn_2(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_3(bool condition); +bool c_prb_warn_3(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_4(bool condition); +bool c_prb_warn_4(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_5(bool condition); +bool c_prb_warn_5(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_6(bool condition); +bool c_prb_warn_6(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_7(bool condition); +bool c_prb_warn_7(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_8(bool condition); +bool c_prb_warn_8(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_9(bool condition); +bool c_prb_warn_9(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_10(bool condition); +bool c_prb_warn_10(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +bool c_prb_warn_11(bool condition); +bool c_prb_warn_11(bool condition) +{ + return WARN_ON_ONCE(condition); +} + +void c_prb_warn_len_zero(unsigned short length); +void c_prb_warn_len_zero(unsigned short length) +{ + pr_warn_once("wrong text_len value (%hu, expecting 0)\n", length); +} + +void c_prb_warn_len_max(unsigned short length, unsigned int max); +void c_prb_warn_len_max(unsigned short length, unsigned int max) +{ + pr_warn_once("wrong text_len value (%hu, expecting <=%u)\n", length, max); +} diff --git a/main.rs b/main.rs index ccc82f2bc8b3ff..2a42ce73281489 100644 --- a/main.rs +++ b/main.rs @@ -7,6 +7,10 @@ #![no_std] +#[cfg(CONFIG_PRINTK)] +#[path = "kernel/printk/printk_ringbuffer.rs"] +pub mod printk_ringbuffer; + mod port_layout { include!(concat!(env!("OBJTREE"), "/include/generated/port-layout.rs")); } diff --git a/rust/Makefile b/rust/Makefile index 8a32436ce1602a..c7e6776b0c1d1e 100644 --- a/rust/Makefile +++ b/rust/Makefile @@ -788,6 +788,7 @@ $(obj)/main.o: private skip_flags = --edition=2021 $(obj)/main.o: private rustc_target_flags = --edition=2024 --extern kr \ @$(objtree)/include/generated/port-layout.cfg \ '--check-cfg=cfg(PORT_LOCKREF_FAST)' \ + '--check-cfg=cfg(CONFIG_PRINTK)' \ '--check-cfg=cfg(CONFIG_ARCH_DMA_ADDR_T_64BIT)' \ '--check-cfg=cfg(CONFIG_PROVE_RCU)' $(obj)/main.o: $(srctree)/main.rs $(obj)/kr.o \ From 363ee4e39f4fb88b23eb367393b3e2635edfa70b Mon Sep 17 00:00:00 2001 From: Bruno Herrera Date: Mon, 21 Sep 2026 14:15:25 -0300 Subject: [PATCH 5/6] printk: use marked scalar snapshots Keep speculative scalar reads inside C READ_ONCE helpers before descriptor validation. This avoids turning ordinary racing memcpy results into Rust scalar values while preserving the ring algorithm and C caller ABI. Test: printk ringbuffer differential harness, ten seeds and both metadata layouts. Test: SMP Rust kernel build with printk KUnit configuration. Assisted-by: LLM [Codex] --- kernel/printk/printk_ringbuffer.rs | 1008 ++++++++++++++++++------- kernel/printk/printk_ringbuffer_ffi.c | 132 ++-- 2 files changed, 800 insertions(+), 340 deletions(-) diff --git a/kernel/printk/printk_ringbuffer.rs b/kernel/printk/printk_ringbuffer.rs index 457c8c2c7dbf35..ff47c92312c07d 100644 --- a/kernel/printk/printk_ringbuffer.rs +++ b/kernel/printk/printk_ringbuffer.rs @@ -3,45 +3,52 @@ //! //! Shared C storage is accessed through the LKMM forwarding boundary. //! Descriptor validation governs whether speculative snapshots may be used. +//! +//! Algorithm comments retain C identifiers and LMM labels for comparison with +//! the reference implementation beside this file. // Descriptor IDs are unsigned long; sequence numbers are always u64. // Keep the widening conversions explicit for 32-bit kernels. #![allow(clippy::unnecessary_cast)] -use core::ffi::{c_char, c_long, c_uint, c_ulong}; -use core::ffi::{c_int, c_void}; -use core::mem::{size_of, MaybeUninit}; +use crate::port_layout::*; + +use core::ffi::{c_char, c_int, c_long, c_uint, c_ulong, c_void}; +use core::mem::size_of; use core::ptr::{addr_of, addr_of_mut, null_mut}; unsafe extern "C" { - fn c_prb_load(p: *const c_long) -> c_long; - fn c_prb_load_acquire(p: *const c_long) -> c_long; - fn c_prb_store(p: *mut c_long, value: c_long); - fn c_prb_cas(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; - fn c_prb_cas_relaxed(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; - fn c_prb_cas_release(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; - fn c_prb_inc(p: *mut c_long); - fn c_prb_rmb(); - fn c_prb_irq_save(flags: *mut c_ulong); - fn c_prb_irq_restore(flags: c_ulong); - fn c_prb_copy(dst: *mut c_void, src: *const c_void, n: usize) -> *mut c_void; - fn c_prb_clear(dst: *mut c_void, value: c_int, n: usize) -> *mut c_void; - fn c_prb_find(src: *const c_void, value: c_int, n: usize) -> *mut c_void; - fn c_prb_panic_cpu() -> bool; - fn c_prb_warn_0(condition: bool) -> bool; - fn c_prb_warn_1(condition: bool) -> bool; - fn c_prb_warn_2(condition: bool) -> bool; - fn c_prb_warn_3(condition: bool) -> bool; - fn c_prb_warn_4(condition: bool) -> bool; - fn c_prb_warn_5(condition: bool) -> bool; - fn c_prb_warn_6(condition: bool) -> bool; - fn c_prb_warn_7(condition: bool) -> bool; - fn c_prb_warn_8(condition: bool) -> bool; - fn c_prb_warn_9(condition: bool) -> bool; - fn c_prb_warn_10(condition: bool) -> bool; - fn c_prb_warn_11(condition: bool) -> bool; - fn c_prb_warn_len_zero(length: u16); - fn c_prb_warn_len_max(length: u16, max: c_uint); + fn c_printk_ringbuffer_load(p: *const c_long) -> c_long; + fn c_printk_ringbuffer_load_acquire(p: *const c_long) -> c_long; + fn c_printk_ringbuffer_store(p: *mut c_long, value: c_long); + fn c_printk_ringbuffer_cas(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; + fn c_printk_ringbuffer_cas_relaxed(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; + fn c_printk_ringbuffer_cas_release(p: *mut c_long, old: *mut c_long, value: c_long) -> bool; + fn c_printk_ringbuffer_inc(p: *mut c_long); + fn c_printk_ringbuffer_rmb(); + fn c_printk_ringbuffer_irq_save(flags: *mut c_ulong); + fn c_printk_ringbuffer_irq_restore(flags: c_ulong); + fn c_printk_ringbuffer_copy(dst: *mut c_void, src: *const c_void, n: usize) -> *mut c_void; + fn c_printk_ringbuffer_clear(dst: *mut c_void, value: c_int, n: usize) -> *mut c_void; + fn c_printk_ringbuffer_read_u8(src: *const u8) -> u8; + fn c_printk_ringbuffer_read_u16(src: *const u16) -> u16; + fn c_printk_ringbuffer_read_u32(src: *const u32) -> u32; + fn c_printk_ringbuffer_read_u64(src: *const u64) -> u64; + fn c_printk_ringbuffer_panic_cpu() -> bool; + fn c_printk_ringbuffer_warn_descriptor_state(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_descriptor_reserve(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_allocation_size(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_reallocation_size(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_block_wrap(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_block_begin_alignment(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_block_next_alignment(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_block_minimum_size(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_block_maximum_size(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_empty_text_length(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_text_length(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_commit(condition: bool) -> bool; + fn c_printk_ringbuffer_warn_len_zero(length: u16); + fn c_printk_ringbuffer_warn_len_max(length: u16, max: c_uint); static debug_non_panic_cpus: bool; static legacy_allow_panic_sync: bool; } @@ -49,16 +56,163 @@ unsafe extern "C" { const RESERVED: i32 = 0; const EINVAL: c_int = -22; const ENOENT: c_int = -2; +const COMMITTED: i32 = 1; +const FINALIZED: i32 = 2; +const REUSABLE: i32 = 3; +const MISS: i32 = -1; +const FLAGS_SHIFT: u32 = c_ulong::BITS - 2; +const ID_MASK: c_ulong = c_ulong::MAX >> 2; +const FAILED_LPOS: c_ulong = 1; +const EMPTY_LINE_LPOS: c_ulong = 3; + +/// Logical position and extent of a C ring data block. +#[repr(C)] +pub struct DataBlkLpos { + begin: c_ulong, + next: c_ulong, +} + +kr::static_assert_layout!(DataBlkLpos, size = PORT_PRB_DATA_BLK_LPOS_SIZE, align = PORT_PRB_DATA_BLK_LPOS_ALIGN, + begin @ PORT_PRB_DATA_BLK_LPOS_BEGIN, + next @ PORT_PRB_DATA_BLK_LPOS_NEXT, +); + +/// C descriptor containing publication state and text positions. +#[repr(C)] +pub struct Desc { + state_var: c_long, + text_blk_lpos: DataBlkLpos, +} + +kr::static_assert_layout!(Desc, size = PORT_PRB_DESC_SIZE, align = PORT_PRB_DESC_ALIGN, + state_var @ PORT_PRB_DESC_STATE_VAR, + text_blk_lpos @ PORT_PRB_DESC_TEXT_BLK_LPOS, +); + +/// C text ring geometry and atomic positions. +#[repr(C)] +pub struct DataRing { + size_bits: c_uint, + data: *mut c_char, + head_lpos: c_long, + tail_lpos: c_long, +} + +kr::static_assert_layout!(DataRing, size = PORT_PRB_DATA_RING_SIZE, align = PORT_PRB_DATA_RING_ALIGN, + size_bits @ PORT_PRB_DATA_RING_SIZE_BITS, + data @ PORT_PRB_DATA_RING_DATA, + head_lpos @ PORT_PRB_DATA_RING_HEAD_LPOS, + tail_lpos @ PORT_PRB_DATA_RING_TAIL_LPOS, +); + +/// C descriptor ring geometry and atomic sequence state. +#[repr(C)] +pub struct DescRing { + count_bits: c_uint, + descs: *mut Desc, + infos: *mut PrintkInfo, + head_id: c_long, + tail_id: c_long, + last_finalized_seq: c_long, +} + +kr::static_assert_layout!(DescRing, size = PORT_PRB_DESC_RING_SIZE, align = PORT_PRB_DESC_RING_ALIGN, + count_bits @ PORT_PRB_DESC_RING_COUNT_BITS, + descs @ PORT_PRB_DESC_RING_DESCS, + infos @ PORT_PRB_DESC_RING_INFOS, + head_id @ PORT_PRB_DESC_RING_HEAD_ID, + tail_id @ PORT_PRB_DESC_RING_TAIL_ID, + last_finalized_seq @ PORT_PRB_DESC_RING_LAST_FINALIZED_SEQ, +); + +/// C printk ring buffer shared by readers and writers. +#[repr(C)] +pub struct Ringbuffer { + desc_ring: DescRing, + text_data_ring: DataRing, + fail: c_long, +} + +kr::static_assert_layout!(Ringbuffer, size = PORT_PRINTK_RINGBUFFER_SIZE, align = PORT_PRINTK_RINGBUFFER_ALIGN, + desc_ring @ PORT_PRINTK_RINGBUFFER_DESC_RING, + text_data_ring @ PORT_PRINTK_RINGBUFFER_TEXT_DATA_RING, + fail @ PORT_PRINTK_RINGBUFFER_FAIL, +); + +/// Writer-owned reservation returned by the C ABI. +#[repr(C)] +pub struct ReservedEntry { + rb: *mut Ringbuffer, + irqflags: c_ulong, + id: c_ulong, + text_space: c_uint, +} + +kr::static_assert_layout!(ReservedEntry, size = PORT_PRB_RESERVED_ENTRY_SIZE, align = PORT_PRB_RESERVED_ENTRY_ALIGN, + rb @ PORT_PRB_RESERVED_ENTRY_RB, + irqflags @ PORT_PRB_RESERVED_ENTRY_IRQFLAGS, + id @ PORT_PRB_RESERVED_ENTRY_ID, + text_space @ PORT_PRB_RESERVED_ENTRY_TEXT_SPACE, +); + +/// Caller-provided metadata and text buffers. +#[repr(C)] +pub struct Record { + info: *mut PrintkInfo, + text_buf: *mut c_char, + text_buf_size: c_uint, +} + +kr::static_assert_layout!(Record, size = PORT_PRINTK_RECORD_SIZE, align = PORT_PRINTK_RECORD_ALIGN, + info @ PORT_PRINTK_RECORD_INFO, + text_buf @ PORT_PRINTK_RECORD_TEXT_BUF, + text_buf_size @ PORT_PRINTK_RECORD_TEXT_BUF_SIZE, +); + +// The record metadata is owned by printk's C writers. Its bitfields and +// configuration-dependent execution context stay opaque to the ring buffer. +// Scalar field offsets are checked against the C-generated values. +/// C record metadata, including opaque execution-context fields. +#[repr(C)] +pub struct PrintkInfo { + seq: u64, + ts_nsec: u64, + text_len: u16, + facility: u8, + flags_level: u8, + caller_id: u32, + execution_context: [u8; PORT_PRINTK_INFO_DEV_INFO - PORT_PRINTK_INFO_CALLER_ID - 4], + dev_info: [u8; PORT_PRINTK_INFO_DEV_INFO_SIZE], +} + +kr::static_assert_layout!(PrintkInfo, + size = PORT_PRINTK_INFO_SIZE, align = PORT_PRINTK_INFO_ALIGN, + seq @ PORT_PRINTK_INFO_SEQ, ts_nsec @ PORT_PRINTK_INFO_TS_NSEC, + text_len @ PORT_PRINTK_INFO_TEXT_LEN, facility @ PORT_PRINTK_INFO_FACILITY, + flags_level @ (PORT_PRINTK_INFO_FACILITY + 1), + caller_id @ PORT_PRINTK_INFO_CALLER_ID, + execution_context @ (PORT_PRINTK_INFO_CALLER_ID + 4), + dev_info @ PORT_PRINTK_INFO_DEV_INFO, +); /// # Safety /// ring's arrays are valid; desc is private writable storage. unsafe fn finalized_seq(ring: *mut DescRing, id: c_ulong, seq: u64, desc: &mut Desc) -> c_int { let mut actual = 0; - // SAFETY: (U1) snapshot access and state validation use the C boundary. + // SAFETY: (U3) the C reader contract keeps ring arrays alive; desc and + // actual are private outputs for desc_read's snapshot. let s = unsafe { desc_read(ring, id, desc, &mut actual, null_mut()) }; + + // An unexpected @id (desc_miss) or @seq mismatch means the record + // does not exist. A descriptor in the reserved or committed state + // means the record does not yet exist for the reader. if s == MISS || s == RESERVED || s == COMMITTED || actual != seq { return EINVAL; } + + // A descriptor in the reusable state may no longer have its data + // available; report it as existing but with lost data. Or the record + // may actually be a record with lost data. if s == REUSABLE || (desc.text_blk_lpos.begin == FAILED_LPOS && desc.text_blk_lpos.next == FAILED_LPOS) { @@ -70,20 +224,14 @@ unsafe fn finalized_seq(ring: *mut DescRing, id: c_ulong, seq: u64, desc: &mut D /// # Safety /// text points to a C ring data range; its contents are speculative. unsafe fn count_lines(text: *const c_char, size: c_uint) -> c_uint { - // SAFETY: (U1) memchr performs the original C speculative text access. + // SAFETY: (U1) each byte snapshot uses the C READ_ONCE access contract. + // SAFETY: (U3) the caller supplies a live range of size bytes. unsafe { - let mut remaining = size as usize; - let mut next = text; let mut count = 1; - while remaining != 0 { - let found = c_prb_find(next.cast(), 10, remaining).cast::(); - if found.is_null() { - break; + for offset in 0..size as usize { + if c_printk_ringbuffer_read_u8(text.cast::().add(offset)) == b'\n' { + count += 1; } - count += 1; - let consumed = found.offset_from(next) as usize + 1; - remaining -= consumed; - next = found.add(1); } count } @@ -99,21 +247,34 @@ unsafe fn copy_data( size: c_uint, lines: *mut c_uint, ) -> bool { + // Caller might not want any data. if (buf.is_null() || size == 0) && lines.is_null() { return true; } // SAFETY: (U1) the C access layer copies only the range validated by get_data. + // SAFETY: (U3) optional output pointers are reader-owned, and ring geometry + // remains immutable throughout the C reader's lifetime. unsafe { let mut available = 0; let data = get_data(ring, lpos, &mut available); + + // Actual cannot be less than expected. It can be more than expected + // because of the trailing alignment padding. + // + // Note that invalid @len values can occur because the caller loads + // the value during an allowed data race. if data.is_null() || available < c_uint::from(len) { return false; } + + // Caller interested in the line count? if !lines.is_null() { *lines = count_lines(data, c_uint::from(len)); } + + // Caller interested in the data content? if !buf.is_null() && size != 0 { - c_prb_copy( + c_printk_ringbuffer_copy( buf.cast(), data.cast(), size.min(c_uint::from(len)) as usize, @@ -127,19 +288,32 @@ unsafe fn copy_data( /// rb is initialized; optional outputs are writable and do not alias shared data. unsafe fn read_record(rb: *mut Ringbuffer, seq: u64, r: *mut Record, lines: *mut c_uint) -> c_int { // SAFETY: (U1) speculative C copies are bracketed by descriptor validation. + // SAFETY: (U3) the C reader owns r and its output buffers; rb arrays remain + // alive while fields and descriptor addresses are accessed. unsafe { let ring = addr_of_mut!((*rb).desc_ring); let info = to_info(ring, seq); let d = to_desc(ring, seq); + + // Extract the ID, used to specify the descriptor to read. let id = load(addr_of!((*d).state_var)) & ID_MASK; let mut desc = empty_desc(); + + // Get a local copy of the correct descriptor (if available). let err = finalized_seq(ring, id, seq, &mut desc); + + // If @r is NULL, the caller is only interested in the availability + // of the record. if err != 0 || r.is_null() { return err; } + + // If requested, copy meta data. if !(*r).info.is_null() { - c_prb_copy((*r).info.cast(), info.cast(), size_of::()); + c_printk_ringbuffer_copy((*r).info.cast(), info.cast(), size_of::()); } + + // Copy text data. If it fails, this is a data-less record. if !copy_data( addr_of!((*rb).text_data_ring), &desc.text_blk_lpos, @@ -150,6 +324,8 @@ unsafe fn read_record(rb: *mut Ringbuffer, seq: u64, r: *mut Record, lines: *mut ) { return ENOENT; } + + // Ensure the record is still finalized and has the same @seq. finalized_seq(ring, id, seq, &mut desc) } } @@ -160,6 +336,7 @@ unsafe fn read_record(rb: *mut Ringbuffer, seq: u64, r: *mut Record, lines: *mut #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_first_seq(rb: *mut Ringbuffer) -> u64 { // SAFETY: (U1) the original C tail/descriptor barriers are preserved. + // SAFETY: (U3) the initialized ring's arrays remain alive for this reader. unsafe { let ring = addr_of_mut!((*rb).desc_ring); loop { @@ -167,9 +344,13 @@ pub unsafe extern "C" fn prb_first_seq(rb: *mut Ringbuffer) -> u64 { let mut seq = 0; let mut desc = empty_desc(); let s = desc_read(ring, id, &mut desc, &mut seq, null_mut()); + + // This loop will not be infinite because the tail is + // _always_ in the finalized or reusable state. if s == FINALIZED || s == REUSABLE { return seq; } + // Guarantee the last state load from desc_read() is before // reloading @tail_id in order to see a new tail in the case // that the descriptor has been recycled. This pairs with @@ -185,7 +366,7 @@ pub unsafe extern "C" fn prb_first_seq(rb: *mut Ringbuffer) -> u64 { // MB from desc_push_tail:B to desc_reserve:F // matching // RMB from prb_first_seq:B to prb_first_seq:A - c_prb_rmb(); // LMM(prb_first_seq:C) + c_printk_ringbuffer_rmb(); // LMM(prb_first_seq:C) } } } @@ -196,24 +377,71 @@ pub unsafe extern "C" fn prb_first_seq(rb: *mut Ringbuffer) -> u64 { #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_next_reserve_seq(rb: *mut Ringbuffer) -> u64 { // SAFETY: (U1) sequence snapshots use the C atomic publication protocol. + // SAFETY: (U3) the initialized ring's arrays remain alive for this reader. unsafe { let ring = addr_of_mut!((*rb).desc_ring); + + // It may not be possible to read a sequence number for @head_id. + // So the ID of @last_finailzed_seq is used to calculate what the + // sequence number of @head_id will be. loop { let seq = last_finalized(rb); + + // @head_id is loaded after @last_finalized_seq to ensure that + // it points to the record with @last_finalized_seq or newer. + // + // Memory barrier involvement: + // + // If desc_last_finalized_seq:A reads from + // desc_update_last_finalized:A, then + // prb_next_reserve_seq:A reads from desc_reserve:D. + // + // Relies on: + // + // RELEASE from desc_reserve:D to desc_update_last_finalized:A + // matching + // ACQUIRE from desc_last_finalized_seq:A to prb_next_reserve_seq:A + // + // Note: desc_reserve:D and desc_update_last_finalized:A can be + // different CPUs. However, the desc_update_last_finalized:A CPU + // (which performs the release) must have previously seen + // desc_read:C, which implies desc_reserve:D can be seen. let head = load(addr_of!((*ring).head_id)); let d = to_desc(ring, seq); + + // Extract the ID, used to specify the descriptor to read. let mut id = load(addr_of!((*d).state_var)) & ID_MASK; let mut desc = empty_desc(); + + // Ensure @last_finalized_id is correct. if finalized_seq(ring, id, seq, &mut desc) == EINVAL { + // Record must have been overwritten. Try again. if seq != 0 { continue; } + + // No record has been finalized or even reserved yet. + // + // The @head_id is initialized such that the first + // increment will yield the first record (seq=0). + // Handle it separately to avoid a negative @diff + // below. let initial = desc_count(ring).wrapping_add(1).wrapping_neg() & ID_MASK; if head == initial { return 0; } + + // One or more descriptors are already reserved. Use + // the descriptor ID of the first one (@seq=0) for + // the @diff below. id = initial.wrapping_add(1); } + + // Diff of known descriptor IDs to compute related sequence numbers. + + // @head_id points to the most recently reserved record, but this + // function returns the sequence number that will be assigned to the + // next (not yet reserved) record. Thus +1 is needed. return seq .wrapping_add(head.wrapping_sub(id) as u64) .wrapping_add(1); @@ -237,18 +465,45 @@ unsafe fn read_valid( return true; } let tail = prb_first_seq(rb); + + // Behind the tail. Catch up and try again. This + // can happen for -ENOENT and -EINVAL cases. if *seq < tail { *seq = tail; - } else if err == ENOENT - || (c_prb_panic_cpu() - && (read(addr_of!(debug_non_panic_cpus).cast::()) == 0 - || read(addr_of!(legacy_allow_panic_sync).cast::()) != 0) - && seq.wrapping_add(1) < prb_next_reserve_seq(rb)) + continue; + } + + // Record exists, but the data was lost. Skip. + if err == ENOENT { + *seq = seq.wrapping_add(1); + continue; + } + + // Non-existent/non-finalized record. Must stop. + // + // For panic situations it cannot be expected that + // non-finalized records will become finalized. But + // there may be other finalized records beyond that + // need to be printed for a panic situation. If this + // is the panic CPU, skip this + // non-existent/non-finalized record unless non-panic + // CPUs are still running and their debugging is + // explicitly enabled. + // + // Note that new messages printed on panic CPU are + // finalized when we are here. The only exception + // might be the last message without trailing newline. + // But it would have the sequence number returned + // by "prb_next_reserve_seq() - 1". + if c_printk_ringbuffer_panic_cpu() + && (read(addr_of!(debug_non_panic_cpus).cast::()) == 0 + || read(addr_of!(legacy_allow_panic_sync).cast::()) != 0) + && seq.wrapping_add(1) < prb_next_reserve_seq(rb) { *seq = seq.wrapping_add(1); - } else { - return false; + continue; } + return false; } } } @@ -258,7 +513,7 @@ unsafe fn read_valid( /// rb is initialized; r and its optional buffers are writable reader-owned storage. #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_read_valid(rb: *mut Ringbuffer, mut seq: u64, r: *mut Record) -> bool { - // SAFETY: (U1) the caller satisfies the C reader output-buffer contract. + // SAFETY: (U3) the caller satisfies the C reader output-buffer contract. unsafe { read_valid(rb, &mut seq, r, null_mut()) } } @@ -277,7 +532,7 @@ pub unsafe extern "C" fn prb_read_valid_info( text_buf: null_mut(), text_buf_size: 0, }; - // SAFETY: (U1) the local record forwards the caller's valid optional outputs. + // SAFETY: (U3) the local record forwards the caller's valid optional outputs. unsafe { read_valid(rb, &mut seq, &mut r, lines) } } @@ -287,7 +542,7 @@ pub unsafe extern "C" fn prb_read_valid_info( #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_first_valid_seq(rb: *mut Ringbuffer) -> u64 { let mut seq = 0; - // SAFETY: (U1) no output buffers are requested. + // SAFETY: (U3) rb remains initialized and alive; no outputs are requested. if unsafe { read_valid(rb, &mut seq, null_mut(), null_mut()) } { seq } else { @@ -300,12 +555,20 @@ pub unsafe extern "C" fn prb_first_valid_seq(rb: *mut Ringbuffer) -> u64 { /// rb and its arrays must remain alive and initialized. #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_next_seq(rb: *mut Ringbuffer) -> u64 { - // SAFETY: (U1) validated sequence reads follow C publication ordering. + // SAFETY: (U3) the C reader lifetime keeps rb and its arrays alive. unsafe { + // Begin searching after the last finalized record. + // + // On 0, the search must begin at 0 because of hack#2 + // of the bootstrapping phase it is not known if a + // record at index 0 exists. let mut seq = last_finalized(rb); if seq != 0 { seq = seq.wrapping_add(1); } + + // The information about the last finalized @seq might be inaccurate. + // Search forward to find the current one. while read_valid(rb, &mut seq, null_mut(), null_mut()) { seq = seq.wrapping_add(1); } @@ -324,26 +587,54 @@ pub unsafe extern "C" fn prb_reserve( r: *mut Record, ) -> bool { // SAFETY: (U1) C operations establish reservation ownership and save IRQ state. + // SAFETY: (U3) e and r are writer-owned. A successful descriptor reservation + // grants exclusive writes to its metadata and block positions. unsafe { let ring = addr_of_mut!((*rb).desc_ring); let data = addr_of_mut!((*rb).text_data_ring); if data_check_size(data, (*r).text_buf_size) { - c_prb_irq_save(addr_of_mut!((*e).irqflags)); + // Descriptors in the reserved state act as blockers to all further + // reservations once the desc_ring has fully wrapped. Disable + // interrupts during the reserve/commit window in order to minimize + // the likelihood of this happening. + c_printk_ringbuffer_irq_save(addr_of_mut!((*e).irqflags)); let mut id = 0; if desc_reserve(rb, &mut id) { let d = to_desc(ring, id as u64); let info = to_info(ring, id as u64); + + // All @info fields (except @seq) are cleared and must be filled in + // by the writer. Save @seq before clearing because it is used to + // determine the new sequence number. let old_seq = read(addr_of!((*info).seq)); - c_prb_clear(info.cast(), 0, size_of::()); + c_printk_ringbuffer_clear(info.cast(), 0, size_of::()); + + // Set the @e fields here so that prb_commit() can be used if + // text data allocation fails. (*e).rb = rb; (*e).id = id; let index = id & (desc_count(ring) - 1); + + // Initialize the sequence number if it has "never been set". + // Otherwise just increment it by a full wrap. + // + // @seq is considered "never been set" if it has a value of 0, + // _except_ for @infos[0], which was specially setup by the ringbuffer + // initializer and therefore is always considered as set. + // + // See the "Bootstrap" comment block in printk_ringbuffer.h for + // details about how the initializer bootstraps the descriptors. let seq = if old_seq == 0 && index != 0 { index as u64 } else { old_seq.wrapping_add(desc_count(ring) as u64) }; write(addr_of_mut!((*info).seq), seq); + + // New data is about to be reserved. Once that happens, previous + // descriptors are no longer able to be extended. Finalize the + // previous descriptor now so that it can be made available to + // readers. (For seq==0 there is no previous descriptor.) if seq != 0 { desc_make_final(rb, id.wrapping_sub(1) & ID_MASK); } @@ -355,16 +646,25 @@ pub unsafe extern "C" fn prb_reserve( begin: read(addr_of!((*d).text_blk_lpos.begin)), next: read(addr_of!((*d).text_blk_lpos.next)), }; + + // Record full text space used by record. (*e).text_space = space_used(data, &pos); return true; } + + // If text data allocation fails, a data-less record is committed. prb_commit(e); + + // prb_commit() re-enabled interrupts. } else { - c_prb_inc(addr_of_mut!((*rb).fail)); - c_prb_irq_restore((*e).irqflags); + // Descriptor reservation failures are tracked. + c_printk_ringbuffer_inc(addr_of_mut!((*rb).fail)); + c_printk_ringbuffer_irq_restore((*e).irqflags); } } - c_prb_clear(r.cast(), 0, size_of::()); + + // Make it clear to the caller that the reserve failed. + c_printk_ringbuffer_clear(r.cast(), 0, size_of::()); false } } @@ -381,19 +681,30 @@ pub unsafe extern "C" fn prb_reserve_in_last( max: c_uint, ) -> bool { // SAFETY: (U1) the C compare-exchange obtains writer ownership before updates. + // SAFETY: (U3) e and r are writer-owned; after reopening, the C reservation + // contract grants exclusive metadata and block-position writes. unsafe { - c_prb_irq_save(addr_of_mut!((*e).irqflags)); + c_printk_ringbuffer_irq_save(addr_of_mut!((*e).irqflags)); let ring = addr_of_mut!((*rb).desc_ring); let data = addr_of_mut!((*rb).text_data_ring); let mut id = 0; + + // Transition the newest descriptor back to the reserved state. let d = desc_reopen_last(ring, caller_id, &mut id); if d.is_null() { - c_prb_irq_restore((*e).irqflags); + c_printk_ringbuffer_irq_restore((*e).irqflags); } else { + // Now the writer has exclusive access: LMM(prb_reserve_in_last:A) let info = to_info(ring, id as u64); + + // Set the @e fields here so that prb_commit() can be used if + // anything fails from now on. (*e).rb = rb; (*e).id = id; let success = 'reserve: { + // desc_reopen_last() checked the caller_id, but there was no + // exclusive access at that point. The descriptor may have + // changed since then. if caller_id != read(addr_of!((*info).caller_id)) { break 'reserve false; } @@ -403,8 +714,8 @@ pub unsafe extern "C" fn prb_reserve_in_last( }; let mut len = read(addr_of!((*info).text_len)); if dataless(&pos) { - if c_prb_warn_9(len != 0) { - c_prb_warn_len_zero(len); + if c_printk_ringbuffer_warn_empty_text_length(len != 0) { + c_printk_ringbuffer_warn_len_zero(len); write(addr_of_mut!((*info).text_len), 0u16); } if !data_check_size(data, (*r).text_buf_size) || (*r).text_buf_size > max { @@ -413,12 +724,15 @@ pub unsafe extern "C" fn prb_reserve_in_last( (*r).text_buf = data_alloc(rb, (*r).text_buf_size, addr_of_mut!((*d).text_blk_lpos), id); } else { + // Increase the buffer size to include the original size. If + // the meta data (@text_len) is not sane, use the full data + // block size. let mut size = 0; if get_data(data, &pos, &mut size).is_null() { break 'reserve false; } - if c_prb_warn_10(c_uint::from(len) > size) { - c_prb_warn_len_max(len, size); + if c_printk_ringbuffer_warn_text_length(c_uint::from(len) > size) { + c_printk_ringbuffer_warn_len_max(len, size); len = size as u16; write(addr_of_mut!((*info).text_len), len); } @@ -444,8 +758,12 @@ pub unsafe extern "C" fn prb_reserve_in_last( return true; } prb_commit(e); + + // prb_commit() re-enabled interrupts. } - c_prb_clear(r.cast(), 0, size_of::()); + + // Make it clear to the caller that the re-reserve failed. + c_printk_ringbuffer_clear(r.cast(), 0, size_of::()); false } } @@ -464,25 +782,26 @@ pub unsafe extern "C" fn prb_init( infos: *mut PrintkInfo, ) { // SAFETY: (U7) the caller supplies exclusively owned C storage for initialization. + // SAFETY: (U1) memset and atomic_long_set receive valid caller-owned ranges. unsafe { let count = 1usize << descbits; let id = (count as c_ulong).wrapping_add(1).wrapping_neg() & ID_MASK; - c_prb_clear(descs.cast(), 0, count * size_of::()); - c_prb_clear(infos.cast(), 0, count * size_of::()); + c_printk_ringbuffer_clear(descs.cast(), 0, count * size_of::()); + c_printk_ringbuffer_clear(infos.cast(), 0, count * size_of::()); (*rb).desc_ring.count_bits = descbits; (*rb).desc_ring.descs = descs; (*rb).desc_ring.infos = infos; - c_prb_store(addr_of_mut!((*rb).desc_ring.head_id), id as c_long); - c_prb_store(addr_of_mut!((*rb).desc_ring.tail_id), id as c_long); - c_prb_store(addr_of_mut!((*rb).desc_ring.last_finalized_seq), 0); + c_printk_ringbuffer_store(addr_of_mut!((*rb).desc_ring.head_id), id as c_long); + c_printk_ringbuffer_store(addr_of_mut!((*rb).desc_ring.tail_id), id as c_long); + c_printk_ringbuffer_store(addr_of_mut!((*rb).desc_ring.last_finalized_seq), 0); (*rb).text_data_ring.size_bits = textbits; (*rb).text_data_ring.data = text; let lpos = (1 as c_ulong).wrapping_shl(textbits).wrapping_neg(); - c_prb_store(addr_of_mut!((*rb).text_data_ring.head_lpos), lpos as c_long); - c_prb_store(addr_of_mut!((*rb).text_data_ring.tail_lpos), lpos as c_long); - c_prb_store(addr_of_mut!((*rb).fail), 0); + c_printk_ringbuffer_store(addr_of_mut!((*rb).text_data_ring.head_lpos), lpos as c_long); + c_printk_ringbuffer_store(addr_of_mut!((*rb).text_data_ring.tail_lpos), lpos as c_long); + c_printk_ringbuffer_store(addr_of_mut!((*rb).fail), 0); let last = descs.add(count - 1); - c_prb_store(addr_of_mut!((*last).state_var), sv(id, REUSABLE) as c_long); + c_printk_ringbuffer_store(addr_of_mut!((*last).state_var), sv(id, REUSABLE) as c_long); (*last).text_blk_lpos.begin = FAILED_LPOS; (*last).text_blk_lpos.next = FAILED_LPOS; (*infos).seq = (count as u64).wrapping_neg(); @@ -498,13 +817,6 @@ pub unsafe extern "C" fn prb_record_text_space(e: *mut ReservedEntry) -> c_uint // SAFETY: (U3) the writer owns this initialized reservation handle. unsafe { (*e).text_space } } -const COMMITTED: i32 = 1; -const FINALIZED: i32 = 2; -const REUSABLE: i32 = 3; -const MISS: i32 = -1; -const FLAGS_SHIFT: u32 = c_ulong::BITS - 2; -const ID_MASK: c_ulong = c_ulong::MAX >> 2; -const FAILED_LPOS: c_ulong = 1; fn dataless(lpos: &DataBlkLpos) -> bool { lpos.begin & lpos.next & 1 != 0 } @@ -513,6 +825,9 @@ fn dataless(lpos: &DataBlkLpos) -> bool { /// ring is an initialized data ring. unsafe fn wrapped(ring: *const DataRing, begin: c_ulong, next: c_ulong) -> bool { // SAFETY: (U3) size_bits is immutable for the lifetime of the ring. + + // Subtract one from next_lpos since it's not actually part of this data + // block. This allows perfectly fitting records to not wrap. unsafe { begin >> (*ring).size_bits != next.wrapping_sub(1) >> (*ring).size_bits } } @@ -522,9 +837,11 @@ unsafe fn next_lpos(ring: *const DataRing, begin: c_ulong, size: c_uint) -> c_ul let next = begin.wrapping_add(c_ulong::from(size)); // SAFETY: (U3) only the ring's immutable geometry is inspected. unsafe { + // First check if the data block does not wrap. if !wrapped(ring, begin, next) { next } else { + // Wrapping data blocks store their data at the beginning. (next & !(data_size(ring) - 1)).wrapping_add(c_ulong::from(size)) } } @@ -539,7 +856,12 @@ unsafe fn data_alloc( id: c_ulong, ) -> *mut c_char { // SAFETY: (U1) payload writes/copies and state changes use the C access layer. + // SAFETY: (U3) immutable ring geometry bounds the addresses; the reserved + // descriptor owns lpos and a successful head update owns the data block. unsafe { + // Data blocks are not created for empty lines. Instead, the + // reader will recognize these special lpos values and handle + // it appropriately. if size == 0 { write(addr_of_mut!((*lpos).begin), EMPTY_LINE_LPOS); write(addr_of_mut!((*lpos).next), EMPTY_LINE_LPOS); @@ -550,16 +872,22 @@ unsafe fn data_alloc( let mut begin = load(addr_of!((*ring).head_lpos)); let next = loop { let next = next_lpos(ring, begin, size); - if c_prb_warn_2(next.wrapping_sub(begin) > data_size(ring)) + + // data_check_size() prevents data block allocation that could + // cause illegal ringbuffer states. But double check that the + // used space will not be bigger than the ring buffer. Wrapped + // messages need to reserve more space, see get_next_lpos(). + // + // Specify a data-less block when the check or the allocation + // fails. + if c_printk_ringbuffer_warn_allocation_size(next.wrapping_sub(begin) > data_size(ring)) || !data_push_tail(rb, next.wrapping_sub(data_size(ring))) { write(addr_of_mut!((*lpos).begin), FAILED_LPOS); write(addr_of_mut!((*lpos).next), FAILED_LPOS); return null_mut(); } - if cas(addr_of_mut!((*ring).head_lpos), &mut begin, next) { - break next; - } + // 1. Guarantee any descriptor states that have transitioned // to reusable are stored before modifying the newly // allocated data area. A full memory barrier is needed @@ -575,11 +903,20 @@ unsafe fn data_alloc( // load a new tail lpos. A full memory barrier is needed // since other CPUs may have updated the tail lpos. This // pairs with data_push_tail:B. - }; // LMM(data_alloc:A) + // LMM(data_alloc:A) + if cas(addr_of_mut!((*ring).head_lpos), &mut begin, next) { + break next; + } + }; let mut block = to_block(ring, begin); write(block, id); // LMM(data_alloc:B) + + // Wrapping data blocks store their data at the beginning. if wrapped(ring, begin, next) { block = to_block(ring, 0); + + // Store the ID on the wrapped block for consistency. + // The printk_ringbuffer does not actually use it. write(block, id); } write(addr_of_mut!((*lpos).begin), begin); @@ -597,36 +934,67 @@ unsafe fn data_realloc( id: c_ulong, ) -> *mut c_char { // SAFETY: (U1) C accessors preserve the ring protocol for recycled memory. + // SAFETY: (U3) the reopened descriptor owns lpos; head CAS extends only + // its newest block within the live C text allocation. unsafe { let ring = addr_of_mut!((*rb).text_data_ring); let begin = read(addr_of!((*lpos).begin)); let end = read(addr_of!((*lpos).next)); let mut head = load(addr_of!((*ring).head_lpos)); + + // Reallocation only works if @blk_lpos is the newest data block. if head != end { return null_mut(); } + + // Keep track if @blk_lpos was a wrapping data block. let was_wrapped = wrapped(ring, begin, end); let next = next_lpos(ring, begin, to_blk_size(size)); + + // Use the current data block when the size does not increase, i.e. + // when @head_lpos is already able to accommodate the new @next_lpos. + // + // Note that need_more_space() could never return false here because + // the difference between the positions was bigger than the data + // buffer size. The data block is reopened and can't get reused. if !need_more_space(ring, head, next) { return to_block(ring, if was_wrapped { 0 } else { begin }) .add(1) .cast(); } - if c_prb_warn_3(next.wrapping_sub(begin) > data_size(ring)) + + // data_check_size() prevents data block reallocation that could + // cause illegal ringbuffer states. But double check that the + // new used space will not be bigger than the ring buffer. Wrapped + // messages need to reserve more space, see get_next_lpos(). + // + // Specify failure when the check or the allocation fails. + if c_printk_ringbuffer_warn_reallocation_size(next.wrapping_sub(begin) > data_size(ring)) || !data_push_tail(rb, next.wrapping_sub(data_size(ring))) { return null_mut(); } + + // The memory barrier involvement is the same as data_alloc:A. if !cas(addr_of_mut!((*ring).head_lpos), &mut head, next) { return null_mut(); } let mut block = to_block(ring, begin); + + // Wrapping data blocks store their data at the beginning. if wrapped(ring, begin, next) { let old = block; block = to_block(ring, 0); + + // Store the ID on the wrapped block for consistency. + // The printk_ringbuffer does not actually use it. write(block, id); + + // Since the allocated space is now in the newly + // created wrapping data block, copy the content + // from the old data block. if !was_wrapped { - c_prb_copy( + c_printk_ringbuffer_copy( block.add(1).cast(), old.add(1).cast(), end.wrapping_sub(begin) as usize - size_of::(), @@ -641,6 +1009,7 @@ unsafe fn data_realloc( /// # Safety /// ring is alive, lpos is a validated local snapshot or reserved block. unsafe fn space_used(ring: *const DataRing, lpos: &DataBlkLpos) -> c_uint { + // Data-less blocks take no space. if dataless(lpos) { return 0; } @@ -649,9 +1018,13 @@ unsafe fn space_used(ring: *const DataRing, lpos: &DataBlkLpos) -> c_uint { let size = data_size(ring); let begin = lpos.begin & (size - 1); let end = lpos.next & (size - 1); + + // Data block does not wrap. if !wrapped(ring, lpos.begin, lpos.next) { end.wrapping_sub(begin) as c_uint } else { + // For wrapping data blocks, the trailing (wasted) space is + // also counted. end.wrapping_add(size).wrapping_sub(begin) as c_uint } } @@ -661,34 +1034,57 @@ unsafe fn space_used(ring: *const DataRing, lpos: &DataBlkLpos) -> c_uint { /// ring is initialized. lpos is a snapshot; returned memory remains speculative. unsafe fn get_data(ring: *const DataRing, lpos: &DataBlkLpos, size: &mut c_uint) -> *const c_char { // SAFETY: (U1) warnings use the C boundary; block geometry is checked before use. + // SAFETY: (U3) immutable C ring geometry and checked positions bound the + // returned pointer; no shared payload is dereferenced here. unsafe { + // Data-less data block description. if dataless(lpos) { + // Records that are just empty lines are also valid, even + // though they do not have a data block. For such records + // explicitly return empty string data to signify success. if lpos.begin == EMPTY_LINE_LPOS && lpos.next == EMPTY_LINE_LPOS { *size = 0; return c"".to_bytes_with_nul().as_ptr().cast(); } + + // Data lost, invalid, or otherwise unavailable. return null_mut(); } let block; + + // Regular data block: @begin and @next in the same wrap. if !wrapped(ring, lpos.begin, lpos.next) { block = to_block(ring, lpos.begin); *size = lpos.next.wrapping_sub(lpos.begin) as c_uint; } else if !wrapped(ring, lpos.begin.wrapping_add(data_size(ring)), lpos.next) { + // Wrapping data block: @begin is one wrap behind @next. block = to_block(ring, 0); *size = (lpos.next & (data_size(ring) - 1)) as c_uint; } else { - c_prb_warn_4(true); + // Illegal block description. + c_printk_ringbuffer_warn_block_wrap(true); return null_mut(); } + + // A valid data block will always be aligned to the ID size. let mask = size_of::() as c_ulong - 1; - if c_prb_warn_5(lpos.begin & mask != 0) || c_prb_warn_6(lpos.next & mask != 0) { + if c_printk_ringbuffer_warn_block_begin_alignment(lpos.begin & mask != 0) + || c_printk_ringbuffer_warn_block_next_alignment(lpos.next & mask != 0) + { return null_mut(); } - if c_prb_warn_7(*size as usize <= size_of::()) { + + // A regular data block will always have an ID and at least + // 1 byte of data. Data-less blocks were handled earlier. + if c_printk_ringbuffer_warn_block_minimum_size(*size as usize <= size_of::()) { return null_mut(); } + + // Subtract block ID space from size to reflect data size. *size -= size_of::() as c_uint; - if c_prb_warn_8(!data_check_size(ring, *size)) { + + // Sanity check the max size of the regular data block. + if c_printk_ringbuffer_warn_block_maximum_size(!data_check_size(ring, *size)) { return null_mut(); } block.add(1).cast() @@ -699,15 +1095,34 @@ unsafe fn get_data(ring: *const DataRing, lpos: &DataBlkLpos, size: &mut c_uint) /// ring is alive; caller requests ownership of its last committed descriptor. unsafe fn desc_reopen_last(ring: *mut DescRing, caller: u32, id_out: &mut c_ulong) -> *mut Desc { // SAFETY: (U1) the C compare-exchange acquires the reserved descriptor state. + // SAFETY: (U3) ring arrays remain live; id_out is private writer storage. unsafe { let id = load(addr_of!((*ring).head_id)); let mut desc = empty_desc(); let mut cid = 0; + + // To reduce unnecessarily reopening, first check if the descriptor + // state and caller ID are correct. if desc_read(ring, id, &mut desc, null_mut(), &mut cid) != COMMITTED || cid != caller { return null_mut(); } let d = to_desc(ring, id as u64); let mut old = sv(id, COMMITTED); + + // Guarantee the reserved state is stored before reading any + // record data. A full memory barrier is needed because @state_var + // modification is followed by reading. This pairs with _prb_commit:B. + // + // Memory barrier involvement: + // + // If desc_reopen_last:A reads from _prb_commit:B, then + // prb_reserve_in_last:A reads from _prb_commit:A. + // + // Relies on: + // + // WMB from _prb_commit:A to _prb_commit:B + // matching + // MB from desc_reopen_last:A to prb_reserve_in_last:A if !cas(addr_of_mut!((*d).state_var), &mut old, sv(id, RESERVED)) { return null_mut(); } @@ -720,8 +1135,13 @@ unsafe fn desc_reopen_last(ring: *mut DescRing, caller: u32, id_out: &mut c_ulon /// rb is an initialized ringbuffer that remains alive. unsafe fn last_finalized(rb: *mut Ringbuffer) -> u64 { // SAFETY: (U1) acquire load pairs with the C release publication protocol. + // SAFETY: (U3) rb remains alive and initialized under the C reader contract. unsafe { - let seq = c_prb_load_acquire(addr_of!((*rb).desc_ring.last_finalized_seq)) as c_ulong; + // Guarantee the sequence number is loaded before loading the + // associated record in order to guarantee that the record can be + // seen by this CPU. This pairs with desc_update_last_finalized:A. + let seq = c_printk_ringbuffer_load_acquire(addr_of!((*rb).desc_ring.last_finalized_seq)) + as c_ulong; #[cfg(target_pointer_width = "64")] { seq as u64 @@ -738,20 +1158,48 @@ unsafe fn last_finalized(rb: *mut Ringbuffer) -> u64 { /// rb remains alive and contains initialized arrays. unsafe fn update_last_finalized(rb: *mut Ringbuffer) { // SAFETY: (U1) release compare-exchange publishes only validated sequence numbers. + // SAFETY: (U3) rb remains alive and initialized during descriptor publication. unsafe { let mut old_seq = last_finalized(rb); loop { let mut finalized = old_seq; let mut next = finalized.wrapping_add(1); + + // Try to find later finalized records. while read_valid(rb, &mut next, null_mut(), null_mut()) { finalized = next; next = next.wrapping_add(1); } + + // No update needed if no later finalized record was found. if finalized == old_seq { return; } let mut old = old_seq as c_long; - if c_prb_cas_release( + + // Set the sequence number of a later finalized record that has been + // seen. + // + // Guarantee the record data is visible to other CPUs before storing + // its sequence number. This pairs with desc_last_finalized_seq:A. + // + // Memory barrier involvement: + // + // If desc_last_finalized_seq:A reads from + // desc_update_last_finalized:A, then desc_read:A reads from + // _prb_commit:B. + // + // Relies on: + // + // RELEASE from _prb_commit:B to desc_update_last_finalized:A + // matching + // ACQUIRE from desc_last_finalized_seq:A to desc_read:A + // + // Note: _prb_commit:B and desc_update_last_finalized:A can be + // different CPUs. However, the desc_update_last_finalized:A + // CPU (which performs the release) must have previously seen + // _prb_commit:B. + if c_printk_ringbuffer_cas_release( addr_of_mut!((*rb).desc_ring.last_finalized_seq), &mut old, finalized as c_long, @@ -775,10 +1223,11 @@ unsafe fn update_last_finalized(rb: *mut Ringbuffer) { /// rb is alive; id refers to a descriptor that may be finalized. unsafe fn desc_make_final(rb: *mut Ringbuffer, id: c_ulong) { // SAFETY: (U1) the C compare-exchange respects descriptor ID and state. + // SAFETY: (U3) immutable ring geometry selects a live descriptor slot. unsafe { let d = to_desc(addr_of_mut!((*rb).desc_ring), id as u64); let mut old = sv(id, COMMITTED) as c_long; - if c_prb_cas_relaxed( + if c_printk_ringbuffer_cas_relaxed( addr_of_mut!((*d).state_var), &mut old, sv(id, FINALIZED) as c_long, @@ -792,13 +1241,40 @@ unsafe fn desc_make_final(rb: *mut Ringbuffer, id: c_ulong) { /// e owns a successfully reserved entry with saved IRQ flags. unsafe fn commit(e: *mut ReservedEntry, s: i32) { // SAFETY: (U1) C atomic commit publishes the writer's data before restoring IRQs. + // SAFETY: (U3) e is a private successful reservation and rb remains live; + // its saved id and IRQ flags remain readable after publication. unsafe { + // Now the writer has finished all writing: LMM(_prb_commit:A) let d = to_desc(addr_of_mut!((*(*e).rb).desc_ring), (*e).id as u64); let mut old = sv((*e).id, RESERVED); + + // Set the descriptor as committed. See "ABA Issues" about why + // cmpxchg() instead of set() is used. + // + // 1 Guarantee all record data is stored before the descriptor state + // is stored as committed. A write memory barrier is sufficient + // for this. This pairs with desc_read:B and desc_reopen_last:A. + // + // 2. Guarantee the descriptor state is stored as committed before + // re-checking the head ID in order to possibly finalize this + // descriptor. This pairs with desc_reserve:D. + // + // Memory barrier involvement: + // + // If prb_commit:A reads from desc_reserve:D, then + // desc_make_final:A reads from _prb_commit:B. + // + // Relies on: + // + // MB from _prb_commit:B to prb_commit:A + // matching + // MB from desc_reserve:D to desc_make_final:A if !cas(addr_of_mut!((*d).state_var), &mut old, sv((*e).id, s)) { - c_prb_warn_11(true); + c_printk_ringbuffer_warn_commit(true); } - c_prb_irq_restore((*e).irqflags); + + // Restore interrupts, the reserve/commit window is finished. + c_printk_ringbuffer_irq_restore((*e).irqflags); } } @@ -808,9 +1284,14 @@ unsafe fn commit(e: *mut ReservedEntry, s: i32) { #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_commit(e: *mut ReservedEntry) { // SAFETY: (U1) the caller holds the reservation through the C commit boundary. + // SAFETY: (U3) e remains caller-owned after commit; its ring remains alive. unsafe { let rb = (*e).rb; commit(e, COMMITTED); + + // If this descriptor is no longer the head (i.e. a new record has + // been allocated), extending the data for this record is no longer + // allowed and therefore it must be finalized. if load(addr_of!((*rb).desc_ring.head_id)) != (*e).id { desc_make_final(rb, (*e).id); } @@ -822,33 +1303,53 @@ pub unsafe extern "C" fn prb_commit(e: *mut ReservedEntry) { /// e must be a successfully reserved entry belonging to this execution context. #[unsafe(no_mangle)] pub unsafe extern "C" fn prb_final_commit(e: *mut ReservedEntry) { - // SAFETY: (U1) the caller owns the reservation being finalized. + // SAFETY: (U3) e remains caller-owned and its ring remains alive after commit. unsafe { commit(e, FINALIZED); update_last_finalized((*e).rb); } } -const EMPTY_LINE_LPOS: c_ulong = 3; -// Snapshots of shared C data must consist only of initialized integer bytes. -// This helper never creates a reference to the C allocation. +// Scalar snapshots use marked C loads. Ordinary racing memcpy loads may +// produce LLVM undef; converting such a snapshot into a Rust integer before +// descriptor validation would not establish Rust value validity. +trait SnapshotScalar: Copy { + /// # Safety + /// src points to a live, initialized C scalar with the required alignment. + unsafe fn snapshot(src: *const Self) -> Self; +} + +macro_rules! snapshot_scalar { + ($ty:ty, $read:ident) => { + impl SnapshotScalar for $ty { + /// # Safety + /// src points to a live, initialized and properly aligned C scalar. + unsafe fn snapshot(src: *const Self) -> Self { + // SAFETY: (U1) the caller supplies a live aligned C scalar; + // READ_ONCE returns its target-defined integer snapshot. + unsafe { $read(src) } + } + } + }; +} + +snapshot_scalar!(u8, c_printk_ringbuffer_read_u8); +snapshot_scalar!(u16, c_printk_ringbuffer_read_u16); +snapshot_scalar!(u32, c_printk_ringbuffer_read_u32); +snapshot_scalar!(u64, c_printk_ringbuffer_read_u64); + /// # Safety -/// Both ranges must be valid for T and T must admit all source bit patterns. -/// The C memcpy access must satisfy the documented LKMM snapshot contract. -unsafe fn read(src: *const T) -> T { - let mut out = MaybeUninit::::uninit(); - // SAFETY: (U1) the caller provides a valid C snapshot source and local output. - unsafe { c_prb_copy(out.as_mut_ptr().cast(), src.cast(), size_of::()) }; - // SAFETY: (U3) the private snapshot is initialized by memcpy above; the - // caller guarantees that every source bit pattern is a valid T. - unsafe { out.as_ptr().read() } +/// src points to a live, initialized C scalar with the required alignment. +unsafe fn read(src: *const T) -> T { + // SAFETY: (U1) the sealed set of integer implementations forwards READ_ONCE. + unsafe { T::snapshot(src) } } /// # Safety /// dst is writable C storage for T; the caller owns the reservation or init. unsafe fn write(dst: *mut T, value: T) { // SAFETY: (U1) memcpy forwards the reservation owner's scalar write. - unsafe { c_prb_copy(dst.cast(), addr_of!(value).cast(), size_of::()) }; + unsafe { c_printk_ringbuffer_copy(dst.cast(), addr_of!(value).cast(), size_of::()) }; } fn sv(id: c_ulong, state: i32) -> c_ulong { @@ -866,7 +1367,7 @@ fn state(id: c_ulong, value: c_ulong) -> i32 { /// p points to an initialized atomic_long_t governed by the C ring protocol. unsafe fn load(p: *const c_long) -> c_ulong { // SAFETY: (U1) atomic_long_read preserves the kernel's LKMM access. - unsafe { c_prb_load(p) as c_ulong } + unsafe { c_printk_ringbuffer_load(p) as c_ulong } } /// # Safety @@ -874,7 +1375,7 @@ unsafe fn load(p: *const c_long) -> c_ulong { unsafe fn cas(p: *mut c_long, old: &mut c_ulong, new: c_ulong) -> bool { let mut expected = *old as c_long; // SAFETY: (U1) C updates only the supplied atomic and local expected value. - let ok = unsafe { c_prb_cas(p, &mut expected, new as c_long) }; + let ok = unsafe { c_printk_ringbuffer_cas(p, &mut expected, new as c_long) }; *old = expected as c_ulong; ok } @@ -936,6 +1437,12 @@ fn to_blk_size(size: c_uint) -> c_uint { /// ring is a valid initialized data ring. unsafe fn data_check_size(ring: *const DataRing, size: c_uint) -> bool { // SAFETY: (U3) size_bits is immutable. + + // Data-less blocks take no space. + + // If data blocks were allowed to be larger than half the data ring + // size, a wrapping data block could require more space than the full + // ringbuffer. size == 0 || c_ulong::from(to_blk_size(size)) <= unsafe { data_size(ring) } / 2 } @@ -957,9 +1464,13 @@ unsafe fn desc_read( ) -> i32 { // SAFETY: (U1) all shared snapshots and barriers use the C LKMM boundary; // outputs belong to this reader. No shared payload reference is formed. + // SAFETY: (U3) immutable ring geometry selects live C descriptor and metadata + // slots; optional outputs are private to this reader. unsafe { let desc = to_desc(ring, id as u64); let info = to_info(ring, id as u64); + + // Check the descriptor state. let mut value = load(addr_of!((*desc).state_var)); // LMM(desc_read:A) let mut s = state(id, value); if s != MISS && s != RESERVED { @@ -977,23 +1488,23 @@ unsafe fn desc_read( // WMB from _prb_commit:A to _prb_commit:B // matching // RMB from desc_read:A to desc_read:C - c_prb_rmb(); // LMM(desc_read:B) + c_printk_ringbuffer_rmb(); // LMM(desc_read:B) + + // Copy the descriptor data. The data is not valid until the + // state has been re-checked. A memcpy() for all of @desc + // cannot be used because of the atomic_t @state_var field. + // LMM(desc_read:C) if !out.is_null() { - c_prb_copy( - addr_of_mut!((*out).text_blk_lpos).cast(), - addr_of!((*desc).text_blk_lpos).cast(), - size_of::(), - ); - // Copy the descriptor data. The data is not valid until the - // state has been re-checked. A memcpy() for all of @desc - // cannot be used because of the atomic_t @state_var field. - } // LMM(desc_read:C) + (*out).text_blk_lpos.begin = read(addr_of!((*desc).text_blk_lpos.begin)); + (*out).text_blk_lpos.next = read(addr_of!((*desc).text_blk_lpos.next)); + } if !seq.is_null() { *seq = read(addr_of!((*info).seq)); } if !caller.is_null() { *caller = read(addr_of!((*info).caller_id)); } + // 1. Guarantee the descriptor content is loaded before re-checking // the state. This avoids reading an obsolete descriptor state // that may not apply to the copied content. This pairs with @@ -1030,13 +1541,17 @@ unsafe fn desc_read( // CPUs. However, the data_alloc:B CPU (which performs the // full memory barrier) must have previously seen // desc_make_reusable:A. - c_prb_rmb(); // LMM(desc_read:D) - // The data has been copied. Return the current descriptor state, - // which may have changed since the load above. + c_printk_ringbuffer_rmb(); // LMM(desc_read:D) + + // The data has been copied. Return the current descriptor state, + // which may have changed since the load above. value = load(addr_of!((*desc).state_var)); // LMM(desc_read:E) s = state(id, value); } if !out.is_null() { + // The descriptor is in an inconsistent state. Set at least + // @state_var so that the caller can see the details of + // the inconsistent state. (*out).state_var = value as c_long; } s @@ -1047,10 +1562,11 @@ unsafe fn desc_read( /// ring and its descriptors remain alive throughout the transition. unsafe fn desc_make_reusable(ring: *mut DescRing, id: c_ulong) { // SAFETY: (U1) the C cmpxchg conditionally invalidates the requested ID. + // SAFETY: (U3) immutable ring geometry selects a live descriptor slot. unsafe { let d = to_desc(ring, id as u64); let mut old = sv(id, FINALIZED) as c_long; - c_prb_cas_relaxed( + c_printk_ringbuffer_cas_relaxed( addr_of_mut!((*d).state_var), &mut old, sv(id, REUSABLE) as c_long, @@ -1074,9 +1590,12 @@ unsafe fn data_make_reusable( out: &mut c_ulong, ) -> bool { // SAFETY: (U1) C accessors read racing payload and preserve descriptor ordering. + // SAFETY: (U3) rb arrays remain alive; out and descriptor snapshots are local. unsafe { let data = addr_of_mut!((*rb).text_data_ring); let ring = addr_of_mut!((*rb).desc_ring); + + // Loop until @lpos_begin has advanced to or beyond @lpos_end. while need_more_space(data, begin, end) { // Load the block ID from the data block. This is a data race // against a writer that may have newly reserved this data @@ -1088,18 +1607,24 @@ unsafe fn data_make_reusable( let mut desc = empty_desc(); match desc_read(ring, id, &mut desc, null_mut(), null_mut()) { FINALIZED => { + // This data block is invalid if the descriptor + // does not point back to it. if desc.text_blk_lpos.begin != begin { return false; } desc_make_reusable(ring, id); } REUSABLE => { + // This data block is invalid if the descriptor + // does not point back to it. if desc.text_blk_lpos.begin != begin { return false; } } _ => return false, } + + // Advance @lpos_begin to the next data block. begin = desc.text_blk_lpos.next; } *out = begin; @@ -1110,12 +1635,15 @@ unsafe fn data_make_reusable( /// # Safety /// rb and arrays remain valid throughout concurrent recycling. unsafe fn data_push_tail(rb: *mut Ringbuffer, lpos: c_ulong) -> bool { + // If @lpos is from a data-less block, there is nothing to do. if lpos & 1 != 0 { return true; } // SAFETY: (U1) the C atomics and barriers preserve the original tail protocol. + // SAFETY: (U3) the C ring lifetime keeps its immutable geometry and arrays live. unsafe { let ring = addr_of_mut!((*rb).text_data_ring); + // Any descriptor states that have transitioned to reusable due to the // data tail being pushed to this loaded value will be visible to this // CPU. This pairs with data_push_tail:D. @@ -1133,8 +1661,17 @@ unsafe fn data_push_tail(rb: *mut Ringbuffer, lpos: c_ulong) -> bool { // thus // READFROM from desc_make_reusable:A to this CPU let mut tail = load(addr_of!((*ring).tail_lpos)); // LMM(data_push_tail:A) + + // Loop until the tail lpos is at or beyond @lpos. This condition + // may already be satisfied, resulting in no full memory barrier + // from data_push_tail:D being performed. However, since this CPU + // sees the new tail lpos, any descriptor states that transitioned to + // the reusable state must already be visible. while need_more_space(ring, tail, lpos) { let mut next = 0; + + // Make all descriptors reusable that are associated with + // data blocks before @lpos. if !data_make_reusable(rb, tail, lpos, &mut next) { // 1. Guarantee the block ID loaded in // data_make_reusable() is performed before @@ -1189,23 +1726,27 @@ unsafe fn data_push_tail(rb: *mut Ringbuffer, lpos: c_ulong) -> bool { // desc_reserve:F CPU (which performs the // full memory barrier) must have previously // seen data_push_tail:D. - c_prb_rmb(); // LMM(data_push_tail:B) + c_printk_ringbuffer_rmb(); // LMM(data_push_tail:B) let new_tail = load(addr_of!((*ring).tail_lpos)); // LMM(data_push_tail:C) if new_tail == tail { return false; } + + // Another CPU pushed the tail. Try again. tail = new_tail; continue; } - if cas(addr_of_mut!((*ring).tail_lpos), &mut tail, next) { - break; - } + // Guarantee any descriptor states that have transitioned to // reusable are stored before pushing the tail lpos. A full // memory barrier is needed since other CPUs may have made // the descriptor states reusable. This pairs with // data_push_tail:A. - } // LMM(data_push_tail:D) + // LMM(data_push_tail:D) + if cas(addr_of_mut!((*ring).tail_lpos), &mut tail, next) { + break; + } + } true } } @@ -1214,22 +1755,37 @@ unsafe fn data_push_tail(rb: *mut Ringbuffer, lpos: c_ulong) -> bool { /// rb and arrays remain valid throughout concurrent recycling. unsafe fn desc_push_tail(rb: *mut Ringbuffer, tail: c_ulong) -> bool { // SAFETY: (U1) all shared state transitions use the original C atomics. + // SAFETY: (U3) rb arrays remain live and descriptor snapshots are reader-local. unsafe { let ring = addr_of_mut!((*rb).desc_ring); let mut desc = empty_desc(); match desc_read(ring, tail, &mut desc, null_mut(), null_mut()) { MISS => { + // If the ID is exactly 1 wrap behind the expected, it is + // in the process of being reserved by another writer and + // must be considered reserved. + + // The ID has changed. Another writer must have pushed the + // tail and recycled the descriptor already. Success is + // returned because the caller is only interested in the + // specified tail being pushed, which it was. return (desc.state_var as c_ulong & ID_MASK) - != (tail.wrapping_sub(desc_count(ring)) & ID_MASK) + != (tail.wrapping_sub(desc_count(ring)) & ID_MASK); } RESERVED | COMMITTED => return false, FINALIZED => desc_make_reusable(ring, tail), _ => {} } + + // Data blocks must be invalidated before their associated + // descriptor can be made available for recycling. Invalidating + // them later is not possible because there is no way to trust + // data blocks once their associated descriptor is gone. if !data_push_tail(rb, desc.text_blk_lpos.next) { return false; } let next = tail.wrapping_add(1) & ID_MASK; + // Check the next descriptor after @tail_id before pushing the tail // to it because the tail must always be in a finalized or reusable // state. The implementation of prb_first_seq() relies on this. @@ -1240,6 +1796,7 @@ unsafe fn desc_push_tail(rb: *mut Ringbuffer, tail: c_ulong) -> bool { let s = desc_read(ring, next, &mut desc, null_mut(), null_mut()); // LMM(desc_push_tail:A) if s == FINALIZED || s == REUSABLE { let mut expected = tail; + // Guarantee any descriptor states that have transitioned to // reusable are stored before pushing the tail ID. This allows // verifying the recycled descriptor state. A full memory @@ -1267,7 +1824,11 @@ unsafe fn desc_push_tail(rb: *mut Ringbuffer, tail: c_ulong) -> bool { // CPUs. However, the desc_reserve:F CPU (which performs // the full memory barrier) must have previously seen // desc_push_tail:B. - c_prb_rmb(); // LMM(desc_push_tail:C) + c_printk_ringbuffer_rmb(); // LMM(desc_push_tail:C) + + // Re-check the tail ID. The descriptor following @tail_id is + // not in an allowed tail state. But if the tail has since + // been moved by another CPU, then it does not matter. if load(addr_of!((*ring).tail_id)) == tail { return false; } @@ -1280,12 +1841,14 @@ unsafe fn desc_push_tail(rb: *mut Ringbuffer, tail: c_ulong) -> bool { /// rb and arrays remain alive; id_out is private storage for the reservation. unsafe fn desc_reserve(rb: *mut Ringbuffer, id_out: &mut c_ulong) -> bool { // SAFETY: (U1) C atomics/barriers implement the reservation ownership protocol. + // SAFETY: (U3) rb arrays remain live; id_out is private writer storage. unsafe { let ring = addr_of_mut!((*rb).desc_ring); let mut head = load(addr_of!((*ring).head_id)); let (id, previous) = loop { let id = head.wrapping_add(1) & ID_MASK; let previous = id.wrapping_sub(desc_count(ring)) & ID_MASK; + // Guarantee the head ID is read before reading the tail ID. // Since the tail ID is updated before the head ID, this // guarantees that @id_prev_wrap is never ahead of the tail @@ -1306,13 +1869,14 @@ unsafe fn desc_reserve(rb: *mut Ringbuffer, id_out: &mut c_ulong) -> bool { // CPUs. However, the desc_reserve:D CPU (which performs // the full memory barrier) must have previously seen // desc_push_tail:B. - c_prb_rmb(); // LMM(desc_reserve:B) + c_printk_ringbuffer_rmb(); // LMM(desc_reserve:B) + + // Make space for the new descriptor by + // advancing the tail. if previous == load(addr_of!((*ring).tail_id)) && !desc_push_tail(rb, previous) { return false; } - if cas(addr_of_mut!((*ring).head_id), &mut head, id) { - break (id, previous); - } + // 1. Guarantee the tail ID is read before validating the // recycled descriptor state. A read memory barrier is // sufficient for this. This pairs with desc_push_tail:B. @@ -1352,157 +1916,35 @@ unsafe fn desc_reserve(rb: *mut Ringbuffer, id_out: &mut c_ulong) -> bool { // 5. Guarantee the head ID is stored before trying to // finalize the previous descriptor. This pairs with // _prb_commit:B. - }; // LMM(desc_reserve:D) + // LMM(desc_reserve:D) + if cas(addr_of_mut!((*ring).head_id), &mut head, id) { + break (id, previous); + } + }; let desc = to_desc(ring, id as u64); + // If the descriptor has been recycled, verify the old state val. // See "ABA Issues" about why this verification is performed. let mut old = load(addr_of!((*desc).state_var)); // LMM(desc_reserve:E) if old != 0 && state(previous, old) != REUSABLE { - c_prb_warn_0(true); + c_printk_ringbuffer_warn_descriptor_state(true); return false; } + + // Assign the descriptor a new ID and set its state to reserved. + // See "ABA Issues" about why cmpxchg() instead of set() is used. + // + // Guarantee the new descriptor ID and state is stored before making + // any other changes. A write memory barrier is sufficient for this. + // This pairs with desc_read:D. + // LMM(desc_reserve:F) if !cas(addr_of_mut!((*desc).state_var), &mut old, sv(id, RESERVED)) { - c_prb_warn_1(true); + c_printk_ringbuffer_warn_descriptor_reserve(true); return false; - // Assign the descriptor a new ID and set its state to reserved. - // See "ABA Issues" about why cmpxchg() instead of set() is used. - // - // Guarantee the new descriptor ID and state is stored before making - // any other changes. A write memory barrier is sufficient for this. - // This pairs with desc_read:D. - } // LMM(desc_reserve:F) + } + + // Now data in @desc can be modified: LMM(desc_reserve:G) *id_out = id; true } } -use crate::port_layout::*; - -/// Logical position and extent of a C ring data block. -#[repr(C)] -pub struct DataBlkLpos { - begin: c_ulong, - next: c_ulong, -} - -kr::static_assert_layout!(DataBlkLpos, size = PORT_PRB_DATA_BLK_LPOS_SIZE, align = PORT_PRB_DATA_BLK_LPOS_ALIGN, - begin @ PORT_PRB_DATA_BLK_LPOS_BEGIN, - next @ PORT_PRB_DATA_BLK_LPOS_NEXT, -); - -/// C descriptor containing publication state and text positions. -#[repr(C)] -pub struct Desc { - state_var: c_long, - text_blk_lpos: DataBlkLpos, -} - -kr::static_assert_layout!(Desc, size = PORT_PRB_DESC_SIZE, align = PORT_PRB_DESC_ALIGN, - state_var @ PORT_PRB_DESC_STATE_VAR, - text_blk_lpos @ PORT_PRB_DESC_TEXT_BLK_LPOS, -); - -/// C text ring geometry and atomic positions. -#[repr(C)] -pub struct DataRing { - size_bits: c_uint, - data: *mut c_char, - head_lpos: c_long, - tail_lpos: c_long, -} - -kr::static_assert_layout!(DataRing, size = PORT_PRB_DATA_RING_SIZE, align = PORT_PRB_DATA_RING_ALIGN, - size_bits @ PORT_PRB_DATA_RING_SIZE_BITS, - data @ PORT_PRB_DATA_RING_DATA, - head_lpos @ PORT_PRB_DATA_RING_HEAD_LPOS, - tail_lpos @ PORT_PRB_DATA_RING_TAIL_LPOS, -); - -/// C descriptor ring geometry and atomic sequence state. -#[repr(C)] -pub struct DescRing { - count_bits: c_uint, - descs: *mut Desc, - infos: *mut PrintkInfo, - head_id: c_long, - tail_id: c_long, - last_finalized_seq: c_long, -} - -kr::static_assert_layout!(DescRing, size = PORT_PRB_DESC_RING_SIZE, align = PORT_PRB_DESC_RING_ALIGN, - count_bits @ PORT_PRB_DESC_RING_COUNT_BITS, - descs @ PORT_PRB_DESC_RING_DESCS, - infos @ PORT_PRB_DESC_RING_INFOS, - head_id @ PORT_PRB_DESC_RING_HEAD_ID, - tail_id @ PORT_PRB_DESC_RING_TAIL_ID, - last_finalized_seq @ PORT_PRB_DESC_RING_LAST_FINALIZED_SEQ, -); - -/// C printk ring buffer shared by readers and writers. -#[repr(C)] -pub struct Ringbuffer { - desc_ring: DescRing, - text_data_ring: DataRing, - fail: c_long, -} - -kr::static_assert_layout!(Ringbuffer, size = PORT_PRINTK_RINGBUFFER_SIZE, align = PORT_PRINTK_RINGBUFFER_ALIGN, - desc_ring @ PORT_PRINTK_RINGBUFFER_DESC_RING, - text_data_ring @ PORT_PRINTK_RINGBUFFER_TEXT_DATA_RING, - fail @ PORT_PRINTK_RINGBUFFER_FAIL, -); - -/// Writer-owned reservation returned by the C ABI. -#[repr(C)] -pub struct ReservedEntry { - rb: *mut Ringbuffer, - irqflags: c_ulong, - id: c_ulong, - text_space: c_uint, -} - -kr::static_assert_layout!(ReservedEntry, size = PORT_PRB_RESERVED_ENTRY_SIZE, align = PORT_PRB_RESERVED_ENTRY_ALIGN, - rb @ PORT_PRB_RESERVED_ENTRY_RB, - irqflags @ PORT_PRB_RESERVED_ENTRY_IRQFLAGS, - id @ PORT_PRB_RESERVED_ENTRY_ID, - text_space @ PORT_PRB_RESERVED_ENTRY_TEXT_SPACE, -); - -/// Caller-provided metadata and text buffers. -#[repr(C)] -pub struct Record { - info: *mut PrintkInfo, - text_buf: *mut c_char, - text_buf_size: c_uint, -} - -kr::static_assert_layout!(Record, size = PORT_PRINTK_RECORD_SIZE, align = PORT_PRINTK_RECORD_ALIGN, - info @ PORT_PRINTK_RECORD_INFO, - text_buf @ PORT_PRINTK_RECORD_TEXT_BUF, - text_buf_size @ PORT_PRINTK_RECORD_TEXT_BUF_SIZE, -); - -// The record metadata is owned by printk's C writers. Its bitfields and -// configuration-dependent execution context stay opaque to the ring buffer. -// Scalar accessors below use the C-generated offsets. -/// C record metadata, including opaque execution-context fields. -#[repr(C)] -pub struct PrintkInfo { - seq: u64, - ts_nsec: u64, - text_len: u16, - facility: u8, - flags_level: u8, - caller_id: u32, - execution_context: [u8; PORT_PRINTK_INFO_DEV_INFO - PORT_PRINTK_INFO_CALLER_ID - 4], - dev_info: [u8; PORT_PRINTK_INFO_DEV_INFO_SIZE], -} - -kr::static_assert_layout!(PrintkInfo, - size = PORT_PRINTK_INFO_SIZE, align = PORT_PRINTK_INFO_ALIGN, - seq @ PORT_PRINTK_INFO_SEQ, ts_nsec @ PORT_PRINTK_INFO_TS_NSEC, - text_len @ PORT_PRINTK_INFO_TEXT_LEN, facility @ PORT_PRINTK_INFO_FACILITY, - flags_level @ (PORT_PRINTK_INFO_FACILITY + 1), - caller_id @ PORT_PRINTK_INFO_CALLER_ID, - execution_context @ (PORT_PRINTK_INFO_CALLER_ID + 4), - dev_info @ PORT_PRINTK_INFO_DEV_INFO, -); diff --git a/kernel/printk/printk_ringbuffer_ffi.c b/kernel/printk/printk_ringbuffer_ffi.c index 92e5808bea8278..f8e0f9184f47b5 100644 --- a/kernel/printk/printk_ringbuffer_ffi.c +++ b/kernel/printk/printk_ringbuffer_ffi.c @@ -13,170 +13,188 @@ EXPORT_SYMBOL_IF_KUNIT(prb_commit); EXPORT_SYMBOL_IF_KUNIT(prb_read_valid); EXPORT_SYMBOL_IF_KUNIT(prb_init); -long c_prb_load(const atomic_long_t *p); -long c_prb_load(const atomic_long_t *p) +long c_printk_ringbuffer_load(const atomic_long_t *p); +long c_printk_ringbuffer_load(const atomic_long_t *p) { return atomic_long_read(p); } -long c_prb_load_acquire(const atomic_long_t *p); -long c_prb_load_acquire(const atomic_long_t *p) +long c_printk_ringbuffer_load_acquire(const atomic_long_t *p); +long c_printk_ringbuffer_load_acquire(const atomic_long_t *p) { return atomic_long_read_acquire(p); } -void c_prb_store(atomic_long_t *p, long value); -void c_prb_store(atomic_long_t *p, long value) +void c_printk_ringbuffer_store(atomic_long_t *p, long value); +void c_printk_ringbuffer_store(atomic_long_t *p, long value) { atomic_long_set(p, value); } -bool c_prb_cas(atomic_long_t *p, long *old, long value); -bool c_prb_cas(atomic_long_t *p, long *old, long value) +bool c_printk_ringbuffer_cas(atomic_long_t *p, long *old, long value); +bool c_printk_ringbuffer_cas(atomic_long_t *p, long *old, long value) { return atomic_long_try_cmpxchg(p, old, value); } -bool c_prb_cas_relaxed(atomic_long_t *p, long *old, long value); -bool c_prb_cas_relaxed(atomic_long_t *p, long *old, long value) +bool c_printk_ringbuffer_cas_relaxed(atomic_long_t *p, long *old, long value); +bool c_printk_ringbuffer_cas_relaxed(atomic_long_t *p, long *old, long value) { return atomic_long_try_cmpxchg_relaxed(p, old, value); } -bool c_prb_cas_release(atomic_long_t *p, long *old, long value); -bool c_prb_cas_release(atomic_long_t *p, long *old, long value) +bool c_printk_ringbuffer_cas_release(atomic_long_t *p, long *old, long value); +bool c_printk_ringbuffer_cas_release(atomic_long_t *p, long *old, long value) { return atomic_long_try_cmpxchg_release(p, old, value); } -void c_prb_inc(atomic_long_t *p); -void c_prb_inc(atomic_long_t *p) +void c_printk_ringbuffer_inc(atomic_long_t *p); +void c_printk_ringbuffer_inc(atomic_long_t *p) { atomic_long_inc(p); } -void c_prb_rmb(void); -void c_prb_rmb(void) +void c_printk_ringbuffer_rmb(void); +void c_printk_ringbuffer_rmb(void) { smp_rmb(); } -void c_prb_irq_save(unsigned long *flags); -void c_prb_irq_save(unsigned long *flags) +void c_printk_ringbuffer_irq_save(unsigned long *flags); +void c_printk_ringbuffer_irq_save(unsigned long *flags) { local_irq_save(*flags); } -void c_prb_irq_restore(unsigned long flags); -void c_prb_irq_restore(unsigned long flags) +void c_printk_ringbuffer_irq_restore(unsigned long flags); +void c_printk_ringbuffer_irq_restore(unsigned long flags) { local_irq_restore(flags); } -void * c_prb_copy(void *dst, const void *src, size_t n); -void * c_prb_copy(void *dst, const void *src, size_t n) +void *c_printk_ringbuffer_copy(void *dst, const void *src, size_t n); +void *c_printk_ringbuffer_copy(void *dst, const void *src, size_t n) { return memcpy(dst, src, n); } -void * c_prb_clear(void *dst, int value, size_t n); -void * c_prb_clear(void *dst, int value, size_t n) +void *c_printk_ringbuffer_clear(void *dst, int value, size_t n); +void *c_printk_ringbuffer_clear(void *dst, int value, size_t n) { return memset(dst, value, n); } -void * c_prb_find(const void *src, int value, size_t n); -void * c_prb_find(const void *src, int value, size_t n) +u8 c_printk_ringbuffer_read_u8(const u8 *src); +u8 c_printk_ringbuffer_read_u8(const u8 *src) { - return memchr(src, value, n); + return READ_ONCE(*src); } -bool c_prb_panic_cpu(void); -bool c_prb_panic_cpu(void) +u16 c_printk_ringbuffer_read_u16(const u16 *src); +u16 c_printk_ringbuffer_read_u16(const u16 *src) +{ + return READ_ONCE(*src); +} + +u32 c_printk_ringbuffer_read_u32(const u32 *src); +u32 c_printk_ringbuffer_read_u32(const u32 *src) +{ + return READ_ONCE(*src); +} + +u64 c_printk_ringbuffer_read_u64(const u64 *src); +u64 c_printk_ringbuffer_read_u64(const u64 *src) +{ + return READ_ONCE(*src); +} + +bool c_printk_ringbuffer_panic_cpu(void); +bool c_printk_ringbuffer_panic_cpu(void) { return panic_on_this_cpu(); } -bool c_prb_warn_0(bool condition); -bool c_prb_warn_0(bool condition) +bool c_printk_ringbuffer_warn_descriptor_state(bool condition); +bool c_printk_ringbuffer_warn_descriptor_state(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_1(bool condition); -bool c_prb_warn_1(bool condition) +bool c_printk_ringbuffer_warn_descriptor_reserve(bool condition); +bool c_printk_ringbuffer_warn_descriptor_reserve(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_2(bool condition); -bool c_prb_warn_2(bool condition) +bool c_printk_ringbuffer_warn_allocation_size(bool condition); +bool c_printk_ringbuffer_warn_allocation_size(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_3(bool condition); -bool c_prb_warn_3(bool condition) +bool c_printk_ringbuffer_warn_reallocation_size(bool condition); +bool c_printk_ringbuffer_warn_reallocation_size(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_4(bool condition); -bool c_prb_warn_4(bool condition) +bool c_printk_ringbuffer_warn_block_wrap(bool condition); +bool c_printk_ringbuffer_warn_block_wrap(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_5(bool condition); -bool c_prb_warn_5(bool condition) +bool c_printk_ringbuffer_warn_block_begin_alignment(bool condition); +bool c_printk_ringbuffer_warn_block_begin_alignment(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_6(bool condition); -bool c_prb_warn_6(bool condition) +bool c_printk_ringbuffer_warn_block_next_alignment(bool condition); +bool c_printk_ringbuffer_warn_block_next_alignment(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_7(bool condition); -bool c_prb_warn_7(bool condition) +bool c_printk_ringbuffer_warn_block_minimum_size(bool condition); +bool c_printk_ringbuffer_warn_block_minimum_size(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_8(bool condition); -bool c_prb_warn_8(bool condition) +bool c_printk_ringbuffer_warn_block_maximum_size(bool condition); +bool c_printk_ringbuffer_warn_block_maximum_size(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_9(bool condition); -bool c_prb_warn_9(bool condition) +bool c_printk_ringbuffer_warn_empty_text_length(bool condition); +bool c_printk_ringbuffer_warn_empty_text_length(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_10(bool condition); -bool c_prb_warn_10(bool condition) +bool c_printk_ringbuffer_warn_text_length(bool condition); +bool c_printk_ringbuffer_warn_text_length(bool condition) { return WARN_ON_ONCE(condition); } -bool c_prb_warn_11(bool condition); -bool c_prb_warn_11(bool condition) +bool c_printk_ringbuffer_warn_commit(bool condition); +bool c_printk_ringbuffer_warn_commit(bool condition) { return WARN_ON_ONCE(condition); } -void c_prb_warn_len_zero(unsigned short length); -void c_prb_warn_len_zero(unsigned short length) +void c_printk_ringbuffer_warn_len_zero(unsigned short length); +void c_printk_ringbuffer_warn_len_zero(unsigned short length) { pr_warn_once("wrong text_len value (%hu, expecting 0)\n", length); } -void c_prb_warn_len_max(unsigned short length, unsigned int max); -void c_prb_warn_len_max(unsigned short length, unsigned int max) +void c_printk_ringbuffer_warn_len_max(unsigned short length, unsigned int max); +void c_printk_ringbuffer_warn_len_max(unsigned short length, unsigned int max) { pr_warn_once("wrong text_len value (%hu, expecting <=%u)\n", length, max); } From 70230b9e99b4be3ffcc9f14ddb2e485592eae680 Mon Sep 17 00:00:00 2001 From: Bruno Herrera Date: Mon, 21 Sep 2026 14:48:44 -0300 Subject: [PATCH 6/6] printk: generate metadata field offsets Emit the packed flags and level byte offset plus execution-context offsets from the active C layout. Use the generated values in the Rust mirror instead of deriving offsets from neighboring fields. Test: Rust main crate builds with execution-context metadata enabled. Assisted-by: LLM [Codex] --- kernel/port-layout.c | 10 ++++++++++ kernel/printk/printk_ringbuffer.rs | 15 ++++++++++++--- rust/Makefile | 1 + 3 files changed, 23 insertions(+), 3 deletions(-) diff --git a/kernel/port-layout.c b/kernel/port-layout.c index 15eb83c2ef4e1e..bb5cf3d8086cee 100644 --- a/kernel/port-layout.c +++ b/kernel/port-layout.c @@ -66,7 +66,17 @@ int main(void) OFFSET(PORT_PRINTK_INFO_TS_NSEC, printk_info, ts_nsec); OFFSET(PORT_PRINTK_INFO_TEXT_LEN, printk_info, text_len); OFFSET(PORT_PRINTK_INFO_FACILITY, printk_info, facility); + DEFINE(PORT_PRINTK_INFO_FLAGS_LEVEL, offsetof(struct printk_info, facility) + 1); OFFSET(PORT_PRINTK_INFO_CALLER_ID, printk_info, caller_id); +#ifdef CONFIG_PRINTK_EXECUTION_CTX + OFFSET(PORT_PRINTK_INFO_CALLER_ID2, printk_info, caller_id2); + OFFSET(PORT_PRINTK_INFO_COMM, printk_info, comm); + DEFINE(PORT_PRINTK_INFO_COMM_SIZE, sizeof(((struct printk_info *)0)->comm)); +#else + DEFINE(PORT_PRINTK_INFO_CALLER_ID2, 0); + DEFINE(PORT_PRINTK_INFO_COMM, 0); + DEFINE(PORT_PRINTK_INFO_COMM_SIZE, 0); +#endif OFFSET(PORT_PRINTK_INFO_DEV_INFO, printk_info, dev_info); DEFINE(PORT_PRINTK_INFO_DEV_INFO_SIZE, sizeof(struct dev_printk_info)); #endif diff --git a/kernel/printk/printk_ringbuffer.rs b/kernel/printk/printk_ringbuffer.rs index ff47c92312c07d..dbb6f31a0c233f 100644 --- a/kernel/printk/printk_ringbuffer.rs +++ b/kernel/printk/printk_ringbuffer.rs @@ -181,7 +181,10 @@ pub struct PrintkInfo { facility: u8, flags_level: u8, caller_id: u32, - execution_context: [u8; PORT_PRINTK_INFO_DEV_INFO - PORT_PRINTK_INFO_CALLER_ID - 4], + #[cfg(CONFIG_PRINTK_EXECUTION_CTX)] + caller_id2: u32, + #[cfg(CONFIG_PRINTK_EXECUTION_CTX)] + comm: [u8; PORT_PRINTK_INFO_COMM_SIZE], dev_info: [u8; PORT_PRINTK_INFO_DEV_INFO_SIZE], } @@ -189,12 +192,18 @@ kr::static_assert_layout!(PrintkInfo, size = PORT_PRINTK_INFO_SIZE, align = PORT_PRINTK_INFO_ALIGN, seq @ PORT_PRINTK_INFO_SEQ, ts_nsec @ PORT_PRINTK_INFO_TS_NSEC, text_len @ PORT_PRINTK_INFO_TEXT_LEN, facility @ PORT_PRINTK_INFO_FACILITY, - flags_level @ (PORT_PRINTK_INFO_FACILITY + 1), + flags_level @ PORT_PRINTK_INFO_FLAGS_LEVEL, caller_id @ PORT_PRINTK_INFO_CALLER_ID, - execution_context @ (PORT_PRINTK_INFO_CALLER_ID + 4), dev_info @ PORT_PRINTK_INFO_DEV_INFO, ); +#[cfg(CONFIG_PRINTK_EXECUTION_CTX)] +kr::static_assert_layout!(PrintkInfo, + size = PORT_PRINTK_INFO_SIZE, align = PORT_PRINTK_INFO_ALIGN, + caller_id2 @ PORT_PRINTK_INFO_CALLER_ID2, + comm @ PORT_PRINTK_INFO_COMM, +); + /// # Safety /// ring's arrays are valid; desc is private writable storage. unsafe fn finalized_seq(ring: *mut DescRing, id: c_ulong, seq: u64, desc: &mut Desc) -> c_int { diff --git a/rust/Makefile b/rust/Makefile index c7e6776b0c1d1e..20fa03e39fa129 100644 --- a/rust/Makefile +++ b/rust/Makefile @@ -789,6 +789,7 @@ $(obj)/main.o: private rustc_target_flags = --edition=2024 --extern kr \ @$(objtree)/include/generated/port-layout.cfg \ '--check-cfg=cfg(PORT_LOCKREF_FAST)' \ '--check-cfg=cfg(CONFIG_PRINTK)' \ + '--check-cfg=cfg(CONFIG_PRINTK_EXECUTION_CTX)' \ '--check-cfg=cfg(CONFIG_ARCH_DMA_ADDR_T_64BIT)' \ '--check-cfg=cfg(CONFIG_PROVE_RCU)' $(obj)/main.o: $(srctree)/main.rs $(obj)/kr.o \