Skip to content

feat(preview): render XLSX spreadsheets in the file-preview overlay - #502

Open
aakhter wants to merge 2 commits into
Ark0N:masterfrom
aakhter:pr/xlsx-preview
Open

aakhter wants to merge 2 commits into
Ark0N:masterfrom
aakhter:pr/xlsx-preview

Conversation

@aakhter

@aakhter aakhter commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

What

.xlsx files were download-only. This adds a read-only preview in the file-preview overlay: sheet tabs, number formats, merged cells and theme colours, virtualized so large sheets scroll smoothly. It works for workspace files, attachments, and xlsx paths printed in the terminal (added to the file-path link pattern and FILE_PREVIEW_EXTENSIONS). .xls and .ods stay download-only, since ExcelJS only reads xlsx; tests pin both.

How

  • All parsing happens in the browser, in a Web Worker (spreadsheet-preview-worker.js + spreadsheet-xlsx-core.js) using exceljs@4.4.0 and fflate@0.8.2 (both MIT, pinned exactly as devDependencies and copied into vendor/ by postinstall and the build, like the other vendored bundles). The server does no parsing.
  • Server side xlsx only joins the existing attachment allowlist and file-content classification, so every request still goes through the existing confinement (resolveFileTarget, resolveServableAttachmentPath). There is no new path handling. ?preview=true is capped at 10 MB (413 above it) on both file-raw and the attachment raw route; downloads are unchanged.
  • Overlay wiring is two small methods on the existing preview (_openSpreadsheetPreview / _disposeSpreadsheetPreview), torn down from _stopFilePreviewMedia on open and close. The worker is created via CodemanBase.url and loads its scripts by relative URL, so --base-url mounts work.

Safety

The workbook is untrusted input:

  • Checked before ExcelJS loads (admitXlsx): at most 5000 ZIP entries, 64 MB inflated in total, 32 MB per entry, a 100:1 compression ratio, 50 worksheets, 250k cells (100k per sheet), 5000 merges per sheet and 5000 styles. Encrypted and ZIP64 files are refused. There is also a 20 s timeout.
  • Rendering: cell text and sheet names are written with textContent. The generated style block only accepts validated #rrggbb colours and a fixed keyword set. At most 2500 cells are drawn per tile.
  • Nothing is evaluated or fetched: formulas show their cached result (or the formula text), and external links, charts, drawings and macros only produce a "not shown" notice.

Cost

Page load gains only spreadsheet-preview.js (5.0 KB gz). Opening a spreadsheet then loads the worker (3.0 KB), core (7.4 KB), fflate (12.5 KB) and exceljs (256 KB gz, 948 KB raw), about 284 KB gz in total. ExcelJS is only fetched after the workbook passes the checks. check:public-assets enforces a 1.1 MB vendor budget, and a content-hash SPREADSHEET_ASSET_VERSION cache-busts the worker.

dependency-security.test.ts gains one exact exemption: exceljs pins uuid@8.3.2 (our own uuid stays >= 14). The advisory is MODERATE, outside that suite's CRITICAL/HIGH policy, and covers v3/v5/v6 with a caller-supplied buffer, while exceljs only calls v4(). The browser also loads exceljs's own prebuilt dist bundle, and nothing server-side imports exceljs.

Testing

  • 43 new tests: xlsx core 20, worker 7, renderer/overlay 10, assets 5, plus 1 Chromium test under a strict CSP (added to BROWSER_TEST_GLOBS). Also 6 new route tests across file-routes and the attachment path guard.
  • Mutation-checked. Each of these fails a test:
    • textContent swapped for innerHTML
    • xlsx removed from the allowlist
    • the classification change reverted
    • the preview cap removed
    • the compression-ratio cap removed
    • the workbook checks skipped
    • no dispose on close
    • the formula-text path broken
  • Related suites pass unchanged, including file-routes, the attachment path guard, sw-precache, base-path and dependency-security.
  • typecheck, lint, check:frontend-syntax, format:check, check:public-assets, check:lockfile and build are clean.
  • Checked in a real browser on an isolated instance: both sheet tabs rendered, a cell containing <img onerror> displayed as text with no image element created, no console errors, and none of the four preview requests were made until a spreadsheet was opened.

xlsx files were download-only. Add a read-only, virtualized preview (sheet
tabs, number formats, merges, theme colours) parsed entirely in a browser
Web Worker with exceljs and fflate, loaded only when a spreadsheet is
opened. The workbook is checked against ZIP-bomb, entry and cell limits
before exceljs loads; cell text is written with textContent, formulas are
never evaluated and nothing referenced by the workbook is fetched. On the
server xlsx only joins the existing allowlist and classification, with a
10 MB cap on ?preview=true. xls and ods stay download-only.
@Ark0N

Ark0N commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Thanks a lot for this, @aakhter. It adds a read-only XLSX preview to the file-preview overlay, parsed entirely in a browser worker, and the overall design is exactly right: no server-side parsing, no new path handling, lazy worker-only vendor bundles, textContent rendering, an allowlisted style block, --base-url support and a content-hashed cache-bust token. Typecheck, lint, format, the asset checks, the full npm test gate and your Chromium test all pass here.

While testing it with real ExcelJS workbooks through the worker I hit three bugs that need fixing before merge:

  1. Date, rich-text, hyperlink and error cells display wrong (src/web/public/spreadsheet-xlsx-core.js:309). ExcelJS turns date-formatted cells into JS Date objects on load, so String(value) shows Mon Jan 15 2024 01:00:00 GMT+0100 (...), and in TZ=America/New_York the same cell reads Sun Jan 14 2024 ..., a day early. Rich text ({richText}), hyperlinks ({text, hyperlink}), error values ({error}) and formulas with an error result all render [object Object]. Please normalize these shapes before formatting (format Date from its UTC components per the numFmt, join richText runs, use .text / .error, recurse on result for formula and sharedFormula) and add worker round-trip tests for each, one of them under a negative-offset TZ.

  2. Filtered or dense sheets fail to preview (src/web/public/spreadsheet-preview-worker.js:179-188). sendTile() includes hidden rows, which are 0 px tall, so the viewport spans all of them, and above 2500 cells it throws tile-limit, which replaces the whole grid with an error. A 1000 x 6 sheet with 980 rows hidden by a filter fails at the renderer's default viewport, and so does a 60 x 60 filled block. Please skip hidden rows and columns in sendTile() and return a truncated tile with a warning at the cap instead of throwing (the renderer already slices to 2500). A hidden-rows worker test would pin it.

  3. admitXlsx() can be bypassed with overlapping ZIP entries (src/web/public/spreadsheet-xlsx-core.js:153-207). Admission follows local headers in file order, while ExcelJS (JSZip) follows the central directory, and the name-count check does not tie the two together. A crafted file with a stored entry that hides a full sheet1.xml, plus a small decoy sheet1.xml later in the stream, was admitted as 1 cell and 611 KB inflated; the worker then parsed 300,000 cells. Impact stays in the viewer's tab, but the PR and the new CLAUDE.md line state that admission caps the ZIP before ExcelJS runs. The simplest robust fix is to hand ExcelJS a STORE-only archive rebuilt from the entries admission already inflated (fflate.zipSync(entries, { level: 0 })); rejecting central entries whose local extents overlap also works. Please add that fixture as a regression test.

Two smaller things that fit in the same round:

  • Row and column headings take their size from CSS (64 x 20 px, styles.css:19176) rather than the axis math (spreadsheet-preview.js:216-234), so they misalign with custom widths and heights, including column B and row 2 of your own fixture. Setting width/height from the same axisOffset differences as the cells fixes it.
  • The new XLSX rule in CLAUDE.md points at docs/architecture-invariants.md#file-path-links-terminal--response-viewer, which was not updated; a short paragraph there (admission caps, worker-only vendor loading, SPREADSHEET_ASSET_VERSION) keeps the two in step. The attachments panel help (panels-ui.js:4793) and codeman attach error text (src/cli.ts:114) could also list .xlsx now.

Everything else (the route changes, the allowlist addition, the packaging and the tests) is in good shape, so once these land it is ready to merge. Thanks again for the careful work on this.

- Normalize the value shapes ExcelJS loads before formatting: Date cells are
  formatted from their serial (UTC), so they no longer render as a local-time
  string a day early at negative UTC offsets; rich text joins its runs,
  hyperlinks show their text, error values show the error, and formula and
  shared-formula results (including error results) recurse. Excel serials are
  rounded to whole milliseconds so 00:05 no longer shows as 00:04.
- sendTile() skips hidden rows and columns, and at the 2500-cell cap returns a
  truncated tile with a warning instead of failing the whole preview.
- ExcelJS now parses a STORE-only archive rebuilt from exactly the entries
  admitXlsx() inflated and counted, never the fetched bytes. Admission walks
  local headers while JSZip reads the central directory, so overlapping
  entries could show the two readers different sheets. A duplicate local
  entry name is refused. The theme fallback reads the admitted entry too.
- Row and column headings take their size from the same axis math as cells.
- Document the admission, worker-only loading and SPREADSHEET_ASSET_VERSION
  rules in architecture-invariants, and list .xlsx in the attachments panel
  help and the `codeman attach` error text (built from the accepted list).
@aakhter

aakhter commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for the careful review, and for testing with real ExcelJS workbooks. All of it is addressed in 4edb7b8.

  1. Cell values. Dates now format from their Excel serial (UTC), so the day-early bug is gone. Rich text joins its runs, hyperlinks show their text, errors show the error string, and formula/sharedFormula recurse on result, including error results. The worker round-trip tests build their workbooks with ExcelJS and run under TZ=America/New_York (the file asserts the zone took effect). They also caught a second bug: excelDate truncated fractional milliseconds, so 00:05 showed as 00:04. It rounds now.
  2. Hidden rows and dense sheets. sendTile() skips hidden rows and columns, and at the cap it returns a truncated tile with a warning instead of throwing. Your 1000x6 sheet with 980 rows filtered out and the 60x60 block are both tests, and both failed with tile-limit on the old code.
  3. Admission bypass. I went with your preferred fix: ExcelJS only ever gets buildAdmittedArchive(), a STORE-only fflate.zipSync(entries, { level: 0 }) of the entries admission inflated, and a name streamed twice is refused. The regression fixture is your construction: a stored carrier entry hiding a full 11,000x10 sheet1.xml, a one-cell decoy later in the stream, and a central directory pointing inside the carrier. On the old code admission counted 1 cell while the worker parsed 110k. Handing ExcelJS the original bytes again fails that test. The cost is one extra copy of the inflated entries, still bounded by the 64 MB cap.

The small ones are done too: headings are sized from the same axisOffset math as the cells (the Chromium test checks column B and row 2 against their cells), docs/architecture-invariants.md has a paragraph under the anchor CLAUDE.md points to, and .xlsx is listed in the attachments panel help and the codeman attach error text. Both of those are now built from one DOCUMENT_ATTACHMENT_EXTENSIONS list, with a test that catches drift.

@Ark0N

Ark0N commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Thanks again, @aakhter. This PR adds a read-only XLSX preview to the file-preview overlay, parsed in a browser worker behind admission caps. I checked every item from the first round at 4edb7b8 and all of them are fixed: the date and value shapes (with the TZ test), hidden rows and the truncated tile, the STORE-only rebuilt archive with its overlapping-entry fixture, heading sizes, and the docs and help text. Typecheck, lint, format, the asset checks, the full npm test gate (436 files, 8355 tests) and your Chromium test all pass here.

One more thing needs fixing before merge, in the same family as the round 1 admission bypass:

  1. Admission does not bound what ExcelJS allocates (src/web/public/spreadsheet-xlsx-core.js:133, parsed at src/web/public/spreadsheet-preview-worker.js:128). createXmlCounter counts <c and <mergeCell occurrences, but ExcelJS 4.4.0 expands three constructs into one object per cell or column at load time:

    • every cell inside a merge range (Worksheet._mergeCellsInternal),
    • every address in a <dataValidation sqref> (DataValidationsXform.parseClose),
    • every index up to <col max> (Column.fromModel, with no clamp to 16384).

    I ran your worker and core through the worker-test harness. A 6.5 KB file that admission counts as 1 cell produced:

    • <mergeCell ref="A1:CV30000"/>: 10.4 s and 1.05 GB of heap,
    • <dataValidation sqref="A1:XFD1048576">: still running after 60 s,
    • <col min="1" max="3000000"/>: 511 MB.

    It also hits ordinary files: a dropdown validation on five whole columns (B2:F1048576) takes 6.9 s and 448 MB. Please:

    • load with xlsx.load(admitted, { ignoreNodes: ['dataValidations'] }). The preview never shows validations, and this takes the full-sheet case from over 60 s to 32 ms here.
    • In createXmlCounter, add each <mergeCell ref> area (rows x cols, via parseRange) to the per-sheet and total cell counts, and refuse a ref that does not parse.
    • Refuse a <col> whose min or max exceeds 16384.
    • Add a worker-harness fixture for each of the three.

Two smaller ones that fit in the same round:

  • src/web/public/spreadsheet-preview.js:85: axisOffset walks the whole override list on every call, and renderTile calls it about six times per cell. On a 100k-row sheet with an explicit height on every row, one tile near the bottom costs about 960 ms on the main thread, and that repeats on every scroll frame. A prefix-sum array per axis, built in selectSheet() and binary-searched, makes each call O(log n).
  • src/web/public/spreadsheet-xlsx-core.js:194: the ratio check divides by the central directory's declared compressedSize, which nothing validates. A file that sets it to 0x7fffffff passes the ratio check (the 32/64 MB caps still hold). Refusing a declared size larger than the file makes the ratio cap real.

At merge time I will move the SPREADSHEET_ASSET_VERSION hash comparison into test/spreadsheet-assets.test.ts, because CI does not run check:public-assets and the vitest test only checks the token's shape. The late-response race in openFilePreview, where a slow file-content reply can start a preview after the overlay has closed, predates this PR and affects every preview type, so I will fix it separately.

Once the admission fix lands, this is ready to merge. Thanks for the careful work on both rounds.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants