Crawls client-rendered Mintlify-on-Next.js documentation sites that ordinary
crawlers can't read, and dumps every page to clean Markdown. Built for
https://docs.onekhusa.com but works on any site with the same architecture.
The site is a Next.js App Router app. A plain HTTP GET returns an empty
streaming shell — the <body> is just <!--$?--> placeholders filled in by
JavaScript — and the usual escape hatches are all gone:
/llms.txt,/sitemap.xml,/robots.txt,page.md→ all 404<meta name="robots" content="noindex">
So a crawler that doesn't execute JS sees nothing.
The full content ships inside each page's React Server Component (RSC)
payload, which you can fetch over plain HTTP by sending an RSC: 1 header.
The payload is a stream of rows <<hexid>>:T<<hexlen>>,<<bytes>>. crwlr:
- Fetches each page with
RSC: 1(text/x-component). - Splits the payload into rows — slicing by byte length (the
T<hexlen>is bytes, not UTF-16 units; char-slicing corrupts multibyte rows). - Prose lives in a compiled MDX module row (a JS function that builds
React elements via
jsx/jsxs). crwlr runs that module in a locked-downvmsandbox withjsx/jsxsstubs that emit Markdown directly — generic across every tag and Mintlify component (Info,Warning, tables, …). - API reference pages render from a flattened OpenAPI spec embedded as a
uuid -> objectmap plus operation objects with UUID refs. crwlr rebuilds the map, resolves the refs, and renders method/path, description, parameters, request body, responses, and the curl sample. (Seeopenapi.js.) - Crawls breadth-first, discovering links from each payload, restricted to
real doc roots (
/api-reference,/guides,/sdk,/changelog,/disclaimer) — the/security/...,/core/...etc. paths in the payload are OpenAPI backend paths from code samples, not pages.
Requires Node 18+ (uses global fetch). No dependencies.
node crawl.js # crawl docs.onekhusa.com -> ./output
node crawl.js --base https://docs.onekhusa.com --out output --delay 200
node crawl.js --roots /api-reference,/guides # limit to specific doc rootsFlags: --base (site root), --out (output dir), --delay (ms between
requests, default 200), --roots (comma-separated allowlisted path prefixes).
output/<path>.md— one Markdown file per page, mirroring the URL structure, withsource:/path:frontmatter.output/index.md— crawl index: every page with size, plus anything discovered but empty.
Last run: 150 pages, 0 empty.
crawl.js— fetch, RSC row parsing, MDX→Markdown renderer, BFS, output.openapi.js— OpenAPI operation reconstruction from the embedded ref map.
- Deep request-body schemas are sometimes lazy-loaded (their UUID isn't in the page payload); in those cases the curl sample still shows the body shape.
- If the site changes its RSC/MDX serialization, the renderer may need tweaks.
- Be polite: keep
--delayreasonable and respect the site's terms of use.