Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

crwlr — RSC-aware docs crawler

Crawls client-rendered Mintlify-on-Next.js documentation sites that ordinary crawlers can't read, and dumps every page to clean Markdown. Built for https://docs.onekhusa.com but works on any site with the same architecture.

Why ordinary crawlers fail here

The site is a Next.js App Router app. A plain HTTP GET returns an empty streaming shell — the <body> is just <!--$?--> placeholders filled in by JavaScript — and the usual escape hatches are all gone:

  • /llms.txt, /sitemap.xml, /robots.txt, page.md → all 404
  • <meta name="robots" content="noindex">

So a crawler that doesn't execute JS sees nothing.

How crwlr gets the content (no headless browser)

The full content ships inside each page's React Server Component (RSC) payload, which you can fetch over plain HTTP by sending an RSC: 1 header. The payload is a stream of rows <<hexid>>:T<<hexlen>>,<<bytes>>. crwlr:

  1. Fetches each page with RSC: 1 (text/x-component).
  2. Splits the payload into rows — slicing by byte length (the T<hexlen> is bytes, not UTF-16 units; char-slicing corrupts multibyte rows).
  3. Prose lives in a compiled MDX module row (a JS function that builds React elements via jsx/jsxs). crwlr runs that module in a locked-down vm sandbox with jsx/jsxs stubs that emit Markdown directly — generic across every tag and Mintlify component (Info, Warning, tables, …).
  4. API reference pages render from a flattened OpenAPI spec embedded as a uuid -> object map plus operation objects with UUID refs. crwlr rebuilds the map, resolves the refs, and renders method/path, description, parameters, request body, responses, and the curl sample. (See openapi.js.)
  5. Crawls breadth-first, discovering links from each payload, restricted to real doc roots (/api-reference, /guides, /sdk, /changelog, /disclaimer) — the /security/..., /core/... etc. paths in the payload are OpenAPI backend paths from code samples, not pages.

Usage

Requires Node 18+ (uses global fetch). No dependencies.

node crawl.js                                    # crawl docs.onekhusa.com -> ./output
node crawl.js --base https://docs.onekhusa.com --out output --delay 200
node crawl.js --roots /api-reference,/guides     # limit to specific doc roots

Flags: --base (site root), --out (output dir), --delay (ms between requests, default 200), --roots (comma-separated allowlisted path prefixes).

Output

  • output/<path>.md — one Markdown file per page, mirroring the URL structure, with source:/path: frontmatter.
  • output/index.md — crawl index: every page with size, plus anything discovered but empty.

Last run: 150 pages, 0 empty.

Files

  • crawl.js — fetch, RSC row parsing, MDX→Markdown renderer, BFS, output.
  • openapi.js — OpenAPI operation reconstruction from the embedded ref map.

Notes / limits

  • Deep request-body schemas are sometimes lazy-loaded (their UUID isn't in the page payload); in those cases the curl sample still shows the body shape.
  • If the site changes its RSC/MDX serialization, the renderer may need tweaks.
  • Be polite: keep --delay reasonable and respect the site's terms of use.

About

RSC-aware crawler for client-rendered Mintlify/Next.js docs sites (no headless browser)

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages