Skip to content
JohanLindvallPublic

About

Bulk downloader in one static binary: resolves albums, folders and playlists from 180+ sites (gofile, bunkr, MEGA, pixeldrain…) and downloads them in parallel, with resumable multi-connection transfers, a web UI, a headless CLI and a Docker image.

Topics

Resources

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

HeapLeach

CI Release License: MIT

HeapLeach is a bulk downloader in one static binary. Paste links from 180+ sites — file hosts such as gofile, bunkr, MEGA and pixeldrain, image galleries, video and broadcaster pages — and it resolves every album, folder and playlist behind them, then downloads the files in parallel: resumable multi-connection transfers, a per-host throttle for overloaded servers, and native decryption of MEGA files and AES-128 HLS streams. The web UI is compiled into the binary, the same binary downloads headless in a terminal, and every release ships builds for Linux, macOS and Windows and a Docker image for amd64 and arm64.

./heapleach                                          # the web UI on a free local port, in your browser
./heapleach https://gofile.io/d/<code> ~/Downloads   # or headless, with progress in the terminal
docker run -p 8080:8080 -v ~/Downloads:/downloads --user "$(id -u):$(id -g)" ghcr.io/johanlindvall/heapleach

The queue in dark mode: header totals with a throughput graph, sidebar filters, and per-file progress with stream counts and ETAs.

  • Backend — Go. Worker pool, per-host extractors, resumable transfers.
  • Frontend — TypeScript + React, compiled and embedded into the Go binary. One file to ship, nothing to serve from disk.
  • Build — runs entirely in Docker; the binary is exported to your host.
  • Two ways to run — serve the UI, or hand it URLs on the command line and watch them download in the terminal.

Screenshots

Light and dark follow the system by default; the header toggle overrides it, and the choice is remembered. The layout collapses to a single column on a phone, with the filters becoming a scrolling strip.

The same queue in light mode. The queue at phone width, single column with a scrolling filter strip.

Given URLs instead of a directory, it skips the browser entirely and animates the same progress on the terminal:

Terminal output: each source announced as it resolves, finished files logged above, and a live block of progress bars with percentage, sizes, rate, stream count and ETA.

Quick start

Grab a build from releases — Linux (x86-64, arm64), macOS (Intel, Apple silicon) and Windows — unpack it and run it. It is one static binary with the UI inside; there is nothing to install alongside.

tar -xzf heapleach_*_linux_amd64.tar.gz
./heapleach

That is the whole thing. Run on its own it takes a free port on your own machine, opens your browser at it, and saves to your Downloads folder — nothing to choose and nothing to collide with, since the kernel picks the port and the process is the only thing that could know which one it got.

It also stops on its own: once nothing is downloading and no browser has been in touch for a minute, it exits. Closing the tab is how you quit, and you never accumulate forgotten copies of it in the background. An open tab counts as being in touch, so leaving one up keeps it running.

Give it any argument and it stops guessing: the settings under Usage below apply as written, it binds :8080, and it runs until you stop it — because that is a way of deploying the program, and outliving the browser may be the whole point. HEAPLEACH_OPEN=0 suppresses the browser if you want the rest of the bare behaviour without it.

Or build it yourself:

make build                  # builds in Docker, writes ./bin/heapleach
./bin/heapleach             # same as above
./bin/heapleach ~/Videos    # explicit: :8080, no browser, saves there

No local Go or Node needed — the toolchain lives in the build image, and only the finished binary lands on your machine.

Prefer to run it as a container — every release is published to the GitHub container registry for amd64 and arm64, tagged with its version (v1.2.3 and 1.2.3 alike), its minor version (v1.2, 1.2) and latest. Pin a version to stay on it; latest follows each release:

docker run -d -p 8080:8080 -v ~/Downloads:/downloads --user "$(id -u):$(id -g)" \
  ghcr.io/johanlindvall/heapleach:latest
make run-image                 # or build the image from this checkout and run it

The container maps a fixed port rather than taking a free one, since a port only the container knows about is one nothing outside it can reach. It saves to ~/Downloads; override with make run-image DOWNLOADS=/mnt/media, and change the mapped port with PORT=9000.

The image carries yt-dlp, ffmpeg, ffprobe, deno and the local CAPTCHA reader beside the binary — the same helpers make dependencies installs — so the hosts that need them work in the container without further setup. It is based on Debian slim because the helpers use glibc. The media helpers are fetched as "latest" and cached with the image, so rebuild with --no-cache (or after docker builder prune) to pick up new releases. OCR dependencies are pinned.

Usage

heapleach                                     serve the UI and open it
heapleach [options] [download-dir]            serve the web UI
heapleach [options] <url>... [download-dir]   download and exit

  -addr string          listen address; :0 takes any free port (default ":8080")
  -concurrency int      parallel transfers, 1-32 (default 4)
  -dir string           directory to download into
  -password string      password for protected sources (headless downloads)
  -retries int          retries per request and per transfer (default 3)
  -streams int          connections to split a slow file across, 1-16 (default 8)
  -slow-speed size      rate below which extra connections open, per second (default 2MB)
  -max-speed size       ceiling on the total download rate, per second (0 is unlimited)
  -min-free size        room to leave at the destination before starting another
                        transfer (0 turns the check off) (default 10GiB)
  -stall-timeout dur    abandon and retry a transfer stuck this long (default 1m30s)
  -debug                verbose logging (default true; -debug=false disables)
  -open                 open the UI in a browser once it is listening
  -version              print the version and exit

Sizes take a unit — 5MB, 1.5GB, 10GiB — or a plain byte count.

On the command line

Hand it URLs and it never opens a socket: the files are downloaded to disk, progress is printed on the terminal, and the process exits.

heapleach https://example.com/d/abc123 ~/Downloads
heapleach -streams 8 https://example.com/a/one https://example.com/a/two
cat urls.txt | heapleach - ~/Downloads          # "-" reads a list from stdin

URLs and the download directory can come in either order — a URL is anything with an http(s) scheme, so the two can never be confused. Lists read from stdin ignore blank lines and # comments.

It exits 0 when every source was accepted, resolved and downloaded, and 1 when an input was rejected, a source could not be read, or a file failed. Failures are named, so it drops straight into a script or a cron job. Ctrl-C exits 130 and leaves the partial files in place; running the same command again continues them rather than starting over.

Everything the server does, this does: the same extractors, the same multi-connection transfers, the same resume. Debug logging is enabled by default, with plain progress lines so logs and the display never fight over the screen. Use -debug=false or HEAPLEACH_DEBUG=0 for quieter output and animated progress on a terminal. Pipes, log files and CI always use plain progress lines.

The download directory can be given three ways. The most explicit wins:

heapleach /mnt/media                    # argument
heapleach -dir /mnt/media               # flag
HEAPLEACH_DIR=/mnt/media heapleach           # environment
heapleach -concurrency 8 -addr :9000 ~/Downloads

It is created if missing, ~ is expanded, and the process refuses to start if the directory cannot be written to — so a permission problem surfaces once, up front, instead of as a wall of failed transfers.

The queue itself is remembered between runs, in ~/.local/state/heapleach/queue.json.zst (XDG_STATE_HOME is honoured, and HEAPLEACH_STATE overrides both). Restarting brings back the list, and anything unfinished comes back held rather than running: press resume, or retry the one job you want, and it picks up from the part files already on disk. Nothing is fetched until you say so, so a machine that reboots at 3 a.m. does not start saturating the line on its own.

An unfinished job is read from its source again rather than replayed. That is not caution but necessity: several hosts sign their media links for twenty minutes or so, and the ones that do carry no link at all until it is minted per attempt, so a stored URL would be a dead one. Re-reading gives fresh links, files already complete are skipped, and part-finished ones resume.

The file is written with mode 0600, because it lists everything you are downloading and carries the password for any source that needed one. Given URLs on the command line the whole mechanism is off: that mode downloads and exits, and has no queue worth outliving it.

If an existing queue cannot be read, the service logs the reason and disables queue saving for that run, preserving the original file. Downloads can still run, but their queue changes will not survive a restart. Stop the service, back up and repair or move the unreadable file, then restart to enable saving again. Automatic migration from queue.json is limited to the default state location; an explicit HEAPLEACH_STATE never adopts a neighbouring file.

A transfer only starts when at least HEAPLEACH_MIN_FREE — 10 GiB by default — is still free there. Below that the queue waits rather than failing: nothing new begins, transfers already running finish normally, and the moment room is made the queue picks up by itself with nothing to re-add. The figure says downloads held while that is the case, so a queue sitting still is never a mystery. Set it to 0 to turn the check off.

The UI shows what is left on that filesystem beside the path, and says it in red once a tenth or less remains: a queue can be larger than the room for it, and the useful moment to notice is before the last block goes rather than after. A destination that cannot be measured shows nothing at all rather than zero, which would read as a disk with no room left.

Network access and trust

HeapLeach is a single-user service with no login or API authentication. Anyone who can reach its API can inspect the queue, submit downloads and change the destination to a directory writable by the service account. Submitted sources can reach local network addresses as well as public sites. Use a trusted network or bind to loopback, for example heapleach -addr 127.0.0.1:8080. The bare desktop launch uses loopback, while the normal -addr default and the container port listen on all interfaces. For remote access, put an authenticated HTTPS reverse proxy in front and restrict direct access to the backend port. Browser cross-origin protection and framing restrictions do not authenticate API callers.

Outbound native HTTP requests honour HTTP_PROXY, HTTPS_PROXY and NO_PROXY. A proxied request uses Go's standard TLS transport rather than the browser-shaped handshake. Helper programs use their own proxy support. The destination directory, optional helper executables and script overrides must be writable only by people you trust to run code as this account.

Supported links

Host Accepts How it resolves
gofile /d/<code> Mints a guest token, then signs each API call with sha256(ua :: lang :: token :: 4h-slot :: secret) in an X-Website-Token header. The secret is rotated server-side, so it is read out of gofile's own obfuscated script rather than pinned: its string table is RC4 under a shuffled base64 alphabet, and the one rotation that decodes the whole file is what proves the answer. A signature gofile rejects is not repeated — it blocks addresses that sign badly. Downloads carry the token as a cookie. Walks nested folders.
bunkr /f/<slug>, /a/<slug>, any bunkr* domain Page → numeric file id → metadata endpoint → separate signing service for the token/ex pair the CDN demands.
pixeldrain /l/<id>, /u/<id>, /f/<id> Public JSON API.
keep2share /file/<id>, with or without a filename, on k2s.cc or keep2share.cc Free downloads through the public API. Reads the image CAPTCHA locally, waits for the host's timer, and reuses the resulting link when resuming. Requires make captcha-helper; included in the container image. Premium-only and private files are reported as restricted.
fileboom fboom.me/file/<id>, with or without a filename Uses FileBoom's own API with the same free-download protocol as Keep2Share, including local OCR and optional proxy routing. Tickets, scores and cooldowns are tracked separately for each service.
filejump filejump.com/s/<id>, including ?access=download Reads the public share's filename and download button, then uses the native ranged downloader. Keeps the stable download endpoint so each transfer obtains a fresh storage signature. Rounded display sizes are estimates; expired or restricted shares fail before queueing.
turbo /embed/<id>, /d/<id> /api/sign issues a short-lived signed URL.
mega /file/<id>#<key>, /folder/<id>#<key>, and the older #!/#F! shapes Nothing about a mega link is legible to the server: names, sizes and bytes are all encrypted under the key in the fragment. Attributes decrypt with AES-CBC; the payload is AES-CTR, undone as the bytes arrive, so ranges, resume and parallel connections all still apply. A link quoted without its # fragment cannot be opened by anyone.
dropbox /s/…, /scl/fi/…, folder shares Asks for dl=1, with the content host kept as a mirror to fail over to. A folder share downloads as the zip dropbox builds for it.
mediafire /file/<key>/<name>, /folder/<key> The download link is on the file page, signed per visit, so the page is re-read when the item starts. Folders come from mediafire's own listing API, recursively.
google drive /file/d/<id>, ?id=<id> Walks past the virus-scan confirmation page — HTML, with a token minted per visit — so the transfer is handed real bytes. Folder links need an API key and are not supported.
svtplay /video/<id>/…, /<programme> Open player API. A programme expands to every episode, grouped into season folders where there is more than one season, read from SVT's GraphQL API rather than the page (which mixes in trailers and recommendations).
youtube videos and playlists Handed to yt-dlp, which is the only practical way in: YouTube withholds playback URLs behind BotGuard attestation on top of its signature and throttling parameters. Needs make dependencies.
bandcamp albums, tracks, a band's discography Read natively, no yt-dlp. Each track is filed as <band>/<album>/<NN> - <title>.mp3; a band's page resolves every release it lists. The MP3 is the 128 kbps stream anyone may play — nothing better is served without a purchase.
odysee, dailymotion, bilibili, niconico, rumble, soundcloud, mixcloud videos and tracks Also handed to yt-dlp, each for a reason recorded in the code so nobody re-derives it: Bilibili signs its playback URLs and ships demuxed DASH; Niconico's HLS is demuxed too, so there is nothing self-contained to fetch; Dailymotion fingerprints TLS and no longer offers a progressive rendition; SoundCloud needs a client id scraped from a rotating bundle; Mixcloud obfuscates its stream URLs; Rumble puts an interstitial in front of its media endpoint. Needs make dependencies.
vimeo /<id>, /<id>/<hash>, channel and group links, player.vimeo.com/video/<id> Everything goes through the embed player rather than the watch page: the watch page answers a non-browser client with a bot check, and yt-dlp's own path through it demands an account, while the player URL an iframe loads is served to anyone. The stream is demuxed — video renditions plus a separate audio group — and carries no progressive MP4, so yt-dlp muxes it. Needs make dependencies.
booru tag searches and posts on 19 image boards, plus the booru.org network One adapter per API family — Danbooru, e621, Moebooru, Gelbooru 0.2, Philomena — which is how a handful of engines covers a lot of sites.
4chan /<board>/thread/<id> The site's read-only JSON API; every attachment in a thread.
vids.st video and embed pages The page names the uploaded file itself, which is downloaded as it is, or an HLS playlist, followed natively. The site sends an incomplete certificate chain, which is completed from the certificate's own issuer address the way a browser does.
lulustream video, embed and download pages The player's setup is packed, and names an HLS playlist whose segments are AES-128-encrypted; they are decrypted natively as they arrive, and the playlist is read again when the file's turn comes, since its links are signed for eight hours.
streamtape, doodstream, mixdrop watch and embed pages Each assembles its link inside the player: two halves joined at an offset the page states, a token endpoint plus a random tail, or a packed script naming the delivery host. Every one of them reads the numbers it needs off the page rather than hard-coding what they were.
pixhost /gallery/<code>, /show/<group>/<file> A gallery's thumbnails say where the full images are: the two links differ only by host prefix and one path segment. Deriving them resolves a gallery of fifty in one request instead of fifty, and a mapping that ever changed would fail visibly with a 404 rather than quietly fetching thumbnails.
suvobox /a/<id>, /f/<id> The album listing states every file's id, its full name with extension and its size, so a whole album resolves in one request with real names in the queue from the start. The bytes come from the media host's ?raw=1, which needs no token — the site's own signed ?dl=1 link would only expire while an item waited its turn.
imagepond /i/<code>, the title-slug form of the same page, /a/<code> albums, and links straight to media.imagepond.net The metadata names the item, but for a video it names the poster frame, and it has been seen misreporting a QuickTime file as MP4 — so the player element's own data-src is read first and the metadata is the fallback. An album arrives whole in one document, but its cards link only to each item's viewer, so the files behind them are resolved one at a time as they download rather than in a burst of fetches up front. An item the host has aged out answers 200 with an expiry notice rather than a 404, and is reported as expired instead of as a parse failure. A link to the media host is the stored file already, and is taken as given — that host is a subdomain of the site, so this extractor claims it and nothing else would. Profile pages list their items client-side and are still not supported.
yandex /video/preview/<id>, on any of its country domains Yandex hosts none of this: a preview is a viewer wrapped around somebody else's video, and the page links the source beside the player. That link is what is followed, and the video is left to yt-dlp, which knows far more hosts than this program does. Needs make dependencies.
ok.ru /video/<id>, /videoembed/<id> The watch page carries an empty rendition list for a logged-out caller, and the player's metadata endpoint answers the same way; the embed page, which exists to be framed elsewhere, carries the same structure filled in. Links are signed with an expiry and the requesting address, so they resolve at download time.
imgur /a/<id>, /gallery/<slug>-<id>, /<id>, /user/<name>, /r/<sub> The album page carries the same JSON the API returns, so no key is needed — and none is accepted either; the header imgur's own web app sends is not validated. A deleted image redirects to a real 503-byte PNG with a correct length that answers a ranged request, which no general rule can catch, so the extractor recognises that dead end itself.
cyberdrop /a/<id>, /f/<slug> bunkr's sibling. The album page states every filename and exact size in one request; the signed link is minted per attempt, which is what lets a file that will not sign fail on its own rather than sinking the album.
archive.org /details/<id>, /download/<id> One metadata call lists every file with an exact length and checksum. It is fetched gently — one connection, one file at a time — because archive.org answers parallelism with a clean 206 served 300× slower, per address, for minutes afterwards. A collection identifier is refused rather than "succeeding" by fetching five site logos.
bitchute /video/<id>, /channel/<name> Two unsigned JSON posts give a permanent MP4. Eleven seed hosts serve each other's paths byte-identically, so they are kept as mirrors — failover onto a set that cannot be wrong.
civitai /images/<id>, /posts/<id>, /models/<id>, /user/<name>/images Documented public JSON, unsigned content-addressed CDN. Collection links are deliberately not matched: the API silently ignores that parameter and returns the site's front page instead, which would look like a successful download of the wrong thing.
pornpics /galleries/<slug>-<id>/, categories, /pornstars/, /channels/ A gallery's full-size URLs are in the page, so one request resolves it. Listings take one of two pagination routes, and the wrong one returns the first 20 items forever while looking like it works.
bandzoogle (family) any install One extractor for the band-website product, which musicians rent and run on their own domains — so unlike the families above it ships with no host list at all, and every site is found by the sniff. The pages label their own player: each track anchor carries the title, the artist and the path the audio comes from, so one fetch yields the lot. Files are named NN Title.ext under a folder per album, taking the position the page prints beside each track and padding it so a directory listing sorts the way the album plays — the folder matters because the numbering restarts with each album. A listing shows twenty and says whether more follow, so the rest are fetched too rather than the album arriving quietly short. The audio link is the site's own /player/…/tracks/….mp3, which redirects to storage signed for a couple of hours; the player link is what is kept, since the signature is minted per request and a stored one would go stale in the queue.
chevereto (family) imgbb, freeimage, gifyu, and any install One extractor for the image-host product, recognised by its own generator tag, so an unlisted install works without a rebuild. imagepond was one of these once and is not any more — it was rewritten onto software of its own and keeps its own extractor, which is why it is listed separately above.
peertube (family) /w/<id>, /videos/watch/<id>, channels, accounts ~1,795 federated instances behind one versioned API, confirmed by a version probe rather than guessed from markup. Unsigned, rangeable, exact lengths. A federated video's file may live on a different instance than the one asked, and the API says which.
fediverse (family) Mastodon, Pleroma, Pixelfed, Lemmy, Misskey Found through /.well-known/nodeinfo, which is a specification rather than a list, so the table cannot go stale. Three API dialects behind one sniff.
foolfuuka (family) /{board}/thread/<n> on the 4chan archives Where threads persist after 4chan drops them, through an API modelled on 4chan's own — and better than it, since each post carries the poster's original filename and an exact size.
mediawiki (family) Category:, File:, articles, on any wiki Commons, every Wikipedia, Fandom, any open wiki. A documented, versioned, keyless API returning untouched originals — the least likely thing here ever to break.
nrk, rúv, ard, zdf, srf, vrt, npo, raiplay, rtve, rtp, pbs, npr, abc listen, bbc sounds programmes, series and episodes svtplay's siblings, thirteen more of them. Each checks the broadcaster's own DRM flag, skips what is protected, and passes the site's own geo-block message through verbatim rather than replacing it with a guess. Whether one of these is here or on the yt-dlp list above is decided by one thing: whether any rendition carries audio and video in the same segments. Those that do are fetched directly; those that do not would concatenate into a silent video, so they are muxed by yt-dlp instead. ARD and RTVE hand back plain rangeable files with exact lengths, so they get the whole engine — splitting, resume and the skip check.
vidmoly, streamable, wetransfer see above Vidmoly is the highest-traffic embed host here; three of its four advertised domains are dead, so everything routes through the one that works.
feeds any RSS or Atom feed One file per <enclosure>, oldest first, since an archive wants the beginning.
a bare .m3u8 any adaptive manifest Joined into a playable file. A .mpd is refused with a reason rather than saved: DASH is usually demuxed, so concatenating it would yield a silent video.
a directory listing Apache, nginx, lighttpd autoindex Walked recursively, with sizes and structure taken from the listing. Also covers IPFS gateway directories.
links:<url> any public page Reads the page and downloads every link a supported host claims, each into its own folder. Aimed at the forum thread with two hundred links in it.
an ordinary webpage any public HTML page without a dedicated extractor Automatically reads links to every registered host, including albums and files, without a links: prefix. Ordinary HTML links are ignored so navigation cannot recursively crawl the web. Links keep their host's resolver and download settings; source and file caps apply.
DarkGram albums any HTML page with a DarkGram player Reads the first embedded album's videos and full-size photos from its data-album metadata. HLS playlists are resolved at download time and refreshed on retries; photos use their original links, never thumbnails. Files keep the page title and album order. File caps and unavailable items are reported.
WordPress categories any domain; category permalinks, custom category bases and ?cat=<id> Recognises classic and block theme markup without a domain list or version check. Follows the selected category's pagination and reads each post's supported file-host links, embedded albums/players and JPEG images, preferring linked originals and full-size image metadata. File links go through the existing host extractors, including deferred K2S/FileBoom resolution and proxy pacing. Each post gets a folder; player thumbnails, navigation and related-post widgets are excluded. Source/file caps and incomplete results are reported.
voyeurking voyeurking.com/categories/<name>, /collection/<slug> and /video/<slug> Walks category and collection pages and reads each video's K2S or FileBoom file from its page data. Files use the corresponding downloader, including its per-address waits and proxy pool, enabled by default. Listing results keep each video in its own folder; source and file caps apply.
balbums balbums.st/?search=<query> An index of somebody else's albums rather than a host: a search is walked page by page and every album it lists is handed to the extractor that does host it, each into its own folder. Asking for a hundred results a page is what makes a search of three hundred albums four requests rather than eighteen — ask for more and the site quietly serves twenty, so nothing trusts the parameter. Only searches are taken. The charts are built in the browser and the front page is the same grid over the whole catalogue.
KVS listings /members/<id>/, /search/<query>/, and any category, model, tag, channel or site-wide list on any KVS install A member's public videos, everything a search turns up, or every video a category, model, tag or channel lists. Sections are recognised by the page rather than by name, since an install may rename them. The platform pages through its own asynchronous block loader, and which parameter pages a block is the block's own business — a wrong one is not refused, it serves page one again — so the walk sends exactly what the pager's own control carries. The last page says it is the last, and the walk stops there without asking.
anything else any http(s) URL Treated as a direct file link after checking for a known platform, a manifest, a directory index, supported download links or a player.

Several of these hand out URLs that expire in minutes, so bunkr, turbo, mega, mediafire, ok.ru, cyberdrop, streamable, wetransfer and the three streaming hosts are resolved at download time, not when the link is queued — otherwise a large queue would start failing halfway down.

Several entries above are platform families: one extractor covering sites running the same software. On unregistered hosts, Pixeldrain-compatible services, booru boards, FoolFuuka archives, Kemono-compatible services and MediaWiki articles are detected through distinctive API responses or platform markup. KVS detection also covers member, search and other listing pages through its player or asynchronous block controls. Existing PeerTube, Chevereto and Bandzoogle detection continues to work the same way.

A familiar URL path selects a possible API; it does not establish support. Guessed API requests are bounded and do not retry. Recognised platforms use their normal extractors, including pagination, deferred resolvers and transfer pacing. Unidentified pages continue through the generic fallback. Known-host registrations remain useful for discovering embedded download links, naming supported sites, and retaining aliases and site-specific settings such as booru filters. They are shortcuts, not a requirement for detecting a compatible platform from a pasted URL.

Other entries are not hosts at all but shapes — an adaptive manifest, an open directory, a page carrying supported download links, and a page carrying a video in its markup. Those cover the sites nobody will ever get round to naming.

Automatic webpage scanning reads anchors, embedded frames and links printed as text, then expands only links claimed by registered extractors. A recognised WordPress category also follows its own post permalinks and pagination; ordinary navigation links are not followed. Exact duplicates are removed; K2S aliases and repeated file IDs are also merged within each service. Recognised platforms keep their own extraction behavior. Use links:<url> to explicitly scan a page that would otherwise go to a dedicated extractor.

WordPress category extraction uses the HTML the site serves. It supports classic post markup, block query loops, older/newer-post navigation, and path or query pagination. A pasted later page starts from the category's first page. Posts without supported file-host links use the shared album and media-player parsers on their content, with DarkGram playlists still resolved at download time. Separate JPEG photos are kept and duplicates removed. Login-only content and listings rendered entirely by scripts are not exposed by this parser. Domain redirects are followed, and relative links use the final page's address.

Debug logs trace WordPress category pages, post counts, pagination and each post's linked sources, including player/proxy progress and failures. Files reach the queue after source resolution finishes; these logs show which post and source are still being resolved during that wait. Player diagnostics also record response status, content type, redirects, challenge markers and connection route for discovery and token refreshes. Debug logging defaults on; use -debug=false or HEAPLEACH_DEBUG=0 to disable it.

Open Logs beside Settings to view the service's recent messages without opening a terminal. The panel updates while visible, with search, severity filters, pause/resume and a download of the filtered messages. Scrolling up stops following new messages; Jump to latest resumes following. The service retains up to 1,000 messages or 4 MiB for the current run, truncating unusually large entries. Older messages roll off and a restart clears this history. Console logging continues as usual.

Every supported site

Generated from the registry by make hosts, so it cannot drift from the code: adding a host to NewRegistry is the only step, and CI fails if this section and the binary disagree.

194 sites across 102 extractors, plus 6 that match by shape rather than by host.

Extractor Sites
4chan 4chan.org, 4channel.org
abclisten abc.net.au/listen
alohatube alohatube.com
archive.org archive.org
ard ardmediathek.de
arte arte.tv
balbums balbums.st
bandcamp bandcamp.com
bbcsounds bbc.co.uk/sounds
bilibili b23.tv, bilibili.com
bitchute bitchute.com
blender video.blender.org
booru aibooru.online, booru.borvar.art, booru.foalcon.com, booru.org, derpibooru.org, e621.net, e6ai.net, e926.net, furbooru.org, hypnohub.net, konachan.com, konachan.net, ponybooru.org, safebooru.org, sakugabooru.com, snootbooru.com, tbib.org, twibooru.org, xbooru.com, yande.re
bunkr bunkr.*
camwhores camwhores.tv
celeb celeb.st
civitai civitai.com
coomerfans coomerfans.com
cyberdrop cyberdrop.cr
dailymotion dai.ly, dailymotion.com
desuarchive desuarchive.org, rbt.asia
doodstream d000d.com, d0o0d.com, do0od.com, dood.la, dood.li, dood.re, dood.sh, dood.so, dood.to, dood.watch, dood.ws, dood.yt, doodstream.com, dooood.com, ds2play.com, playmogo.com, vidply.com
drive docs.google.com, drive.google.com, drive.usercontent.google.com
dropbox dropbox.com, dropboxusercontent.com
drtuber drtuber.com
drtv dr.dk
eporner eporner.com
erome erome.com
fapello fapello.com
fapster fapster.xyz
fileboom fboom.me
filejump filejump.com
filester filester.*
framatube framatube.org
francetv france.tv
freeimage freeimage.host
gifyu gifyu.com
gofile gofile.io
imagepond imagepond.net
imgbb ibb.co, imgbb.com
imgur imgur.com
keep2share k2s.cc, keep2share.cc
kemono coomer.party, coomer.st, coomer.su, kemono.cr, kemono.party, kemono.st, kemono.su
knuthing knuthing.com
kolektiva kolektiva.media
lulustream lulustream.com, luluvdo.com, luluvido.com
makertube makertube.net
mediafire mediafire.com
mega mega.co.nz, mega.nz
mixcloud mixcloud.com
mixdrop mdbekjwqa.pw, mixdrop.ag, mixdrop.bz, mixdrop.ch, mixdrop.club, mixdrop.co, mixdrop.gl, mixdrop.is, mixdrop.my, mixdrop.ps, mixdrop.sx, mixdrop.to, mixdrp.co, mixdrp.to
moannest moannest.com
niconico nico.ms, nicovideo.jp
npo npo.nl, npostart.nl
npr npr.org
nrk tv.nrk.no
odysee lbry.tv, odysee.com
ok.ru odnoklassniki.ru, ok.ru
orf on.orf.at, orf.at, tvthek.orf.at
palanq archive.palanq.win
pbs pbs.org
pixeldrain nova.storage, pixeldrain.com
pixhost pixhost.to
pornhits pornhits.com
pornhub pornhub.com
pornone pornone.com
pornpics pornpics.com
porntrex porntrex.com
pornzog pornzog.com
radio-canada ici.radio-canada.ca, radio-canada.ca
raiplay raiplay.it
redgifs redgifs.com
rtbf auvio.rtbf.be, rtbf.be
rtpplay rtp.pt
rtve rtve.es
rumble rumble.com
ruv ruv.is
sexvid sexvid.xxx
soundcloud soundcloud.com
spectra spectra.video
srf playsuisse.ch, srf.ch
streamable streamable.com
streamtape strcloud.link, streamta.pe, streamtape.cc, streamtape.com, streamtape.net, streamtape.site, streamtape.to, streamtape.xyz
suvobox suvobox.com
svtplay svtplay.se
tchncs tube.tchncs.de
thisvid thisvid.com
tilvids tilvids.com
tnaflix tnaflix.com
turbo turbo.cr, turbocdn.st
vidmoly vidmoly.biz, vidmoly.me, vidmoly.net, vidmoly.to
vids.st vids.st
vimeo player.vimeo.com, vimeo.com
voyeurking voyeurking.com
vrtmax vrt.be
wetransfer we.tl, wetransfer.com
xhamster xhamster.*
xxthots xxthots.com
yandex yandex.*
yle areena.yle.fi, arenan.yle.fi
youtube youtu.be, youtube-nocookie.com, youtube.com
zdf zdf.de

And these match a URL or document shape, so the set of sites they reach is open-ended:

Extractor Recognises
autoindex an open directory listing (Apache, nginx, lighttpd)
fediverse an instance publishing /.well-known/nodeinfo — Mastodon, Pleroma, Pixelfed, Lemmy, Misskey
feed an RSS or Atom feed, by its enclosures
hls an adaptive manifest (.m3u8), from any host
links a links:<url> prefix — every supported link on the page
mediawiki any wiki with an open api.php — Commons, Wikipedia, Fandom

Where a host offers the same file from more than one place — gofile's storage servers, dropbox's two front ends — the alternatives are kept as mirrors and a failed attempt lands on a different one.

SVT Play is an adaptive stream, so it has no single file to fetch. The extractor picks a rendition carrying audio and video together, which means the segments join into a playable file with no muxing at all.

Any finished .ts is rewrapped to .mp4 when ffmpeg is available. The rewrap is always a stream copy, so nothing is ever re-encoded: if the streams cannot enter MP4 untouched, the .ts is kept instead. What comes over is the best video plus the best audio and subtitle track of each language, so a recording that shipped in several languages stays usable in all of them without also carrying the duplicate encodes within each. Without ffmpeg the .ts is kept and plays fine.

Features

  • Parallel downloads with a live worker count you can change while transfers are running.
  • Multi-connection transfers. A file that is downloading slowly is split across more connections, up to a configurable ceiling. Ranges are chosen by bisecting whatever is still outstanding — halfway, then the quarters, then the eighths — so the extra connections share the remaining work rather than duplicating it. Throughput is averaged over a window before any decision is made, so one noisy interval never triggers a change.
  • Host-aware. Hosts cap how many connections they will accept. Extra connections are budgeted per host across all downloads, and a host that turns one away has its range handed straight back and its budget lowered, so a working download is never failed by trying to go faster.
  • Patient with busy hosts. Gofile answers a request for a file on a busy storage server by redirecting to its own web page. That is detected rather than saved to disk, and retried with backoff — rotating through the file's other storage servers where it has them — because a busy host has not failed, it has asked us to come back. Patience is spent only where trying again can change the answer: a URL with nothing to re-resolve, which is usually a page no extractor recognised, is told so after a few attempts instead of retrying forever.
  • Forgives a dropped connection that was getting somewhere. A transfer that loses its connection after moving bytes resumes from its part file thirty seconds later, and that attempt is not counted against the retries: the budget counts attempts in a row that move nothing, and one that moved anything starts the count over. A stall after progress is treated the same way. So a CDN that drops every few dozen megabytes still finishes a large file instead of failing it three drops in — and since every such attempt leaves more on disk than it found, a finite file cannot cycle forever.
  • Pause and resume the queue. Native transfers park inside their reads rather than being torn down, so a short pause costs nothing and a long one falls back on the same resume every other interruption uses. Already-running external helpers continue; pausing prevents new ones from starting.
  • A ceiling on native download throughput, set from the header or with -max-speed. It is one shared budget across native connections, not a per-file allowance — and the code that opens extra connections knows about it, so a transfer held at the ceiling is not mistaken for a slow one and split eight ways for nothing.
  • Skips what is already there. A file whose name and length already match the destination is not downloaded again — checked against the length the server reports, so it works even for hosts that publish no sizes. Sizes read off listing pages are rounded, and are never used to make that call. Entries in one listing that share a sanitized destination get distinct, stable numbered names, including when their lengths match.
  • Notices a stalled transfer — and steps around it. A connection that stops delivering without closing is invisible to a read timeout; if the byte counter has not moved for -stall-timeout, the attempt is abandoned and the item goes to the back of the queue with its partial file intact, so a host that has stopped serving does not pin a worker while everything behind it waits. Its next turn resumes from disk; a host that never resumes still fails the item once the retry budget is spent. Playlist downloads are watched the same way, against the bytes actually arriving rather than against whole parts landing — a large part fetched slowly is progress, a silent connection is not. And a stalled transfer is never "helped" by splitting it further: slow means more room than one connection uses, stalled means the host is serving nothing, and the two get opposite answers.
  • Live progress over server-sent events: per-file bytes, rate, ETA, and parts joined for a file that arrives as a playlist and so has no byte total until its last part lands. A transfer waiting on purpose says so, rather than looking like one that has died.
  • Connection recovery. If an event stream opens without delivering its first snapshot, the UI retries and polls state in the meantime. A partially accepted submission keeps rejected URLs in the form for correction.
  • A searchable queue. Type / and filter hundreds of jobs by title, source, host or filename; Esc clears. Scrolling is not a retrieval strategy.
  • The destination is changeable while it runs. Click the path in the header, type another. It is expanded, created and proved writable exactly as a directory named on the command line is, and it says why if it cannot be used. Transfers already running keep the destination they started with — their path was settled when they began, and a part file that moved mid-flight could not be resumed — so the change takes hold from the next queued file onward.
  • Cancel and retry a whole job or a single file.
  • Resume — partial files are kept and continued with a Range request; a cancelled 100 MB download restarts where it stopped.
  • Password-protected folders (gofile).
  • Safe filenames — remote names are reduced to one portable path component, so nothing can be written outside the download directory.
  • A light progression layer: ranks, session totals and a few badges.

Image boards

Boards sharing an API are covered by one adapter each, in the style gallery-dl uses:

Family Boards
Danbooru aibooru, booruvar
e621 e621, e926, e6ai
Moebooru yande.re, konachan, sakugabooru
Gelbooru 0.2 safebooru, tbib, hypnohub, xbooru
Philomena derpibooru, ponybooru, furbooru, twibooru
szurubooru snootbooru, foalcon
Gelbooru 0.1 the booru.org network, one board per subdomain, scraped — it has no JSON API

Paste a tag search or a post link. A listing with no tags fetches the board's latest posts; a bare domain is rejected, since that is far more likely a mis-paste than a request for everything.

Every board above was checked against its live API. Ones that now demand an API key (gelbooru.com), sit behind a challenge (danbooru.donmai.us) or have switched their API off (realbooru) are deliberately absent rather than listed as supported and quietly broken.

Optional helpers

YouTube — and every host marked "Needs make dependencies" above — is fetched by yt-dlp, which is asked for every audio language a video carries rather than only the one it ranks first: a dubbed release puts its original track at the top, so the English beside it would otherwise be dropped. The languages arrive as separate tracks in one file, labelled, for the player to choose between. ffmpeg rewraps any finished .ts as .mp4 and lets yt-dlp merge separate video and audio tracks. deno runs the player JavaScript YouTube signs its media URLs with: yt-dlp still extracts without a runtime, but warns that doing so is deprecated and that formats may be missing. Everything else needs none of them.

make dependencies    # yt-dlp + ffmpeg + deno + local CAPTCHA reader into ./bin

The service looks for these next to the heapleach binary first, then on PATH — so the copies in ./bin are picked up without touching the system. That is also why deno is passed to the helper by path: yt-dlp finds one on PATH by itself, and ./bin is the place it would not look.

Keep2Share and FileBoom free downloads use heapleach-ocr, a separate helper containing the ddddocr recognition model and its runtime. Recognition happens locally, without an account, API key or paid solver. make captcha-helper builds just this helper with Docker on Linux (glibc 2.36 or newer); make dependencies includes that step. The container image already carries it. Native builds on other platforms are described in helpers/captcha/README.md.

The free flow tries at most three CAPTCHA images, and at most three readings of each, ranked by the reader's confidence. It shows the host's waiting time, and allows one file and one connection at a time per service, shared across its aliases. Unreadable challenges fail with a retryable error; OCR is not always correct. With proxies enabled (the default), the one-file limit applies to each proxy address, allowing several free downloads at once without a subscription, up to -concurrency across both services. Each address may carry one K2S transfer and one FileBoom transfer simultaneously. The hosts' speed limits and cooldowns still apply per address and service. A file held back by the wait between free downloads goes back to the queue as waiting, with the time left, and its download slot goes to other files meanwhile. Waiting is cancelable and bounded to two hours per attempt; retrying within the same process preserves an accepted ticket and reuses unexpired download links.

Open Settings in the UI to change files downloaded at once, streams per file, and the speed limit. The same panel enables or disables download proxies and edits proxy endpoints and discovery feeds without restarting HeapLeach. These settings apply to the current session; flags and environment variables set the startup defaults. Active transfers keep their assigned route. When switching routing modes, a service's existing transfers finish before new ones use the new mode, preserving the per-address limit and download tickets.

The proxy list shows the selection score, measured average and current throughput, recent success rate, request count, active transfers, cooldowns, and last successful request. Choose Keep2Share, FileBoom or WAF recovery to view the corresponding scores. The WAF view also shows challenge success separately from ordinary request success. Search, status filters and sorting work across the entire inventory, with 50 rows per page. Scores are estimated useful bytes per second using the selector's mean success probability; actual selection also explores untried routes. The list refreshes every five seconds while open. Discovery status is shown with each feed.

Free proxy routes are enabled at startup. Disable them with -proxies=false, HEAPLEACH_PROXIES=0, or the Settings switch. The default pool combines the normal outbound connection with Proxifly's free proxy feed. HTTP, HTTPS and SOCKS5 endpoints are supported; SOCKS5 resolves destination names through the proxy. Plain host:port entries mean HTTP. Set HEAPLEACH_PROXY_ENDPOINTS to a comma-separated list of explicit endpoints, including direct if the normal connection should participate, and HEAPLEACH_PROXY_FEEDS to plain-text or Proxifly JSON feed URLs. An explicitly empty feed setting disables discovery; an empty endpoint setting excludes the normal connection. Explicit proxies take precedence over NO_PROXY.

Each transfer attempt keeps its route through the CAPTCHA, ticket, redirects and file requests. Keep2Share and FileBoom API calls and CAPTCHA image requests have a 20-second timeout. A network timeout penalises and cools the route before another is tried. Status notes show the connection attempt number on retries, separately from the CAPTCHA image count. A refused or broken route returns the file to the queue to try another; other hosts continue downloading while routes are busy or cooling down. Keep2Share and FileBoom still require the local OCR helper, and premium-only files remain restricted.

Every host using HeapLeach's HTTP client can recover through proxies after a Cloudflare WAF challenge, including unregistered pages, direct files and Doodstream/Playmogo. Ordinary requests use the normal connection first. Challenges during extraction retry the source operation; challenges during downloads return the file to the queue. A route change re-reads the source to obtain fresh signed links and keeps the resolver, cookies and media on the chosen route. Attempts use the configured proxy retry budget, and waiting for an extraction route is bounded by the request timeout.

Doodstream also switches to this proxy pool after temporary network/HTTP failures or repeated RELOAD replies from its token endpoint. It tries the normal connection first. Each retry refreshes the player and token together on one reserved route, using the same random draw for unmeasured candidates as Keep2Share. Source discovery and file downloads both skip dead or dropped proxies without spending the file's retry budget, including timeouts waiting for a proxy connection. Response-header timeouts also penalise the proxy and retry another route, within the configured retry budget. User cancellation does not penalise a route. Missing videos and parser errors remain final. Progress messages show the player/token stage and connection attempt in the browser and terminal, including while the source is still being resolved.

WAF refusals affect only the WAF success score and cooldown. They do not reduce ordinary service scores, erase speed measurements or mark a proxy broken. WAF selection still uses ordinary reliability and measured transfer speed, excluding user-throttled samples and free-host caps when estimating unrestricted capacity. Keep2Share and FileBoom retain their per-service address limits even when WAF recovery is needed.

This tries other addresses; it does not solve interactive challenges. With proxies disabled, a challenge remains an error. External downloaders such as yt-dlp manage their own HTTP requests and are outside this recovery path.

The pool adapts amzscrape's proxy logic and stores inventory, source memberships, request outcomes, measured bytes per second, and cooldowns in a private bbolt database. Faster, reliable routes score higher; bodies under 64 KiB do not train bandwidth, so a quick CAPTCHA response cannot stand in for a fast file transfer. Recent measurements carry more weight. Refusals cool only that service; connection failures cool the endpoint across services. HTTPS proxies may use self-signed certificates on the connection to the proxy. Destination certificates are still verified inside the tunnel. An explicit proxy never silently falls back to a direct connection.

The scheduler maximises expected useful bytes per second within the live file-concurrency setting and each service's one-file-per-address rule:

  1. For each free worker, score available routes for the next file's remaining bytes B: p × B / (setup + B / speed + (1 − p) × recovery). p is sampled from recent success/failure evidence, speed is measured throughput, and setup includes CAPTCHA and ticket waits. Unknown setup starts at 30 seconds; unknown speed uses the pool's learned average. A healthy preferred route can reuse the file's existing ticket. The UI shows the deterministic mean score for a 32 MiB transfer; an actual dispatch uses that file's remaining size.
  2. Update speed from useful-byte windows every five seconds once at least 64 KiB has arrived. Resumed bytes are excluded. Pause and speed-limit periods do not train throughput or setup estimates. These live samples update both route selection and the proxy list before the file finishes, and the background task saves them to BoltDB.
  3. Give unmeasured routes one shared exploration opportunity, choosing its candidate uniformly at random from the available unmeasured endpoints. Busy and cooling addresses are excluded before the draw; selection and reservation are atomic. While measured routes are free, allow at most one active unmeasured transfer and space exploration starts by 30 seconds. If no measured route is available, use idle workers to discover capacity instead of leaving them unused.
  4. Reassess slow transfers after 20 seconds of samples. A proven alternative must predict at least 25% higher useful throughput, save at least 15 seconds after setup and recovery costs, and win two checks ten seconds apart. Switch only after the server has confirmed range support and at least 4 MiB remains. Reserve the new address first, stop and flush the old attempt, then obtain a ticket on the new route and resume the same file. An unknown route never displaces a working transfer just for exploration.

This is an adaptive estimate: public proxy capacity changes and cannot be known in advance. Setup costs, confirmation windows and ticket affinity keep small, noisy speed changes from spending time on repeated route switches.

Saved inventory is available immediately after restart. Feed requests, including GitHub, are normally 24 hours apart. An earlier top-up requires runnable K2S, FileBoom or WAF recovery work and too few distinct usable addresses for its current concurrency, with at least five minutes between attempts. New active transfers count while their first speed sample is pending; measured and standby routes must have succeeded and score at least as well as an untried route's learned prior. Throughput, reliability and cooldowns therefore matter, not the raw list size. An idle or paused queue cannot trigger an early reload. The refresh schedule survives restarts. A failed feed keeps its last good list and refreshes preserve health. Entries absent from feeds are retired after 30 days without success; current feed entries and explicit endpoints are retained. Public proxy availability and bandwidth vary, so extra parallel downloads depend on finding usable routes. The database allows one HeapLeach process at a time; use a separate -proxy-db for a second instance, and mount its directory on a persistent volume in containers.

The YouTube download itself runs through yt-download.sh rather than inline Go, so the recipe is in one readable place. A copy of that script placed beside the binary overrides the built-in one, so it can be adjusted without rebuilding. yt-dlp's pieces (each stream, the merge in progress) are kept in a hidden .heapleach directory beside the destination, so a media library watching that folder sees only the finished file. The directory is removed once the download completes, and kept after a failure so the next attempt can resume.

Metadata probes have a two-minute deadline and an 8 MiB output limit; helper diagnostics are bounded too. YouTube playlist enumeration uses HEAPLEACH_MAX_FILES at the helper, and a truncated result is labelled partial. On Unix, cancellation kills the helper's process group, including children. The bundled download script needs a POSIX shell; it is not a native Windows helper launcher. External transfers use yt-dlp's own retries and transfer controls, so the native speed ceiling, mid-transfer pause and stall watchdog do not apply to them.

Configuration

Every setting has an environment variable; the common ones also have a flag, and a flag beats the environment. Sizes and rates take a unit — 5MB, 1.5GB, 10GiB — or a plain byte count.

Variable Default Meaning
HEAPLEACH_ADDR :8080 Listen address. Flag: -addr.
HEAPLEACH_DIR your Downloads folder Where files are written. Defaults to the platform's own download folder — ~/Downloads on macOS and Windows, and on Linux whatever the desktop's XDG user-dirs file says, which is where a relocated or localised folder is recorded. The container image uses /downloads instead, having no home directory to speak of. Flag: -dir, or the positional argument.
HEAPLEACH_CONCURRENCY 4 Parallel transfers (1–32). Flag: -concurrency. Also settable live in Settings.
HEAPLEACH_PROXIES on Enable proxy routes for Keep2Share, FileBoom and WAF challenge recovery on any host. Set to 0 or use -proxies=false to disable. Also settable live in Settings.
HEAPLEACH_PROXY_DB ~/.local/state/heapleach/proxies.db Persistent bbolt inventory and health; honours XDG_STATE_HOME on Linux. Independent of queue persistence, including in CLI mode. Flag: -proxy-db.
HEAPLEACH_PROXY_ENDPOINTS direct Comma- or whitespace-separated HTTP, HTTPS or SOCKS5 proxy URLs; direct means the normal outbound connection, including environment proxy settings. Empty excludes that connection. Also editable live in Settings.
HEAPLEACH_PROXY_FEEDS Proxifly's global text feed Comma- or whitespace-separated feed URLs. Empty disables discovery. Also editable live in Settings.
HEAPLEACH_PROXY_RETRIES 20 Route changes after failures that make no disk progress. A proxy that cannot be reached, drops the connection or has its address refused does not count. A host's free-download cooldown does not spend this budget.
HEAPLEACH_MAX_RETRIES 3 Retries per request and per native transfer, counting attempts in a row that moved nothing: an attempt that downloaded anything before failing resumes after 30s and starts the count over. Flag: -retries. Busy responses and rate limits have separate bounded patience; a resolvable busy storage link can be refreshed repeatedly.
HEAPLEACH_STREAMS 8 Connections one slow file may be split across (1–16). Flag: -streams. Also settable live in the UI.
HEAPLEACH_SLOW_SPEED 2MB Rate per second below which extra connections are opened. Flag: -slow-speed.
HEAPLEACH_MAX_SPEED 0 Shared ceiling on the native download rate per second; 0 is unlimited. Does not throttle external helpers. Flag: -max-speed. Also settable live in the UI.
HEAPLEACH_STALL_TIMEOUT 90s How long a transfer may make no progress before the attempt is retried. Flag: -stall-timeout.
HEAPLEACH_MIN_FREE 10GiB Room that must be left at the destination before another transfer starts. Below it the queue waits rather than filling the disk; 0 turns the check off. Flag: -min-free.
HEAPLEACH_STATE ~/.local/state/heapleach/queue.json.zst ($XDG_STATE_HOME when set) Where the queue is written so a restart can pick it up, as zstd-compressed JSON (zstdcat reads it). A plain-JSON file from an earlier version is still read, and a queue.json beside the default path is picked up once and retired. Unfinished jobs come back held, and are re-read when retried or when the queue is resumed; nothing is fetched until then. Empty disables it. A run given URLs on the command line never writes one.
HEAPLEACH_USER_AGENT a current desktop Chrome UA Sent on every request. Gofile mixes it into its signature, so it must match what signs.
HEAPLEACH_LANGUAGE en-US Accept-Language, and part of the gofile signature.
HEAPLEACH_GOFILE_SECRET read from gofile The secret gofile signs requests with. It is normally recovered from gofile's own script and cached for as long as that script says it is good for, so this is only needed if that ever stops working — setting it overrides the lookup entirely.
HEAPLEACH_EXTRA_HOSTS unset Extra hosts for a platform family, as family:host,host;family:host — for example peertube:tube.example;kvs:tube2.example. Every family here is software many sites run, so a list compiled into a binary can only ever trail them; this adds installs without a rebuild. Families: kvs, peertube, chevereto, foolfuuka, fediverse, mediawiki, bandzoogle.
HEAPLEACH_KVS_HOSTS unset The original KVS-only form of the above, still honoured.
HEAPLEACH_IA_FORMATS unset Archive.org format labels to keep, overriding the per-mediatype rendition policy — the escape for when an item's interesting rendition is one the policy passes over.
HEAPLEACH_UTLS unset HEAPLEACH_UTLS=0 turns the browser-shaped TLS handshake off and uses Go's standard one. A few hosts (wiki.gg) challenge the default fingerprint; a few others require it.
HEAPLEACH_DEBUG 1 Debug logging. Set 0, false, no or off to disable it. Flag: -debug=false.
HEAPLEACH_MAX_SOURCES 500 How many of a page's links are followed: the albums of an index search, the links of a harvested thread. Flag: -max-sources.
HEAPLEACH_MAX_FILES 20000 How many files one submitted URL may resolve to. A listing past it stops at a whole album and says so beside the job's name — never in its folder, so raising the cap and adding the URL again files the rest beside what is already downloaded. Flag: -max-files.
HEAPLEACH_PPROF unset Serve Go's runtime profiles at this address, on a listener of their own — go tool pprof http://127.0.0.1:6060/debug/pprof/profile?seconds=30 then shows where the CPU goes. Flag: -pprof. Off unless set. Name a loopback address: the profiles describe the process from the inside, and a non-loopback one is logged as a warning. In a container, the container's own 127.0.0.1 is reached with docker exec.
HEAPLEACH_RESUME unset Resume the jobs a previous run left unfinished at startup, instead of holding them for Resume or a retry. Flag: -resume. For a service that is redeployed without anyone watching — held would mean stopped until noticed.
HEAPLEACH_OPEN unset Open a browser once listening. Flag: -open. A bare run does this anyway, so this is mostly how to say no: HEAPLEACH_OPEN=0 (also false, no, off) suppresses it, for a machine with no desktop or a session over SSH.

HTTP API

Method Path Purpose
GET /api/health Liveness.
GET /api/state Current snapshot.
GET /api/events Complete initial snapshot, then item patches; missed updates are replaced with a complete snapshot.
GET /api/logs Recent service logs. Pass the returned session and next as session and after for new entries; reset replaces prior history after a restart or retention gap.
POST /api/downloads {"urls": "…", "password": "…"} — newline-separated or an array.
GET /api/settings Current runtime settings, including configured proxy endpoints and discovery feeds.
POST /api/settings Any of {"concurrency": n, "streams": n, "paused": bool, "speedLimit": n, "downloadDir": "…", "proxies": bool, "proxyEndpoints": ["…"], "proxyFeeds": ["…"]} — each optional, so a request carries only what changed.
GET /api/proxies Proxy measurements and feed status. Query: site (keep2share, the default, fileboom, or cloudflare), offset, limit (1–200, default 50), search, status (all, available, active, busy, cooling, untested, finishing), sort (score, throughput, success, address).
POST /api/clear Forget finished jobs.
POST /api/jobs/{id}/cancel · /retry Whole job.
DELETE /api/jobs/{id} Cancel and forget.
POST /api/jobs/{id}/items/{itemId}/cancel · /retry One file.

Settings updates are validated together: an invalid field leaves the current settings unchanged, including when opening the proxy database fails. Proxy inventory is paged separately from queue snapshots; endpoint passwords are redacted there and only included when explicitly fetching settings for editing. In an SSE job marked patch, merge items by ID and retain unmentioned items; otherwise replace its item list. Every frame contains the complete job list and current aggregates.

Oversized JSON requests return 413; malformed JSON returns 400. Unknown API routes return 404, and a known route with an unsupported method returns 405. Ordinary HTTP requests have 30-second read/write deadlines. Event streams use their own per-write deadline and remain open while idle.

curl -X POST localhost:8080/api/downloads \
  -H 'Content-Type: application/json' \
  -d '{"urls":"https://pixeldrain.com/l/<id>"}'

Architecture

Two diagrams, for the two questions worth answering: what happens to a link, and what is allowed to depend on what.

What happens to a link

The middle of this is the part worth knowing. An extractor downloads nothing — it returns one File per download, and the shape of that File is what decides everything downstream. A plain URL reaches the segmented engine. A Resolve closure instead of a URL means the host signs its links and they expire, so one is minted per attempt rather than while the item sat in a queue. Segments means the media arrives as an ordered list of parts with no length to range over. External means reaching it needs more than HTTP. Cipher means the bytes arrive encrypted and are decrypted on the way in.

Adding a host is choosing among those shapes; nothing below the extractor has to be told which host it is talking to.

%%{init: {"flowchart": {"wrappingWidth": 420}}}%%
flowchart TD
    paste(["a link — pasted into the UI, or given on the command line"])

    subgraph resolve["extractor — what is behind it"]
        claim{"a registered host claims it?"}
        host["that host's extractor"]
        direct["Direct — the catch-all, after sniffing for a player, a manifest or a directory index"]
        res["Result: one File per download<br>URL · Resolve · Segments · External · Cipher"]
        claim -- yes --> host --> res
        claim -- no --> direct --> res
    end

    queue[("the queue — an item starts once a worker is free, the queue is unpaused, and the host's own pace allows")]

    subgraph move["download — how the bytes arrive"]
        pick{"what does the File carry?"}
        multi["segmented — the remaining span bisected, one connection per range"]
        seq["sequential — a single stream"]
        playlist["playlist — parts fetched several at a time, appended in order"]
        ytdlp["yt-dlp, through the helper script"]
        bucket["one shared token bucket — the rate ceiling, and where a pause parks"]
        part[("a .part file keyed on the job's own URL, with a sidecar recording each range · ciphertext decrypted on the way in")]
        pick -- "URL, length known,<br>server honours ranges" --> multi --> bucket
        pick -- "URL, no length" --> seq --> bucket
        pick -- Segments --> playlist --> bucket
        pick -- External --> ytdlp
        bucket --> part
    end

    fin["renamed into place · any .ts rewrapped to .mp4"]

    paste --> claim
    res --> queue --> pick
    part --> fin
    ytdlp --> fin
Loading

Two details in there are load-bearing rather than incidental. The .part file is named from a hash of the job's own source URL rather than from anything per-run or per-item, which is what lets an interrupted transfer be recognised by the next run instead of restarting from zero — a signed media link would change every time and resume nothing. And the token bucket is genuinely one object shared by every connection, which is why pausing costs nothing and why the code that opens extra connections has to ask whether the ceiling, rather than the host, is what is holding a transfer back.

What may depend on what

Arrows are imports — the layering rather than every edge, since a package may reach anything below it. What matters is that nothing points back up. webui sits outside the stack altogether: it carries the compiled frontend and no logic, so only the entry point touches it.

flowchart TD
    main["cmd/heapleach<br>flags, signals, and which of the two modes to run"]
    main --> server
    main --> cli
    main --> webui

    server["server<br>JSON API · SSE stream · serves the embedded UI"]
    cli["cli<br>the animated terminal display"]
    webui["webui<br>the compiled frontend (go:embed)"]

    server --> download
    cli --> download
    download["download<br>worker pool · resumable transfers · progress"]
    download --> extractor
    download --> proxy
    proxy["proxy<br>egress leases · throughput scoring · Bolt inventory"]
    proxy --> httpx
    extractor["extractor<br>one file per host, and a fallback that sniffs"]
    extractor --> httpx
    httpx["httpx<br>browser-shaped client · redirects · retry and backoff"]
    httpx --> base

    download --> tools
    extractor --> tools
    tools["tools<br>locates yt-dlp, ffmpeg and deno"]
    base["config · util<br>every tunable, and the dependency-free helpers"]
Loading

Both halves of the program drive the same download.Manager: the server subscribes to it and streams whole state snapshots to the browser, and the headless run polls it on a ticker of its own. Nothing about a transfer knows which of the two is watching.

Development

The 2026-10-03 service audit records the hardening changes, validation results and remaining operating limits.

make run          # build the standalone binary, serve locally and open the browser
make run-image    # build and run the container image
make dev          # Go API on :8080 + Vite dev server on :5173 (hot reload)
make dev-backend  # API only
make test         # Go unit tests, with the race detector
make check        # formatting, vet, host inventory, Go, UI, OCR and release tests
make test-live    # include local, gitignored live extractor tests (when present)
make frontend     # build the UI into the Go embed directory
make lock         # regenerate frontend/package-lock.json
make dist         # cross-compile the release archives into ./dist
make tag V=v0.1.0 # manually choose a release version (main pushes bump the patch)
make help         # every target

make build, make run and make image build through Docker. Frontend builds and tests use local npm when available and Docker otherwise, with locked dependencies installed through npm ci. Go tests, vet, host-list generation and release cross-compilation need local Go; Go downloads the toolchain pinned in backend/go.mod automatically. make dev needs both Go and Node locally.

The checked-in tests use synthetic fixtures and local HTTP servers. Live extractor tests belong in gitignored *live_test.go files behind the live build tag and take their source URLs and passwords from environment variables. A fresh clone has no live-site fixtures.

Layout

backend/
  cmd/heapleach/        entry point
  internal/config/  settings and every shared tunable constant
  internal/util/    shared helpers (no dependencies)
  internal/httpx/   HTTP client: browser headers, redirects, retry/backoff
  internal/tools/   locates the optional yt-dlp, ffmpeg and deno
  internal/extractor/  one file per host + a direct-link fallback
  internal/download/   worker pool, resumable transfers, progress
  internal/server/     JSON API, SSE, embedded-asset serving
  internal/cli/        the headless run's terminal display
  internal/webui/dist/ the compiled frontend (go:embed)
frontend/           React + TypeScript (strict), built by Vite

Adding a host

Implement extractor.Extractor (Name, Match, Extract) and register it in NewRegistry. Embed a hostSet naming the domains it claims: that one declaration supplies Match and the inventory entry above both, so the two cannot disagree, and a Match of your own is only for what a host list cannot say. Return a File per download; set Resolve instead of URL when the host issues links that expire.

Releases

Every push to main and pull request runs the tests, go vet and a gofmt check, with the UI compiled first so the binary is built against the real embedded frontend rather than the placeholder. The independent OCR checks run alongside the main checks, and both must pass before a release starts.

After a push to main passes CI, it automatically tags the tested commit with the next patch version and publishes the archives and container images. Versions follow the highest stable vX.Y.Z tag. Publishing is queued, and a rerun reuses its tag or skips a commit already included in a newer release. make tag V=vX.Y.Z remains available for choosing a version explicitly; pushing that tag uses the same publisher.

Automatic releases reuse the caller's tests after verifying the checked-out commit; standalone tags run the Go suite themselves. A separate release cache retains compilation output for all five targets between commits. Container images use the published Linux binaries after checking their checksums and version, so the archives and images carry the same executable and embedded UI.

One Linux runner produces every archive, because the program is pure Go with cgo off and the targets differ only by GOOS and GOARCH:

Archive For
heapleach_<tag>_linux_amd64.tar.gz Linux, x86-64
heapleach_<tag>_linux_arm64.tar.gz Linux, arm64
heapleach_<tag>_darwin_amd64.tar.gz macOS, Intel
heapleach_<tag>_darwin_arm64.tar.gz macOS, Apple silicon
heapleach_<tag>_windows_amd64.zip Windows, x86-64

Each carries the binary, the README and the licence, and SHA256SUMS covers the set. make dist builds exactly the same archives locally, which is the way to check a release before tagging one.

About the name

Heap leaching is a mining method: crushed ore is piled into a heap and irrigated from above, and the solution percolates down through it, dissolving out the metal as it goes and draining to a pad at the bottom. This does the same to a heap of links. The extractors percolate through whatever the pages are made of — players, listings, signed redirects, encrypted payloads — and what is worth keeping drains out into a folder.

Notes

  • The queue is written to HEAPLEACH_STATE every ten minutes when it has changed, and on every clean shutdown, so a restart finds its unfinished jobs held and re-reads them on retry; finished files stay put. A crash can lose queue changes from the last ten minutes: the part files, not the queue file, are what resume a transfer. A .part file carries a .part.state sidecar recording per-connection progress, so an interrupted multi-connection transfer resumes rather than starting over. Both can be deleted safely.
  • These sites change their plumbing without warning. make test-live is the fastest way to find out which extractor broke when the ignored live test file and its environment-provided sources are present. A fresh clone has fixture tests only; a green suite alone does not verify live host support.
  • Be a good citizen: the defaults are deliberately modest, and the client honours Retry-After and backs off on 429s.

Star history

HeapLeach's GitHub stars over time

About

Bulk downloader in one static binary: resolves albums, folders and playlists from 180+ sites (gofile, bunkr, MEGA, pixeldrain…) and downloads them in parallel, with resumable multi-connection transfers, a web UI, a headless CLI and a Docker image.

Topics

Resources

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages