Skip to content
Scorch

Search and extract the public web at Rust speed

Scorch gives applications one guarded service for search, scraping, mapping, and crawling—without a database, broker, browser farm, or worker fleet.

API · 127.0.0.1:33000   /   Renderer · Obscura

2 purpose-built binaries
0 required runtime services
6 focused Rust crates
MIT open-source license

One guarded runtime

Total visibility from source to document

One deliberate surface

A single system that finds, reads, and follows the web

Runtime and routing policy remain on the server. API, CLI, and MCP callers get a compact contract without gaining access to forbidden engines or backends.

Unified signal, less noise

Query a server-bounded engine subset concurrently, normalize and merge URLs, then rerank agreement while policy remains operator-controlled.

Search · reciprocal-rank fusion →

Clean content on demand

Render every page through isolated embedded Obscura with JavaScript execution and Chrome-like stealth transport statically enabled.

Extract · Markdown · text →

Bounded discovery

Discover sitemaps and same-site links, then run robots-aware crawl jobs with pagination, cancellation, retained-byte limits, and expiry.

Map · breadth-first crawl →
Purple Scorch pipeline transforming web signals into clean documents

Why self-contained

A complete web pipeline gives applications an edge, but stitching services together creates tradeoffs

The deployment gap
Every queue, browser sidecar, cache, and worker adds another service to secure, observe, and recover.
The policy gap
Separate fetch and browser paths drift unless they share URL validation, DNS controls, redirects, limits, and deadlines.

Measured, not asserted

Against Firecrawl on one host, the same scrape costs far less

Both products rendered the same four public pages in a browser and returned Markdown, collected side by side on August 14, 2026. Every request succeeded on both.

1.5× higher scrape throughput 5.56 vs 3.61 req/s at concurrency four
82× less warm-idle memory 38 MiB vs 3,128 MiB resident
1 vs 6 long-running units One service, no queue or browser sidecar
1 vs 55 processes at peak The renderer stays in-process

A deployment-footprint and browser-scrape microbenchmark, not a feature-parity or extraction-quality comparison. Where two collections disagreed these are the figures least favourable to Scorch. See the full comparison and method.

Operationally compact

Clear boundaries from the first request

The client only speaks HTTP. The service owns execution, isolation, policy, and ephemeral crawl state.

2

Strictly separated executables

CLIENT / SERVICE
1

Shared guarded egress policy

DIRECT / BROWSER
0

Durable services required

SELF / CONTAINED

From one release build to a local API, Scorch keeps the path short without hiding the security boundary.

Enter the reproducible environment, build two optimized binaries, start scorchd, and connect through the CLI, JSON API, or MCP.

Read the quickstart

Frequently asked questions

The constraints that matter before putting Scorch to work.

Does Scorch need a database or queue?+

No. Crawl jobs live in bounded memory and expire. Restarting scorchd removes them, so Scorch is not a durable crawl archive.

Can callers choose search engines or a browser?+

The server controls the engine allowlist; callers may select an allowed subset, while omitted selection uses DuckDuckGo. Every page scrape uses embedded Obscura with stealth statically enabled; callers cannot switch rendering or transport modes.

Is authentication built in?+

Not in the local-first release. Keep the loopback bind, or place an authenticated TLS gateway in front of the API.

Does Obscura run as a sidecar?+

No. It is embedded in scorchd as a Rust library, with fresh context, page, and JavaScript state per request.

Bring the public web closer to the systems that need it

Build Scorch, start scorchd, and make the first request through a runtime you control.