Documentation menu
Documentation
How Scorch works
A small HTTP boundary separates callers from one bounded runtime that finds, fetches, renders, extracts, and follows public web content.
Process boundary
Scorch is a six-crate Rust workspace that produces two executables. The split is architectural, not cosmetic: client operations always cross HTTP, while runtime policy stays inside the service.
scorch
CLI and MCP adapter. It only speaks HTTP.
scorchd
Axum API, policy, metasearch, browser runtime, extraction, and crawl jobs.
Guarded egress
Pinned metasearch and mapping fetches plus isolated browser scrapes through validating egress.
| Crate | Responsibility |
|---|---|
metasearch | Engine adapters, concurrent routing, normalization, cache, circuit breakers, and reranking. |
scorch-types | Shared HTTP request and response contracts. |
scorch-engine | Network policy, validating proxy, fetch, browsers, extraction, search, mapping, and crawl jobs. |
scorch-api | Axum routes, strict request parsing, response headers, and HTTP error mapping. |
scorch-server | The scorchd executable, service configuration, logging,
and lifecycle. |
scorch-cli | The scorch HTTP client and API-backed MCP adapter. |
The client does not link scorch-engine, scorch-api, metasearch, or Obscura. This prevents a CLI or MCP caller from
accidentally bypassing service policy.
Scrape pipeline
- Parse a strict contract. Unknown JSON fields are rejected and option ranges are checked before work starts.
- Check extracted-result memory. An identical fresh result returns
immediately. Set
maxAgeMsto0to bypass this lookup. - Join equivalent in-flight work. Concurrent misses for the same browser-affecting options share one bounded render. Each caller keeps its own timeout, extraction formats, and elapsed time.
- Validate the target. Scorch accepts only HTTP and HTTPS URLs without embedded credentials, blocks unsafe ports, and resolves the host against local and special-purpose address policy.
- Enter a render slot. Bounded admission assigns the request to a long-lived Obscura slot whose outbound connections pass through Scorch's validating proxy.
- Navigate with stealth. Obscura loads the page, executes JavaScript, blocks configured media and trackers, and serializes the resulting DOM. This policy is static rather than caller-selectable.
- Extract. The resulting HTML becomes requested Markdown, HTML, text, links, and metadata.
- Return only requested large fields. The response identifies the final URL, Obscura engine, elapsed time, metadata, warnings, and formats.
- Warm other formats after returning. A bounded background queue extracts the remaining formats into an up-to-five-minute, 256-entry, 64 MiB weighted cache. Origin freshness can shorten retention; redirects, cookies, author scripts or inline event handlers, private responses, and restrictive cache directives are excluded.
Chrome-like TLS and HTTP identity is always active.
Every scrape executes page JavaScript in-process.
Requests cannot switch rendering or transport modes.
Browser rendering
Obscura is the only renderer. It runs in-process as a Rust library—there is no browser executable, sidecar, CDP server, or subprocess.
Each Obscura request receives a fresh page and V8 state on a persistent render slot. Slots retain connection pools and explicitly public anonymous scripts, while cookies are cleared between renders and document bodies are not shared. Scorch retains network metadata for cache safety but disables Obscura's unused CDP response-body buffers. Browser concurrency is bounded, and one absolute deadline covers validation, queueing, navigation, explicit waiting, and serialization.
Stealth transport is statically enabled. It provides browser-aligned TLS and HTTP behavior, ordered headers, aligned user-agent/client-hint/JavaScript identity, and blocking of common tracker or fingerprinting domains. It does not hide the source IP, provide anonymity, guarantee CAPTCHA bypass, or weaken SSRF controls.
Why Obscura replaced the older renderer
Before removing the legacy renderer, both paths were measured in the same optimized build on the same 16-thread Linux host. Three alternating-order trials covered four public pages, with four sequential requests followed by eight requests at concurrency four.
| Measurement | Obscura stealth | Legacy Chromium | Difference |
|---|---|---|---|
| First warm render | 0.266 s | 0.661 s | 2.5× faster startup |
| PSS after warm-up | 38.3 MiB | 407.9 MiB | 10.7× less memory |
| Parallel peak PSS | 90.6 MiB | 467.7 MiB | 5.2× less memory |
| Parallel throughput | 1.41 req/s | 3.03 req/s | Legacy renderer was 2.2× faster |
| Peak process count | 1 | 14 | Obscura stayed in-process |
The legacy path delivered higher steady-state throughput, but its warm start was slower and its runtime substantially larger. Scorch prioritizes the smaller self-contained renderer for local Pi agent workflows and predictable deployment. That throughput row predates the render slot pool, which was added afterwards and raised browser throughput substantially; the figures below were collected after it landed.
Measured footprint
Unreleased Scorch after the render slot pool was compared with Firecrawl 2.11.0 at commit ef12eb36. Both columns were collected together on August 14, 2026, on one 16-thread
Linux host. Firecrawl ran its pinned, unmodified full Docker Compose
configuration; Scorch ran as one process at maximum concurrency four to
match it.
| Measured result | Firecrawl | Scorch |
|---|---|---|
| Long-running deployment units | 6 containers | 1 service |
| Sequential scrape latency, median | 0.892 s | 0.638 s |
| Throughput at concurrency four | 3.61 req/s | 5.56 req/s |
| Warm-idle cgroup memory | 3,128 MiB | 38 MiB |
| Peak cgroup memory | 3,329 MiB | 75 MiB |
| OS processes, warm / peak | 49 / 55 | 1 / 1 |
Both products scraped all 24 parallel requests successfully. This is an end-to-end browser-rendered Markdown scrape and deployment-footprint microbenchmark over four public pages, not a feature-parity, extraction-quality, crawl-speed, or maximum-throughput comparison. The benchmark was collected twice and these are the figures least favorable to Scorch. Memory is cgroup v2 working set sampled every 50 ms, summed across Firecrawl's running containers. Full methodology, exclusions, and host details are in the project README.
Metasearch
Metasearch is Scorch's only search provider. The server allowlist remains
the policy boundary; API, CLI, and MCP requests may narrow it to an explicit
subset or omit selection to use DuckDuckGo. The server may allow the 19
live-validated credential-free engines, plus credential-backed official
Brave or Google adapters. Brave Web parses Brave Search's public HTML;
Google CSE uses Blackle's public Programmable Search Element. Both are
request-explicit, credential-free, and best-effort. Search categories are
typed query restrictions: the GitHub category adds site:github.com before dispatch and labels matching results.
- Start every selected, allowed, and healthy engine concurrently under per-engine limits.
- Accept the first useful result, then hold a short collection window for additional agreement.
- Cancel unfinished work rather than letting one slow engine determine total latency.
- Normalize URLs by removing fragments and common tracking parameters.
- Merge duplicates and rank with weighted reciprocal-rank fusion.
- Report contributing sources and completed-engine failures in warnings.
The runtime maintains a bounded 60-second in-memory result cache and temporary circuit breakers after repeated failures. Search is best-effort: provider markup changes, rate limits, and challenges are treated as failures rather than authoritative empty results.
Mapping and crawl
Mapping is sitemap-first. Scorch checks robots-declared sitemaps, the conventional sitemap URL, bounded nested sitemap indexes, and finally links on the root page. URLs are normalized, deduplicated, path-filtered, and restricted to the requested site.
Crawling starts from mapped URLs, follows same-origin links breadth-first,
applies include and exclude paths, and respects robots.txt as ScorchBot. A work-conserving bounded frontier starts the next eligible URL whenever
a page completes, and jobs preserve individual page failures beside
successful documents.
- Maximum 100 pages and discovery depth 5 per crawl.
- Maximum four active crawl jobs.
- Five-minute absolute crawl deadline.
- 32 MiB retained data per job and 128 retained jobs.
- Paginated status responses, cancellation, and configurable completed-job TTL.
Security boundary
Pages, search results, browser output, sitemaps, robots files, and extracted links are untrusted data. They cannot change process configuration or trigger command execution.
Manual redirects, DNS validation and pinning, HTTP(S)-only URLs, unsafe-port rejection, and response caps.
The same target checks plus an embedded HTTP/CONNECT proxy that validates browser redirects and subresources.
Obscura receives only the validating proxy endpoint and retains its own private-address guard. Browser redirects and subresources remain inside the same proxy-enforced network boundary.
Runtime state
scorchd owns one Axum server, one globally bounded direct-fetch runtime,
bounded isolated Obscura work, one validating proxy, and a bounded in-memory crawl
registry.
There is no durable state. A restart clears crawl jobs, extracted scrape and metasearch cache entries, in-flight render coordination, and circuit-breaker state. This keeps deployment self-contained but means clients must persist any results they need to retain.