Skip to content
Documentation menu

Documentation

How Scorch works

A small HTTP boundary separates callers from one bounded runtime that finds, fetches, renders, extracts, and follows public web content.

Process boundary

Scorch is a six-crate Rust workspace that produces two executables. The split is architectural, not cosmetic: client operations always cross HTTP, while runtime policy stays inside the service.

Client process

scorch

CLI and MCP adapter. It only speaks HTTP.

Service process

scorchd

Axum API, policy, metasearch, browser runtime, extraction, and crawl jobs.

Public web

Guarded egress

Pinned metasearch and mapping fetches plus isolated browser scrapes through validating egress.

CrateResponsibility
metasearchEngine adapters, concurrent routing, normalization, cache, circuit breakers, and reranking.
scorch-typesShared HTTP request and response contracts.
scorch-engineNetwork policy, validating proxy, fetch, browsers, extraction, search, mapping, and crawl jobs.
scorch-apiAxum routes, strict request parsing, response headers, and HTTP error mapping.
scorch-serverThe scorchd executable, service configuration, logging, and lifecycle.
scorch-cliThe scorch HTTP client and API-backed MCP adapter.

The client does not link scorch-engine, scorch-api, metasearch, or Obscura. This prevents a CLI or MCP caller from accidentally bypassing service policy.

Scrape pipeline

  1. Parse a strict contract. Unknown JSON fields are rejected and option ranges are checked before work starts.
  2. Check extracted-result memory. An identical fresh result returns immediately. Set maxAgeMs to 0 to bypass this lookup.
  3. Join equivalent in-flight work. Concurrent misses for the same browser-affecting options share one bounded render. Each caller keeps its own timeout, extraction formats, and elapsed time.
  4. Validate the target. Scorch accepts only HTTP and HTTPS URLs without embedded credentials, blocks unsafe ports, and resolves the host against local and special-purpose address policy.
  5. Enter a render slot. Bounded admission assigns the request to a long-lived Obscura slot whose outbound connections pass through Scorch's validating proxy.
  6. Navigate with stealth. Obscura loads the page, executes JavaScript, blocks configured media and trackers, and serializes the resulting DOM. This policy is static rather than caller-selectable.
  7. Extract. The resulting HTML becomes requested Markdown, HTML, text, links, and metadata.
  8. Return only requested large fields. The response identifies the final URL, Obscura engine, elapsed time, metadata, warnings, and formats.
  9. Warm other formats after returning. A bounded background queue extracts the remaining formats into an up-to-five-minute, 256-entry, 64 MiB weighted cache. Origin freshness can shorten retention; redirects, cookies, author scripts or inline event handlers, private responses, and restrictive cache directives are excluded.
TRANSPORT / STEALTH

Chrome-like TLS and HTTP identity is always active.

RUNTIME / OBSCURA

Every scrape executes page JavaScript in-process.

POLICY / STATIC

Requests cannot switch rendering or transport modes.

Browser rendering

Obscura is the only renderer. It runs in-process as a Rust library—there is no browser executable, sidecar, CDP server, or subprocess.

Each Obscura request receives a fresh page and V8 state on a persistent render slot. Slots retain connection pools and explicitly public anonymous scripts, while cookies are cleared between renders and document bodies are not shared. Scorch retains network metadata for cache safety but disables Obscura's unused CDP response-body buffers. Browser concurrency is bounded, and one absolute deadline covers validation, queueing, navigation, explicit waiting, and serialization.

Stealth transport is statically enabled. It provides browser-aligned TLS and HTTP behavior, ordered headers, aligned user-agent/client-hint/JavaScript identity, and blocking of common tracker or fingerprinting domains. It does not hide the source IP, provide anonymity, guarantee CAPTCHA bypass, or weaken SSRF controls.

Requests do not select a browser backend, rendering mode, or transport. Every scrape gets a fresh page and an empty cookie jar, with stealth transport always enabled. Renders are dispatched onto a fixed set of long-lived render slots, so two renders on the same slot reuse its connections and TLS sessions; no page, DOM, or cookie crosses between them.

Why Obscura replaced the older renderer

Before removing the legacy renderer, both paths were measured in the same optimized build on the same 16-thread Linux host. Three alternating-order trials covered four public pages, with four sequential requests followed by eight requests at concurrency four.

MeasurementObscura stealthLegacy ChromiumDifference
First warm render0.266 s0.661 s2.5× faster startup
PSS after warm-up38.3 MiB407.9 MiB10.7× less memory
Parallel peak PSS90.6 MiB467.7 MiB5.2× less memory
Parallel throughput1.41 req/s3.03 req/sLegacy renderer was 2.2× faster
Peak process count114Obscura stayed in-process

The legacy path delivered higher steady-state throughput, but its warm start was slower and its runtime substantially larger. Scorch prioritizes the smaller self-contained renderer for local Pi agent workflows and predictable deployment. That throughput row predates the render slot pool, which was added afterwards and raised browser throughput substantially; the figures below were collected after it landed.

Measured footprint

Unreleased Scorch after the render slot pool was compared with Firecrawl 2.11.0 at commit ef12eb36. Both columns were collected together on August 14, 2026, on one 16-thread Linux host. Firecrawl ran its pinned, unmodified full Docker Compose configuration; Scorch ran as one process at maximum concurrency four to match it.

Measured resultFirecrawlScorch
Long-running deployment units6 containers1 service
Sequential scrape latency, median0.892 s0.638 s
Throughput at concurrency four3.61 req/s5.56 req/s
Warm-idle cgroup memory3,128 MiB38 MiB
Peak cgroup memory3,329 MiB75 MiB
OS processes, warm / peak49 / 551 / 1

Both products scraped all 24 parallel requests successfully. This is an end-to-end browser-rendered Markdown scrape and deployment-footprint microbenchmark over four public pages, not a feature-parity, extraction-quality, crawl-speed, or maximum-throughput comparison. The benchmark was collected twice and these are the figures least favorable to Scorch. Memory is cgroup v2 working set sampled every 50 ms, summed across Firecrawl's running containers. Full methodology, exclusions, and host details are in the project README.

Metasearch

Metasearch is Scorch's only search provider. The server allowlist remains the policy boundary; API, CLI, and MCP requests may narrow it to an explicit subset or omit selection to use DuckDuckGo. The server may allow the 19 live-validated credential-free engines, plus credential-backed official Brave or Google adapters. Brave Web parses Brave Search's public HTML; Google CSE uses Blackle's public Programmable Search Element. Both are request-explicit, credential-free, and best-effort. Search categories are typed query restrictions: the GitHub category adds site:github.com before dispatch and labels matching results.

  1. Start every selected, allowed, and healthy engine concurrently under per-engine limits.
  2. Accept the first useful result, then hold a short collection window for additional agreement.
  3. Cancel unfinished work rather than letting one slow engine determine total latency.
  4. Normalize URLs by removing fragments and common tracking parameters.
  5. Merge duplicates and rank with weighted reciprocal-rank fusion.
  6. Report contributing sources and completed-engine failures in warnings.

The runtime maintains a bounded 60-second in-memory result cache and temporary circuit breakers after repeated failures. Search is best-effort: provider markup changes, rate limits, and challenges are treated as failures rather than authoritative empty results.

Mapping and crawl

Mapping is sitemap-first. Scorch checks robots-declared sitemaps, the conventional sitemap URL, bounded nested sitemap indexes, and finally links on the root page. URLs are normalized, deduplicated, path-filtered, and restricted to the requested site.

Crawling starts from mapped URLs, follows same-origin links breadth-first, applies include and exclude paths, and respects robots.txt as ScorchBot. A work-conserving bounded frontier starts the next eligible URL whenever a page completes, and jobs preserve individual page failures beside successful documents.

  • Maximum 100 pages and discovery depth 5 per crawl.
  • Maximum four active crawl jobs.
  • Five-minute absolute crawl deadline.
  • 32 MiB retained data per job and 128 retained jobs.
  • Paginated status responses, cancellation, and configurable completed-job TTL.

Security boundary

Pages, search results, browser output, sitemaps, robots files, and extracted links are untrusted data. They cannot change process configuration or trigger command execution.

Direct path

Manual redirects, DNS validation and pinning, HTTP(S)-only URLs, unsafe-port rejection, and response caps.

Browser path

The same target checks plus an embedded HTTP/CONNECT proxy that validates browser redirects and subresources.

Obscura receives only the validating proxy endpoint and retains its own private-address guard. Browser redirects and subresources remain inside the same proxy-enforced network boundary.

Authentication is outside this boundary. The local-first API has no built-in authentication or TLS. Use loopback or an authenticated TLS gateway when serving other machines.

Runtime state

scorchd owns one Axum server, one globally bounded direct-fetch runtime, bounded isolated Obscura work, one validating proxy, and a bounded in-memory crawl registry.

There is no durable state. A restart clears crawl jobs, extracted scrape and metasearch cache entries, in-flight render coordination, and circuit-breaker state. This keeps deployment self-contained but means clients must persist any results they need to retain.