Documentation menu
Documentation
HTTP API
A small JSON API for liveness, readiness, scraping, metasearch, URL discovery, and bounded crawl jobs.
Connection
The default base URL is http://127.0.0.1:33000. All request
bodies use JSON with camel-case property names and content-type: application/json.
curl -sS http://127.0.0.1:33000/health
Request objects reject unknown fields. Every response includes an x-request-id header suitable for correlation with service logs.
Health and readiness
GET /health
Process liveness. A healthy process returns HTTP 200:
{ "status": "ok" } GET /ready
Reports runtime policy and Obscura availability. It returns HTTP 200 when the embedded renderer is available, or 503 when unavailable.
{
"status": "ready",
"browserAvailable": true,
"browser": "obscura",
"obscuraStealth": true,
"maxConcurrency": 4,
"searchProvider": "metasearch",
"searchEngines": ["bing", "brave-web", "crates-io", "crossref", "docker-hub", "duckduckgo", "github", "google-cse", "hacker-news", "hugging-face", "mwmbl", "npm", "nvd", "openalex", "open-library", "pubmed", "wikidata", "wikipedia", "yahoo"]
} Scrape
POST /v1/scrape
{
"url": "https://example.com",
"options": {
"formats": ["markdown", "html", "text", "links", "metadata"],
"timeoutMs": 30000,
"waitForMs": 0,
"onlyMainContent": true,
"blockMedia": true,
"maxAgeMs": 300000,
"storeInCache": true
}
} | Field | Default | Notes |
|---|---|---|
url | required | Public HTTP or HTTPS URL without embedded credentials. |
formats | ["markdown"] | markdown, html, text, links, and metadata. |
timeoutMs | 30000 | Absolute scrape budget; 100–120,000 ms. |
waitForMs | 0 | Explicit post-load wait; capped at 60,000 ms. |
onlyMainContent | true | Focus readable extraction on the main page content. |
blockMedia | true | Suppress browser media requests where supported. |
maxAgeMs | 300000 | Accept an eligible identical extracted result from the bounded
in-memory cache up to this age. Set to 0 to force a fresh
scrape and refresh the cache. Only same-URL successful navigations without
author scripts or inline event handlers and with safe freshness, cookie,
redirect, and Vary provenance are retained. |
storeInCache | true | Read and populate the extracted-result cache. Set to false to bypass both cache reads and writes. |
The response engine is always obscura.
{
"url": "https://example.com",
"finalUrl": "https://example.com/",
"engine": "obscura",
"elapsedMs": 143,
"metadata": {
"title": "Example Domain",
"statusCode": 200,
"contentType": "text/html"
},
"markdown": "# Example Domain\n\n...",
"links": [],
"warnings": []
} Search
POST /v1/search
{
"query": "rust async runtime",
"limit": 5,
"country": "us",
"language": "en",
"categories": ["github"],
"scrapeOptions": null
}
Limits range from 1 to 20 and queries are capped at 512 bytes. Metasearch is
the only public provider. Omit engines to use DuckDuckGo, or select
a subset of the server allowlist. The github category restricts the
upstream web query to site:github.com. Select brave-web
for credential-free Brave website results or google-cse for credential-free
Google-derived Programmable Search results. Both public frontend integrations
are best-effort. Results identify contributing sources
and matching categories.
{
"query": "rust async runtime",
"provider": "metasearch",
"engines": ["duckduckgo"],
"results": [
{
"position": 1,
"title": "Repository title",
"url": "https://github.com/example/repository",
"description": "Result description",
"sources": ["duckduckgo"],
"category": "github"
}
],
"elapsedMs": 911,
"warnings": []
}
Set scrapeOptions to a normal scrape-options object to enrich each
result. Per-result enrichment errors are retained in that result's error field without discarding successful results.
Map
POST /v1/map
{
"url": "https://example.com",
"limit": 100,
"includeSubdomains": false,
"includePaths": [],
"excludePaths": []
} Limits range from 1 to 1,000. Scorch checks robots-declared and conventional sitemaps, bounded nested sitemap indexes, and root-page links. A successful map may contain no links.
{
"url": "https://example.com",
"links": ["https://example.com/", "https://example.com/docs"],
"elapsedMs": 287,
"sources": ["sitemap"]
} Crawl
POST /v1/crawls
Starts an in-memory crawl and returns HTTP 202 with a job summary.
{
"url": "https://example.com",
"limit": 20,
"maxDepth": 2,
"concurrency": 4,
"includePaths": [],
"excludePaths": [],
"scrapeOptions": {
"formats": ["markdown"]
}
} Configured defaults and ceilings:
- Page limit 20 by default, maximum 100.
- Depth 2 by default, maximum 5.
- Concurrency 4 by default and no higher than global concurrency.
- Five-minute absolute job deadline and 32 MiB retained data per job.
- 128 retained jobs, four active crawls, and 15-minute completed-job TTL by default.
GET /v1/crawls/{id}
GET /v1/crawls/550e8400-e29b-41d4-a716-446655440000?cursor=0&pageSize=10
Page sizes range from 1 to 50. Status values are queued, running, completed, cancelled, and failed.
Follow nextCursor until it is absent.
DELETE /v1/crawls/{id}
Cancels active work and removes the job. Success returns:
{ "id": "550e8400-e29b-41d4-a716-446655440000", "deleted": true } Errors
{
"code": "invalid_request",
"message": "invalid request: ..."
} | Status | Category |
|---|---|
400 | Invalid request or unsafe target URL. |
404 | Unknown route or crawl job. |
408 | Outer HTTP request timeout. |
413 | Remote response exceeded the configured byte limit. |
415 | Unsupported remote content. |
429 | Runtime or job capacity exhausted. |
502 | DNS, fetch, extraction, or search upstream failure. |
503 | Browser unavailable or browser operation failed. |
504 | Engine deadline exceeded. |
Security behavior
All target-bearing endpoints accept only HTTP and HTTPS URLs without embedded credentials. Scorch rejects unsafe ports and hostnames that resolve to local, private, link-local, reserved, or other special-purpose addresses.
Direct redirects are manual, limited to five, and independently revalidated. Browser connections route through an embedded validating HTTP/CONNECT proxy. Response bytes, request bodies, queueing, rendering, crawl growth, retained job data, and job lifetimes are bounded.