Developer documentation
Web data, one API.
Send a URL. Get HTML, text, Markdown, JSON, links, images and structured fields in one result. Execution is automatic.
1. Get a price
Use your deployment URL and API key. This shell example uses curl and jq.
Shell · quote and approve the current price
export SCRAPER_URL="https://YOUR_SCRAPER_HOST"
export SCRAPER_API_KEY="sg_YOUR_KEY"
request='{"target_url":"https://example.com/","idempotency_key":"example-001"}'
quote=$(curl --fail-with-body -sS "$SCRAPER_URL/api/v1/pricing/quote" \
-H "Authorization: Bearer $SCRAPER_API_KEY" \
-H "Content-Type: application/json" -d "$request")
echo "$quote" | jq .
request=$(echo "$request" | jq --argjson cost "$(echo "$quote" | jq '.cost')" '. + {max_cost:$cost}')
2. Create the job
Shell · keep this response and idempotency key
job=$(curl --fail-with-body -sS "$SCRAPER_URL/api/v1/jobs" \
-H "Authorization: Bearer $SCRAPER_API_KEY" \
-H "Content-Type: application/json" -d "$request")
job_id=$(echo "$job" | jq -r '.job_id')
echo "$job" | jq .
3. Read the result
Shell · result endpoint
curl -i "$SCRAPER_URL/api/v1/jobs/$job_id/result" \
-H "Authorization: Bearer $SCRAPER_API_KEY"
202 means still working: wait for the seconds in Retry-After. 200 returns all data. Stop polling on 422 (failed or canceled) or 410 (expired).
Authentication
Create a key in API Keys. Send it in the authorization header on customer API requests.
HTTP header
Authorization: Bearer sg_YOUR_KEY
Keep keys on your server or in your integration's credential store. Access follows the key's permissions. Documentation, OpenAPI and the n8n template are public.
Pricing & retries
Auto jobs charge for success. Failed and canceled Auto jobs are free; internal retries are included. Screenshot and XHR are optional paid extras. A job that sets request_payload.headers is billed per_request: charged once it has run, even when it fails.
Methods, headers and bodies
Requests use GET unless request_payload.method is POST. request_payload.body carries up to 64 KiB of UTF-8 text with POST; without a Content-Type header, valid JSON is sent as application/json and anything else as application/x-www-form-urlencoded. A POST runs like a GET: in a browser it is sent as a same-origin form submission to the target, so screenshots, XHR capture, saved logins and page actions work with it. In a browser a POST needs an https:// URL: browsers upgrade plain-http page loads, and the job fails without sending rather than send a GET in its place. PDF extraction and sitemap discovery need GET. A POST is sent at most once: it is retried only when the request provably never reached the target, so a timeout after sending fails the job rather than repeating it.
request_payload.headers takes up to 32 headers; the Content-Type a body gets by default is added on top. Names are unique regardless of case, and values lose surrounding spaces and tabs. The platform manages Host, Content-Length, Transfer-Encoding, TE, Trailer, Connection, Keep-Alive, Upgrade, Proxy-Connection, Expect, Accept-Encoding, User-Agent, the headers a browser sends on every page (Accept, Accept-Language, Cache-Control, Pragma, Priority, Upgrade-Insecure-Requests, DNT, Origin, X-Requested-With), client-IP headers (X-Forwarded-For, X-Forwarded-Host, Forwarded, Via, X-Real-IP, True-Client-IP, CF-Connecting-IP) and every header starting with Proxy-, X-Proxy- or Sec-. A job with a saved login cannot also set Cookie or Accept-Language. When a site's configuration fixes the method or a header, a request that would change it is rejected rather than silently rewritten. Headers and bodies are stored with the job as sent, so avoid long-lived secrets; job lists omit the body and GET /jobs/{jobID} returns it.
| Option | Behavior |
|---|
max_cost | Approved price in credits. Use the quote's cost. A higher current price returns 409 before admission. |
proxy_policy.country | Optional country from the quote's countries. Omit for Auto. Only offer a country selector when at least two are available. |
idempotency_key | Keep the same key and request when retrying a lost response. Use a new key for a different job. |
For a batch, quote the same items you submit to POST /job-batches (up to 100 URLs). A successful batch returns 201 with batch, items and pagination; created:true marks a new batch. Resending the same items with the same keys in the same order returns that batch with 200 and created:false. Refused items return 422 with error, the first refusal, and errors: one {item_no, field, code, message} per item; already_exists means the key is already used; to read a batch back, resend all of its items unchanged. The whole request must stay under 1 MiB; a larger one returns 413. Use bounded retries for 429 and temporary 5xx, honoring Retry-After.
Set sticky:true on a batch and on its quote to send every item through the same IP. Sticky jobs run one at a time at the site's pacing, so the batch takes longer, and the quote includes the sticky supplement. Items using a saved login with a sticky IP keep that login's IP. A task opts in with definition.sticky:true; quote its pages as items with sticky:true.
Content & fields
All result formats are always included. Use request_payload.content to select the content you need.
Job request · optional content settings
{
"target_url": "https://example.com/",
"idempotency_key": "example-content-001",
"request_payload": {
"content": {
"main": true,
"exclude": ["nav", ".advertisement"],
"fields": {
"title": {"selector": "h1", "type": "string", "required": true},
"image": {"selector": "img", "attribute": "src", "all": true}
}
}
}
}
| Setting | Effect |
|---|
main | Keep the main content. Falls back to the page with a warning if detection is unavailable. |
include / exclude | Arrays of CSS selectors. Keep matching elements or remove them. |
fields | Sources: css (selector and attribute), page (title, description, canonical, language), jsonld or json (property path such as offers.price), or ai (plain-language description). Set all:true for an array. Valid JSON-LD is also returned in data.structured_data. |
type | string, number or boolean. Missing or invalid required fields fail the job. |
Fields are extracted after include/exclude filtering and before main-content selection. CSS extraction requires HTML and does not use AI inference.
PDF & OCR
For a PDF URL, set request_payload.document:{"ocr":false}. The first 10 pages, up to 4 MiB, are included in the quoted scrape price. Set ocr:true to recognize scanned pages in English and Russian: the quote includes one OCR supplement covering up to 10 pages. HTML is a text representation of the PDF; all text formats and field extraction remain available. Screenshots, XHR, saved sessions, page actions, sitemap discovery and POST are not supported for PDF jobs.
result.document reports pages, total_pages, ocr and truncated. More than 10 pages produces document_page_limit_reached; oversized files, encrypted or unreadable documents fail. OCR is bounded by the job deadline and does not guarantee layout/table reconstruction. Without OCR, a fully scanned document fails with guidance to enable OCR.
Saved logins & page actions
In the panel, open Saved logins, choose Sign in to website, enter its login page URL, sign in (including 2FA), then choose Save login. In a task, enable Use saved login and select it. You can also sign in directly from the task form. Without this option, tasks start a fresh session and do not save login state. The browser closes after saving; stored profiles do not reserve a worker.
API: create a profile with POST /sessions: {"name":"Store account","origin":"https://example.com","login_url":"https://example.com/login","retention_days":30,"idempotency_key":"profile-1"}. Pass its session_id in request_payload for subsequent jobs. All profiles use Firefox. Only metadata is returned by GET /sessions; cookies and tokens cannot be downloaded through this API.
Retention is 30 or 90 days after use, or 0 until deletion. Website authorization can expire sooner. Cookies, localStorage and IndexedDB are saved; enable session_storage:true for temporary tab data. Storage is bounded (512 cookies, 256 local entries, 10,000 database records and 2 MiB of database/tab state); unsupported storage values fail explicitly. This is portable authorization state, not a complete disk profile. Optionally set auth_selector to an element visible only after sign-in. A missing marker or HTTP 401 sets status:needs_login and fails the task with instructions to sign in again.
POST /sessions/{sessionID}/login opens an owned interactive browser, or resumes the one still open; a sign-in nobody polled for 90 seconds is closed and replaced. GET /sessions names the open one in sign_in_job_id. Poll GET /profile-logins/{jobID}: expires_at says when the browser closes, and passing the last frame.hash as seen answers frame.same:true with an empty image while the picture is unchanged. Send UUID-identified commands to POST /profile-logins/{jobID}/input; its reply carries only status. Commands: click, text, key, scroll, save, cancel. Input order is preserved. The browser closes after 10 minutes, or 90 seconds without panel activity. Input and frames are encrypted temporary data, never job artifacts.
Jobs using one profile run sequentially and update its saved state. Profile country stays fixed. With sticky_ip:true, the IP pin lasts one hour and the quoted sticky supplement applies. An expired pin or a detected IP change stops the job. POST /sessions/{sessionID}/renew-ip explicitly starts a new pin while retaining authorization. Residential IP availability remains provider-dependent; an IP check cannot guarantee an address forever.
POST /sessions/{sessionID}/revoke deletes stored login data. Known saved credential values and sensitive HTTP headers are redacted from profile results and XHR. Do not use page actions to enter passwords: action values are part of the saved job request; use the interactive login window.
Use request_payload.actions for up to 10 ordered steps within 10 seconds total: {"type":"fill","selector":"#search","value":"book"}, {"type":"click","selector":"#search-button"}, {"type":"scroll","pixels":800}, {"type":"wait_for","selector":".results"}. A click/fill selector must match exactly one visible element. Values are stored in the job request. Actions run before extraction; failure fails the job. Scripts are not accepted; actions are restricted to the initial origin.
Results & XHR
GET /jobs/{jobID}/result returns a single response with every representation.
200 · successful result (abbreviated)
{
"job_id": "...",
"status": "succeeded",
"data": {
"html": "<main>...</main>", "text": "...", "markdown": "# ...",
"json": {"text": "...", "fields": {}, "links": [], "images": []},
"fields": {"title": "Example"}, "links": [], "images": []
},
"metadata": {"url": "https://example.com/", "status_code": 200},
"xhr": [], "warnings": [], "billing": {},
"url": "/api/v1/jobs/.../artifact"
}
To test fields before a full run, use Content → Extract fields → Preview 1 URL in the panel. It quotes and creates one normal job; refresh reuses its result. In the API or SDK, submit one URL with your content options and read data.fields. For a JSON source, data.json contains the parsed original JSON. Download the original artifact at GET /jobs/{jobID}/artifact.
Screenshot and network capture
Set capture_screenshot:true or capture_xhr:true inside request_payload, and include them in the price quote. Download the screenshot from the artifact endpoint with ?view=screenshot.
| XHR field | Meaning |
|---|
response_body | Captured response body. |
response_body_encoding | utf8 text or base64 binary. May be absent on legacy text captures. |
response_body_truncated | The captured body was cut short. |
response_body_error | The body could not be fully collected, for example because it exceeded the capture limit. |
Firefox captures response bodies as well as metadata. Capture has size and timing limits; check availability and truncation fields. It is not a complete network archive.
Crawl a website
Start a bounded run with POST /task-runs. Each page creates a regular job and uses the same content and extra settings.
POST /api/v1/task-runs
{
"idempotency_key": "catalog-001",
"definition": {
"urls": ["https://example.com/"], "crawl": true,
"max_pages": 100, "max_depth": 3,
"exclude_paths": ["/login", "/account/*"],
"request": {
"max_cost": 3,
"request_payload": {"content": {"main": true}}
}
}
}
Replace the example max_cost:3 with your approved per-page quote. Maximum cost per run is max_pages × max_cost; only successful pages are charged.
- Discovery follows HTML links on the same origins as the seed URLs, deduplicates URLs and removes fragments.
include_paths and exclude_paths use shell globs: * matches within a path segment.- Limits: 100 seed URLs, 1,000 pages, depth 10. Use
crawl:false for a URL list: every URL runs once, max_depth and sitemap are refused, and max_pages may be omitted. - Crawls send GET requests, because discovered pages reuse the request.
Read GET /task-runs/{runID} for pages and job IDs. Terminal states are completed, completed_with_errors and canceled. Fetch page results through the regular result endpoint. Cancel with POST /task-runs/{runID}/cancel.
Sitemap discovery
For sitemap discovery, POST /maps with a regular GET job request and poll GET /maps/{jobID}. Quote first with request_payload.discover_sitemap:true; a map uses the normal successful-page price. The executor reads same-origin robots.txt, sitemap.xml and nested sitemap indexes, bounded to 10 documents, 1 MiB per document, 1,000 URLs and 10 seconds. Cross-origin sitemap URLs and redirects are excluded. Read errors and truncated; this is a bounded preview, not a guarantee of every URL. Repeat include_path/exclude_path query parameters to filter without another scrape. A crawl definition with "sitemap":true discovers sitemap URLs on seed pages and applies its existing page and path limits.
Change monitoring
Saved tasks can set monitor:true and optional monitor_fields:["price"]. Text is compared when no fields are selected. GET /task-runs/{runID}/changes returns paginated before/after observations with offset and limit (default 20, maximum 50). before_text and after_text preserve exact display values, including integers beyond JavaScript precision. First observations and unchanged pages do not notify; changed pages emit signed page.changed webhooks. Missing results and comparisons exceeding 128 KiB are reported as errors. The last 20 versions per URL are retained, plus versions with unfinished webhook deliveries. Changing the extraction schema starts a new baseline.
Saved tasks & schedules
Save a reusable definition with POST /saved-tasks, or use Automations in the panel.
Saved task envelope · reuse your definition from the crawl example
{
"idempotency_key": "daily-catalog-001",
"name": "Daily catalog",
"definition": {"urls": ["https://example.com/"], "request": {"max_cost": 3}},
"schedule": "0 9 * * *",
"timezone": "Europe/Moscow",
"enabled": true
}
Use a real quoted per-page ceiling. Schedules use five-field cron and an IANA timezone (UTC by default). The example runs at 09:00 Moscow time daily.
- Each occurrence authorizes a new run with the saved price ceiling. A price or balance change stops further page admissions.
- Overlapping runs are skipped; missed occurrences are coalesced into one.
- Run now with
POST /saved-tasks/{taskID}/runs and a new idempotency_key. - Send the updated saved task to
PUT /saved-tasks/{taskID} with enabled:false to pause. Cancel an active run separately. DELETE /saved-tasks/{taskID} deletes the task and its schedule. Its past runs stay, no longer linked to it; its change-monitoring history is deleted.
Webhooks & delivery
Set callback_url on a job or on a saved task's definition.request. We POST a completion event to your endpoint.
Verify the signature
Retrieve your private secret with GET /webhooks/signing-key. Verify the exact raw request body before parsing JSON.
Signature contract
signed_payload = timestamp + "." + raw_body
expected = "sha256=" + HMAC_SHA256_HEX(secret, signed_payload)
X-Scraper-Timestamp: Unix timestamp in seconds
X-Scraper-Signature: sha256=...
X-Scraper-Event-ID: stable delivery event ID
Use constant-time comparison, allow at most five minutes of timestamp skew, and deduplicate event_id. Return 2xx after durable acceptance. Delivery is at least once; retries keep the same event ID. Redirects are not followed.
Find delivery history
Open Automations → Webhooks to see delivery attempts and retry failures. History appears after a job with a callback completes. The API exposes the same journal through GET /webhooks; use POST /webhooks/{deliveryID}/retry for a failed delivery.
Python & TypeScript
Clients handle quoting, submission, bounded polling and webhook verification. SDK source is included in the repository; packages are not yet published to PyPI or npm.
Python 3.10+
From a repository checkout, install with pip install ./sdk/python.
Python · server-side
import os
from scraper_client import Client
client = Client(os.environ["SCRAPER_URL"], os.environ["SCRAPER_API_KEY"])
result = client.scrape("https://example.com/",
request_payload={"content": {"main": True}})
print(result["data"]["markdown"])
TypeScript / Node.js 20+
Build the source package with npm --prefix sdk/typescript ci and npm --prefix sdk/typescript run build. In your application, install its local path with npm install /path/to/scraper-go/sdk/typescript.
TypeScript · server-side
import { Client } from '@scraper-go/client'
const client = new Client(process.env.SCRAPER_URL!, process.env.SCRAPER_API_KEY!)
const result = await client.scrape('https://example.com/', {
request_payload: { content: { main: true } }
})
console.log(result.data.markdown)
No repository access? Use the HTTP quickstart or import the OpenAPI specification into your API client.
n8n integration
n8n is a separate visual workflow service. Connect steps such as scraping a page, transforming its data and saving it to another application without building a full integration yourself.
Our integration is an importable workflow template. It requests a price, creates an Auto job, waits for completion and returns all result formats. n8n is not installed inside this panel.
- Download the JSON, open your n8n instance, and import the workflow from the file.
- Open Configure. Set
base to your externally reachable API URL ending in /api/v1, and set target_url. - Create an HTTP Header Auth credential: name
Authorization, value Bearer sg_YOUR_KEY. Select it in Quote, Create job, Job status and Result. - Execute the workflow manually. Read the response from Result, then connect your destination step.
The workflow contains no credentials. Polling stops on failure or after five minutes. If a response is lost, resume the existing job through the API instead of starting a new workflow execution.
Errors
| Status | Next action |
|---|
400 | Correct the request body or content selectors. |
401 / 403 | Check the API key and its permissions. |
402 | Add balance before starting another run. |
404 | Check the identifier and account ownership. |
409 | Approve a fresh price, or use a new key for changed intent. |
410 | The result is no longer retained. Stop polling. |
413 | The request exceeds 1 MiB. Send fewer batch items or shorter request bodies. |
422 | The job failed or was canceled. Read the job's error; stop polling. For a batch or batch quote, correct the items listed in errors. |
429 / 5xx | Honor Retry-After and retry with bounded backoff. Preserve the request's idempotency key. |
API reference
All paths below are relative to /api/v1. OpenAPI 3.0 contains request schemas, response types and authentication requirements.
| Endpoint | Purpose |
GET / POST/sessions | List or create saved sessions |
POST/sessions/{sessionID}/revoke | Forget saved browser state |
POST/sessions/{sessionID}/login | Open or resume interactive login |
POST/sessions/{sessionID}/renew-ip | Renew IP without deleting login data |
GET/profile-logins/{jobID} | Read the owned login frame |
POST/profile-logins/{jobID}/input | Send input, save or cancel |
POST/pricing/quote | Price and available countries |
POST/jobs | Create a job |
GET/jobs | List your jobs |
GET/jobs/{jobID} | Job status and billing |
GET/jobs/{jobID}/result | All result representations |
GET/jobs/{jobID}/artifact | Original artifact or screenshot |
POST/jobs/{jobID}/cancel | Cancel a job |
POST/job-batches | Submit a URL batch |
GET/job-batches/{batchID} | Batch status |
GET/saved-tasks | List saved tasks |
POST/saved-tasks | Save a task |
PUT/saved-tasks/{taskID} | Update or pause a task |
DELETE/saved-tasks/{taskID} | Delete a task and its schedule |
POST/saved-tasks/{taskID}/runs | Run a saved task |
POST/task-runs | Start a crawl or URL list |
GET/task-runs | Recent runs |
GET/task-runs/{runID} | Run status and pages |
POST/task-runs/{runID}/cancel | Cancel a run |
GET/webhooks | Delivery history and attempts |
GET/webhooks/signing-key | Your private signing secret |
POST/webhooks/{deliveryID}/retry | Retry a failed delivery |
GET/integrations/n8n.json | Public workflow download |
POST/maps | Discover sitemap URLs |
GET/maps/{jobID} | Read or filter discovered URLs |
GET/task-runs/{runID}/dataset | Paginated results or CSV, Excel and JSONL exports |
GET/task-runs/{runID}/changes | Compare page versions |
POST/task-runs/{runID}/retry | Retry failed pages |
GET/runs | All runs, batches and single jobs in one list |
GET/runs/{kind}/{runID} | One run by kind and ID |
POST/runs/hide | Hide runs from the list |
POST/runs/unhide | Show hidden runs again |
POST/runs/cancel | Stop runs |
POST/runs/retry | Retry the failed pages of runs |
POST/runs/rerun | Run runs again |
GET /runs lists task runs, batches and single jobs newest first, 50 per page (limit up to 100). Each run has a kind, a type (site, document, pages or page), a status (running, succeeded, partial, failed or canceled), its source, page counts and charged credits. Filter with view (all, failed_today, crawls, scheduled), comma-separated status and type, period (today, 7d, 30d) and search; counts per view ignore the filters; include_counts=false leaves them out. GET /runs/{kind}/{runID} reads one run the same way, hidden or not. POST /runs/hide with {"runs":[{"kind":"job","id":"…"}]} removes up to 500 runs from the list without touching billing or results; POST /runs/unhide undoes it.
include_totals=true adds totals: the pages and failed pages of every run the filters match. Each run's retry_cost_per_page is the per-page price cap its retry runs under, so pages.failed × retry_cost_per_page bounds what retrying it may cost. POST /runs/cancel, /runs/retry and /runs/rerun take the same runs list plus an idempotency_key (up to 64 characters, required for retry and rerun) and act on each run through its own endpoint's rules: a cancel stops a task run, a batch's unfinished jobs or a job; a retry starts a run's failed pages as a new URL-list run within that cap; a rerun starts a run again as a new run of its kind with the same requests. The answer lists the runs stopped or started, counts skipped runs (unknown, not yours, nothing to do) and names failed runs with the reason admission refused them. Repeating a retry or rerun with the same key starts nothing twice.
Run results and exports
GET /task-runs/{runID}/dataset returns columns, rows and total, with offset=0&limit=50 (maximum 100). Rows contain URL, job ID, status, error and values. Field columns use fields.name. Failed pages and unavailable results stay visible. Inline dataset pages are limited to 8 MiB (reduce columns/limit on HTTP 413); stored results larger than 16 MiB must be downloaded as individual artifacts. Pagination during a running crawl is a live view; wait for completion for a stable dataset.
After completion, add format=csv, format=xlsx or format=jsonl to download every page. Optionally select columns using a URL-encoded JSON array, such as columns=["url","status","fields.Price, USD"]; this also supports names containing commas. JSONL preserves nested values and large text; Excel rejects cells longer than 32,767 characters. CSV protects text that spreadsheet apps could interpret as formulas.
POST /task-runs/{runID}/retry with {"idempotency_key":"unique-retry-key"} creates a new run containing only failed or canceled pages. Successful pages are not repeated. The original per-page price ceiling applies, successful retries are charged normally, and repeating the same key returns the same run. In the panel, open Automations → Runs to choose columns, export or retry.