Hive

Documentation

Build with Hive.

Guides and API reference for browser automation, proxy connections, and research with Hive.

Getting started

Authenticate your requests.

Use your Hive API key in the Authorization: Bearer $HIVE_API_KEY header. Send JSON request bodies with Content-Type: application/json. Keep your key on your server.

Examples use https://hive.arglegal.live as the base URL. Check GET /capabilities before requesting a service. Use a unique Idempotency-Key for tunnel creation and Books submissions; reuse it only when repeating the same request.

Quick start

Create your first browser session

Set HIVE_API_KEY in your environment, then run this request. The response includes its selected browser family, profile ID and automation endpoint.

Shell · create a browser
curl --fail-with-body -X POST "https://hive.arglegal.live/browser/session" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{}'

Browsers use direct access from Hive by default. To enable a managed proxy, send {"use_proxy":true,"country":"AR"}. Country, region, ASN, mobile network, and identity constraints apply to proxy connections. A country value alone does not change a direct session’s exit IP.

Next, follow the browser automation guide to open a page with Playwright.

Browser transport

Connect to your profile.

Create a reusable profile with POST /browser/profiles and {"browser":"chrome"} or {"browser":"firefox"}. Omit the family for a random choice. Chrome uses Donut; Firefox uses Camoufox.

Save the returned id, then create sessions with {"profile_id":"YOUR_PROFILE_ID"}. Cookies, origin storage and fingerprint survive session closure and service replacement. Use the existing context shown below.

Add {"network":{"country":"AR","asn":"7303","region":"Buenos Aires"}} when creating a profile to save its proxy policy. A profile also binds on its first managed proxy launch. Later sessions need only the profile ID: Hive restores the network settings, verifies a matching exit and applies its locale and timezone. Replacements retain the saved ASN, region and timezone; unavailable matches return 503. Inspect network.last_exit on the profile for the last verified IP. Sticky routing does not guarantee a permanent IP.

Profiles belong to your API key. Active reuse and conflicting selections return 409; unknown profiles return 404. Close the session before reuse or deletion. Unavailable engines fail without switching your identity.

Python · Playwright
import os
from playwright.sync_api import sync_playwright

auth = {"Authorization": "Bearer " + os.environ["HIVE_API_KEY"]}
# session is the result of POST /browser/session.
with sync_playwright() as p:
    browser = (p.firefox.connect(session["ws_endpoint"], headers=auth)
        if session["browser"] == "firefox" else
        p.chromium.connect_over_cdp(session["cdp_url"], headers=auth))
    context = browser.contexts[0]
    page = context.pages[0] if context.pages else context.new_page()
    page.goto("https://example.com")
    browser.close()
# DELETE /browser/session?session_id=... to release the lease.

Network transport

Use your own client.

Python · HTTPX
import httpx
from urllib.parse import quote

# tunnel is the result of POST /network/tunnels.
ep = tunnel["endpoints"]["http"]
credentials = tunnel["credentials"]
user = quote(credentials["username"], safe="")
password = quote(credentials["password"], safe="")
proxy = f"http://{user}:{password}@{ep['host']}:{ep['port']}"
with httpx.Client(proxy=proxy, timeout=30) as client:
    response = client.get("https://example.com")
    response.raise_for_status()
# SOCKS clients use endpoints.socks5 with socks5h:// for remote DNS.

Resources & exchanges

Request. Poll. Release.

Create a resource with POST /resources/phone, /resources/mail, /resources/captcha, or /resources/network. Keep the returned resource_id. Read the recorded state with GET, poll for progress with POST, and DELETE the resource when finished.

Phone requests use service and a two-letter country. Mail can specify a domain. Captcha requests use captcha_type with site_key and page_url, or image_b64 for an image. For a browser-bound challenge, supply browser_context with the submitting browser’s user_agent and relevant cookies. Cookie entries use name, value, domain, path, and optional http_only, secure, and expires. Hive forwards that context for the solve without storing cookie values in resource records. A token is ready to submit when the resource becomes ready; the target website still decides whether to accept it. The reference below lists each field.

Exchanges return one result: send messages to POST /exchange/llm or a prompt to POST /exchange/image.

Completion request body
{
  "messages": [
    {
      "role": "user",
      "content": "Summarize this paragraph."
    }
  ],
  "max_tokens": 200
}

Research

Public LinkedIn data.
Search through Hive.

Search for people, find company employees, and retrieve public profiles, posts, comments, and reactions. Use filters and pagination to select the results you need.

Check /capabilities for service availability and the API reference for each operation’s request fields. Pagination is finite: limit is 1–1000, page is 1–100, and pages is 1–20, with tighter limits on individual target and filter lists.

  • POST /scrapers/linkedin/people/search
  • POST /scrapers/linkedin/company/employees
  • POST /scrapers/linkedin/profile/details
  • POST /scrapers/linkedin/profile/posts
  • POST /scrapers/linkedin/post/search
  • POST /scrapers/linkedin/post/comments
  • POST /scrapers/linkedin/profile/comments
  • POST /scrapers/linkedin/profile/reactions
Profile details
{
  "profiles": [
    "https://www.linkedin.com/in/<profile-id>/"
  ],
  "detail_level": "full"
}

Public research

People, posts, web, and images.

Find candidate people and retrieve public profiles and posts on Instagram, X, Facebook, and TikTok; search Facebook groups, Reddit posts and comments, X quotes and Spaces, and saved Instagram Highlight stories. Discover username candidates with Sherlock, search Google, and find image matches or text with Google Lens. Use the same Hive bearer key and strict snake_case JSON bodies for every operation. Unknown fields and invalid targets are rejected before work begins.

Profile requests accept 1–10 handles or canonical public URLs in profiles. Facebook also accepts numeric IDs and profile.php URLs. Results contain items, count, and unresolved. Each entry echoes the exact target and zero-based input_index, in input order, including duplicates. Unresolved reasons are not_found, private, unavailable, malformed_data, or ambiguous; omitted profiles are unavailable, not proof of nonexistence. Resolved entries include public source links, and available identity, biography, image, verification, and count fields. Missing optional fields are omitted; profiles can be private, sparse, renamed, or unavailable.

Sherlock accepts 1–5 usernames of up to 64 characters and a limit of 1–100 per distinct username (default 20). Every result has username, site, profile_url, and status: candidate. Matching handles do not establish identity. truncated: true means a work or result limit omitted results; null means completeness is unknown.

Google accepts 1–5 queries, each up to 512 characters and 32 words; pages is 1–3 per query (default 1), and limit is 1–100 results per distinct query (default 20). Optional country and language select a supported two-letter country and search interface language. Organic results preserve query, rank, title, url, and an optional snippet.

Google Lens accepts one public HTTPS image_url or a private uploaded asset_id without credentials or private addresses, plus distinct search_types: visual, exact, and/or text. The URL is at most 4096 characters. limit is 1–100 total records (default 20), interleaved in requested type order. Visual and exact matches carry a title and public URL with optional image metadata; text records carry extracted text. A valid search may return no matches. Results do not promise completeness, totals, or a next page.

Google and Sherlock return up to 500 items for five inputs, grouped in input order; duplicate inputs are collected once. Sherlock scans each username independently so earlier inputs cannot consume later inputs’ collection windows. X account creation dates use ISO 8601 UTC; an ambiguous verification flag without badge type or blue status becomes null.

operation_health provides each operation’s reason, observed_at (Unix time), and retry_after_s. Reasons distinguish rate limiting, timeout, invalid response, limited capacity, service unavailability, missing resources and invalid requests. Capacity is shared across operations and bearer keys on the deployment. Retry timing is a minimum recheck delay; a null value means no delay is known. Elapsed time does not establish recovery. Accounting recovery can clear a capacity failure to unverified until a new search succeeds. Job recovery still uses the original key.

Read scrapers.<name>.status and operation_status from GET /capabilities. unconfigured means no active configuration; configured means all operations are unverified; usable means at least one observed success with no recorded failure; degraded means an operation has a known failure. A usable LinkedIn service does not establish that all eight operations work. Observations survive restart and reset when service configuration changes. Discovery never launches research.

Calls have finite execution limits and create no lease. Do not automatically repeat a timed-out or disconnected research call: work may already have completed. Social posts and people search support same-key recovery, as do Lens jobs. The older synchronous LinkedIn, profile/details, Sherlock, Google and Lens search routes do not support idempotent replay. Errors use bad_request, unauthorized, not_found, idempotency_conflict, expired, rate_limited, unavailable, or upstream_timeout. Invalid Hive authentication returns 401. Invalid fields return 400 with safe field detail before execution. Service failures return 503 unavailable; retained job errors and capability operation status use the same code. Keep using your Hive key. Respect retry_after_s when present.

Upload private image bytes with POST /scrapers/assets and Content-Type image/jpeg, image/png, or image/webp (10 MiB maximum). Use the returned asset_id instead of image_url. Images are encrypted, bearer-scoped, and retained for 24 hours. Their bytes are sent to the image-search service for the requested search. Download with GET /scrapers/assets/{asset_id}; delete with DELETE /scrapers/assets/{asset_id}.

Submit Lens work through POST /scrapers/google_lens/jobs with a required Idempotency-Key. The 202 response contains a poll_url for GET /scrapers/google_lens/jobs/{job_id}. Repeat the same key and body after a lost submission to retrieve the same job. Poll until completed (with result), failed, or outcome_unknown. An interrupted run is never automatically restarted. Results expire after 24 hours (410); keys stay reserved for seven days. Different bodies under one key return 409.

Lens unwraps supported Google redirect links. Embedded thumbnails have a bearer-authenticated Hive image_url, image_sha256, and image_kind: thumbnail. These hashes identify the returned thumbnail bytes, not full-resolution originals. Resolve relative image paths against your Hive origin; send your Hive key only to Hive, never to public image hosts.

Instagram, X, Facebook and TikTok expose native people/search: a free-text query (1–200 characters), up to five locations hints (1–100 characters), and limit 1–50. Hints are appended in caller order to one query; ambiguous names remain text, and city and employer filters are not enforced. The response reports effective_query, location_handling and filters_enforced: false. Instagram commas become spaces to keep one search. Candidates reuse profile fields and add the exact query, absolute rank, status: candidate and missing_fields. Missing biography/avatar fields remain null; collection-wide partial status is carried on every page. Facebook returns people with canonical profile URLs and nullable usernames or IDs; its location line is not a biography. Opaque people links retain their source URL without an inferred handle or numeric ID; details/posts require supported handle or numeric targets. Facebook cards without a usable source profile URL are omitted and make the whole collection partial; if every returned card is unusable, the search fails. TikTok returns user cards. Candidates are not verified identities and have no lookup target attribution.

People search uses the same job, key and cached-page workflow as posts, with at most 50 candidates in native order after identity deduplication. Keep the same query, locations and limit on later pages. No extra profile enrichment occurs. Replays and cached pages never repeat a search. A valid empty result differs from a failed search; completeness remains unknown. Check discovery health separately from profile lookup and posts in operation_status.

Instagram, X, Facebook and TikTok also expose profile/posts. Initial requests require Idempotency-Key; send Prefer: respond-async for a prompt 202 job, or wait up to 20 seconds for a page. Poll GET /scrapers/jobs/{job_id} with your bearer. Repeat the original key and body to recover the same job; results expire after 24 hours.

Posts accept 1–10 profiles and a total page limit of 1–50. Later pages send the same body plus result_id and page, without requiring a key or starting another collection. Recent jobs collect up to 50 posts per target by default (40 for Facebook, up to 100 for X) and 500 attributed items. Target outcomes preserve input order and duplicates. Source authors, carousel order, video thumbnails, actual tags and separate caption mentions are preserved. Facebook covers public people and pages, including numeric profiles. TikTok preserves ordered slideshow photos, playable references when supplied, and separate cover thumbnails. A post page URL is not video bytes. Null metadata and history_complete: null indicate unknown information. Public media references can expire; send Hive credentials only to Hive.

For older posts on any of the four services, submit one profile with archive: true. X, TikTok, Instagram and Facebook Pages use archive_end_date and archive_window_days (1–30); Pages also require archive_kind: page. X supports media_filter: images or videos. Facebook personal profiles use archive_kind: person and a cursor without dates; set per_target_limit to at least 5 (default 100) so cursors can advance. Read all cached pages first. Continue Instagram within a date window and Facebook personal profiles by submitting next_archive_cursor as archive_cursor in a new job. Hive filters repeated Facebook personal posts and marks a stalled cursor partial. When an Instagram window has no cursor, use next_archive_end_date to search older dates. X, TikTok and Facebook Pages continue directly with next_archive_end_date. Use a new idempotency key per job. Narrow a truncated date window when no cursor is available. X and TikTok allow up to 1,000 posts per archive window; Instagram and Facebook allow up to 500. Provider coverage cannot prove a complete platform history.

For deeper X collection, submit POST /scrapers/x/profile/dumps with one profiles entry and an Idempotency-Key. Poll GET /scrapers/x/profile/dumps/{job_id}, then page deduplicated posts at GET /scrapers/x/profile/dumps/{job_id}/posts with after=0 and each returned next_after. A running job can add posts after has_more: false. Continue partial jobs with POST /scrapers/x/profile/dumps/{job_id}/resume; retry_stalled=true revisits a cursor cycle. Posts include photo and video source URLs, including video variants; clients download media bytes themselves. history_coverage compares the collected post count with X's displayed count, and timelines_exhausted means X stopped offering cursors. Neither proves complete account history.

Set retain_media: true to preserve returned photos and video thumbnails as private JPEG/PNG/WebP evidence. Each media item reports retention_status, a fixed retention_reason, and an optional asset with its authenticated relative URL, exact-byte SHA-256, size, MIME and expiry. The original source URL and original/thumbnail role remain separate. Recent jobs allow 20 unique images, 32 MiB and 30 seconds; archive jobs allow 100 images, 96 MiB and 120 seconds. Each image is capped at 10 MiB and 40 million pixels. Assets share the bearer/global image quotas and expire after 24 hours. Videos remain public references. media_count_reported and media_complete identify known missing images: Instagram archive carousels currently supply only their first slide and mark the page partial. Unresolved targets, media failures and retention limits also set collection-wide partial status. Duplicate sources reuse assets; replay never redownloads, including after deletion or expiry. Keep the same retention value on cached page requests.

Recover a submission and read its cached pages

Read GET /scrapers/capacity for shared research availability and retry guidance. Recheck after a known retry delay; elapsed time does not prove recovery. Polling and replay preserve existing work, and individual requests can still be unavailable.

Keep the collection-wide partial status on every cached page when interpreting results. The last cached page does not establish complete account history.

Use an Idempotency-Key of 1–200 printable ASCII characters without spaces. It binds your bearer, service, operation, ordered business request and limit. A different body under the same key, or a changed body/limit with result_id, returns 409 idempotency_conflict. Foreign or wrong-operation references return 404 not_found; expired results return 410 expired. Replay tombstones reserve keys for seven days. Reads never renew expiry. A terminal failed or outcome_unknown job cannot relaunch through replay; keep the key after an interruption.

A 202 submission includes Location, Retry-After: 2 and, when requested, Preference-Applied: respond-async. Polling returns HTTP 200 with the job envelope; completion puts the first page under result. Completed replays and cached pages return HTTP 200 even with Prefer: respond-async. Initial collection starts at page 1; later pages are 1–500 and must stay inside the collected window. An out-of-window page returns 400 bad_request. Follow next_page until null. Polls, replays and cached pages perform zero additional work.

The Bash workflow below needs curl, jq and sha256sum. Choose one service/operation and preserve its key and request. It demonstrates an explicit same-key replay, bounded polling, a second cached page when available, and an authenticated image download. A thumbnail hash identifies that thumbnail’s exact bytes; an original image hash identifies the retained original bytes.

Bash · recovery, cached page and image hash
# Run in Bash with curl, jq and sha256sum installed.
set -euo pipefail
: "${HIVE_API_KEY:?Set your Hive bearer in the environment}"
HIVE_ORIGIN="https://hive.arglegal.live"
HIVE_SERVICE="instagram"  # instagram, x, facebook or tiktok
HIVE_OPERATION="profile/posts"
# Keep this key and the exact request after a disconnect; use a new key only
# for an intentional new collection. Do not enable shell tracing.
HIVE_RESEARCH_KEY="investigation-social-001"
request='{"profiles":["nasa"],"limit":2,"retain_media":true}'
# For native people discovery instead, set these before calling submit:
# HIVE_OPERATION="people/search"
# request='{"query":"José García","locations":["Madrid"],"limit":2}'
submit() {
  curl --fail-with-body --max-time 30 -sS \
    "$HIVE_ORIGIN/scrapers/$HIVE_SERVICE/$HIVE_OPERATION" \
    -H "Authorization: Bearer $HIVE_API_KEY" \
    -H "Content-Type: application/json" \
    -H "Idempotency-Key: $HIVE_RESEARCH_KEY" \
    -H "Prefer: respond-async" --data "$request"
}
reply="$(submit)"
# If the response was lost, run submit again with the SAME key and request.
# This explicit replay also demonstrates recovery without another collection.
reply="$(submit)"
for ((poll=0; poll<32; poll++)); do
  status="$(jq -r '.status // "completed"' <<< "$reply")"
  case "$status" in
    completed) break ;;
    failed|outcome_unknown)
      jq '{job_id,status,error}' <<< "$reply"
      exit 1 ;;  # Keep the key; a new key would start separate work.
    pending|running) ;;
    *) exit 1 ;;
  esac
  poll_url="$(jq -r '.poll_url' <<< "$reply")"
  [[ "$poll_url" =~ ^/scrapers/jobs/[a-zA-Z0-9_-]+$ ]]
  reply="$(curl --fail-with-body --max-time 30 -sS \
    "$HIVE_ORIGIN$poll_url?wait_seconds=25" -H "Authorization: Bearer $HIVE_API_KEY")"
done
[[ "$(jq -r '.status // "completed"' <<< "$reply")" == completed ]] || exit 1
page="$(jq '.result // .' <<< "$reply")"
jq '{result_id,page,count,collected_count,next_page,partial}' <<< "$page"
# The first completed page is stable. Follow next_page only when present.
if [[ "$(jq -r '.next_page' <<< "$page")" != null ]]; then
  next_request="$(jq --argjson page "$page" \
    '. + {result_id:$page.result_id,page:$page.next_page}' <<< "$request")"
  curl --fail-with-body --max-time 30 -sS \
    "$HIVE_ORIGIN/scrapers/$HIVE_SERVICE/$HIVE_OPERATION" \
    -H "Authorization: Bearer $HIVE_API_KEY" \
    -H "Content-Type: application/json" --data "$next_request"
fi
# Hash one retained original image or thumbnail from the first page.
asset="$(jq -c '[.items[].media[]? | select(.retention_status == "retained")][0] // null' <<< "$page")"
if [[ "$asset" != null ]]; then
  asset_url="$(jq -r '.asset.url' <<< "$asset")"
  [[ "$asset_url" =~ ^/scrapers/assets/[a-zA-Z0-9_-]+$ ]]
  curl --fail-with-body --max-time 30 -sS \
    "$HIVE_ORIGIN$asset_url" -H "Authorization: Bearer $HIVE_API_KEY" \
    --output retained-image.bin
  expected="$(jq -r '.asset.sha256' <<< "$asset")"
  printf '%s  retained-image.bin\n' "$expected" | sha256sum --check -
  jq '{role,sha256:.asset.sha256,expires_at:.asset.expires_at}' <<< "$asset"
fi

Examples for each operation

These response shapes are illustrative. Social submission examples show the initial 202 job separately from the completed page under the polled result. A completed replay can return the page directly with 200. Public data and optional fields can differ.

POST /scrapers/instagram/profile/details

JSON request body
{
  "profiles": [
    "nasa"
  ]
}
Illustrative result
{
  "items": [
    {
      "username": "nasa",
      "profile_url": "https://www.instagram.com/nasa/",
      "display_name": "NASA",
      "target": "nasa",
      "input_index": 0
    }
  ],
  "count": 1,
  "unresolved": []
}

POST /scrapers/instagram/profile/posts

Request headers
{
  "Idempotency-Key": "instagram-profile-posts-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/profile/posts" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: instagram-profile-posts-001" \
  -H "Prefer: respond-async" \
  --data '{"profiles":["nasa"],"limit":2}'
JSON request body
{
  "profiles": [
    "nasa"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "instagram",
  "operation": "profile/posts",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "target": "nasa",
      "input_index": 0,
      "post_id": "123",
      "url": "https://www.instagram.com/p/example/",
      "author": {
        "username": "nasa",
        "profile_url": "https://www.instagram.com/nasa"
      },
      "timestamp": "2026-09-01T12:00:00+00:00",
      "text": "Public post example.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "context": []
    }
  ],
  "count": 1,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 500,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": false,
  "targets": [
    {
      "target": "nasa",
      "input_index": 0,
      "status": "resolved",
      "collected_count": 1,
      "reason": null
    }
  ],
  "unresolved": [],
  "per_target_limit": 50,
  "history_complete": null
}

POST /scrapers/instagram/people/search

Request headers
{
  "Idempotency-Key": "instagram-people-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/people/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: instagram-people-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"José García","locations":["Madrid"],"limit":2}'
JSON request body
{
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "instagram",
  "operation": "people/search",
  "status": "pending",
  "created_at": 1789905600,
  "expires_at": 1789992000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "query": "José García",
      "rank": 1,
      "status": "candidate",
      "username": "example_person",
      "profile_url": "https://www.instagram.com/example_person",
      "profile_id": "123456",
      "display_name": "José García",
      "biography": null,
      "image_url": "https://images.example.com/avatar.jpg",
      "missing_fields": [
        "biography"
      ]
    }
  ],
  "count": 1,
  "result_id": "rresult_22222222222222222222222222222222",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1789992000,
  "truncated": null,
  "partial": true,
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "effective_query": "José García Madrid",
  "location_handling": "query_hint",
  "filters_enforced": false
}

POST /scrapers/instagram/profile/followers

Request headers
{
  "Idempotency-Key": "instagram-profile-followers-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/profile/followers" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: instagram-profile-followers-001" \
  -H "Prefer: respond-async" \
  --data '{"profile":"nasa","collection_size":100,"limit":2}'
JSON request body
{
  "profile": "nasa",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "instagram",
  "operation": "profile/followers",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "username": "exampleperson",
      "profile_url": "https://www.instagram.com/exampleperson/",
      "profile_id": "123",
      "display_name": null,
      "image_url": null,
      "biography": null,
      "location": null,
      "created_at": null,
      "is_private": null,
      "is_verified": null,
      "source_profile": "nasa",
      "relation": "followers",
      "relation_evidence": "source_attribution"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 100,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "source_profile": "nasa",
  "relation": "followers",
  "upstream_continuation_available": false
}

Synthetic attributed followers edge. Returned source and direction must prove each edge; the source root handle and observed root ID are excluded. Biography, location and account creation remain null, with other optional identity fields nullable. One finite first collection supports cached pages only; no continuation credentials are accepted or returned. Graph completeness remains unknown. Account restrictions, caps and private/hidden profiles can limit results; public identity fields do not grant private content access.

POST /scrapers/instagram/profile/following

Request headers
{
  "Idempotency-Key": "instagram-profile-following-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/profile/following" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: instagram-profile-following-001" \
  -H "Prefer: respond-async" \
  --data '{"profile":"nasa","collection_size":100,"limit":2}'
JSON request body
{
  "profile": "nasa",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "instagram",
  "operation": "profile/following",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "username": "exampleperson",
      "profile_url": "https://www.instagram.com/exampleperson/",
      "profile_id": "123",
      "display_name": null,
      "image_url": null,
      "biography": null,
      "location": null,
      "created_at": null,
      "is_private": null,
      "is_verified": null,
      "source_profile": "nasa",
      "relation": "following",
      "relation_evidence": "source_attribution"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 100,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "source_profile": "nasa",
  "relation": "following",
  "upstream_continuation_available": false
}

Synthetic attributed following edge. Returned source and direction must prove each edge; the source root handle and observed root ID are excluded. Biography, location and account creation remain null, with other optional identity fields nullable. One finite first collection supports cached pages only; no continuation credentials are accepted or returned. Graph completeness remains unknown. Account restrictions, caps and private/hidden profiles can limit results; public identity fields do not grant private content access.

POST /scrapers/instagram/profile/tagged

Request headers
{
  "Idempotency-Key": "instagram-profile-tagged-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/profile/tagged" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: instagram-profile-tagged-001" \
  -H "Prefer: respond-async" \
  --data '{"profile":"nasa","collection_size":50,"limit":2}'
JSON request body
{
  "profile": "nasa",
  "collection_size": 50,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "instagram",
  "operation": "profile/tagged",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "456",
      "url": "https://www.instagram.com/p/ExamplePost/",
      "author": {
        "username": "examplephotographer",
        "profile_url": "https://www.instagram.com/examplephotographer/"
      },
      "timestamp": "2020-06-01T12:00:00+00:00",
      "text": "Synthetic science post.",
      "media": [],
      "geotag": null,
      "tagged_users": [
        {
          "username": "nasa",
          "profile_url": "https://www.instagram.com/nasa/"
        }
      ],
      "mentions": null,
      "tagged_profile": "nasa",
      "relation_evidence": "tagged_users"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 50,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "tagged_profile": "nasa"
}

Synthetic incoming post with an actual returned NASA tag, retaining its source author. Caption mentions and echoed request fields do not establish tags. Ambiguous output cannot establish no tags; retained valid rows may be partial. Collection completeness is unknown. Source media references can expire and these operations retain no media bytes.

POST /scrapers/instagram/profile/highlights

Request headers
{
  "Idempotency-Key": "instagram-profile-highlights-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/profile/highlights" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: instagram-profile-highlights-001" \
  -H "Prefer: respond-async" \
  --data '{"profile":"nasa","collection_size":10,"limit":2}'
JSON request body
{
  "profile": "nasa",
  "collection_size": 10,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "instagram",
  "operation": "profile/highlights",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "highlight_id": "123",
      "source_profile": "nasa",
      "title": "Synthetic saved Highlight",
      "cover_image_url": null,
      "reported_media_count": null,
      "created_at": null,
      "updated_at": null,
      "latest_story_at": null,
      "content_retrieved": false,
      "relation_evidence": "profile_container"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 10,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 10,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "source_profile": "nasa",
  "profile_status": "partial",
  "profile_reason": "partial_response",
  "content_available": false
}

Synthetic saved Highlight metadata only: observed ID with nullable cover, count and dates. profile_status distinguishes success, partial, private, not_found, unavailable and rate_limited; a public profile with zero summaries differs from a failed or missing profile. Story content is unavailable, no content arrays are requested, and unexpected content is rejected. These summaries do not establish complete history. Covers may expire and no media bytes are retained.

POST /scrapers/instagram/profile/highlights/content

Request headers
{
  "Idempotency-Key": "hive-doc-instagram-profile-highlights-content",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/instagram/profile/highlights/content" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-instagram-profile-highlights-content" \
  -H "Prefer: respond-async" \
  --data '{"limit":2,"profile":"nasa","collection_size":5,"highlight_limit":1}'
JSON request body
{
  "limit": 2,
  "profile": "nasa",
  "collection_size": 5,
  "highlight_limit": 1
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "instagram",
  "operation": "profile/highlights/content",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "story_id": "3975089864169193671_528817151",
      "media_id": "3975089864169193671",
      "source_profile": "nasa",
      "owner_id": "528817151",
      "highlight_id": "18195781759377100",
      "highlight_title": "Roman",
      "highlight_url": "https://www.instagram.com/stories/highlights/18195781759377100/",
      "timestamp": "2026-08-30T11:00:53+00:00",
      "expires_at": null,
      "media_type": "video",
      "media_url": "https://cdn.example.com/story.mp4",
      "thumbnail_url": "https://cdn.example.com/image.jpg",
      "width": 640,
      "height": 1136,
      "caption": null,
      "music_title": null,
      "music_artist": null,
      "mentions": [
        "nasakennedy"
      ],
      "hashtags": [],
      "links": [],
      "position": 1,
      "relation_evidence": "observed_owner_and_highlight",
      "media_retained": false
    },
    {
      "story_id": "3975097618556613637_528817151",
      "media_id": "3975097618556613637",
      "source_profile": "nasa",
      "owner_id": "528817151",
      "highlight_id": "18195781759377100",
      "highlight_title": "Roman",
      "highlight_url": "https://www.instagram.com/stories/highlights/18195781759377100/",
      "timestamp": "2026-08-30T11:16:12+00:00",
      "expires_at": null,
      "media_type": "video",
      "media_url": "https://cdn.example.com/story.mp4",
      "thumbnail_url": "https://cdn.example.com/image.jpg",
      "width": 640,
      "height": 1136,
      "caption": null,
      "music_title": null,
      "music_artist": null,
      "mentions": [],
      "hashtags": [],
      "links": [],
      "position": 2,
      "relation_evidence": "observed_owner_and_highlight",
      "media_retained": false
    }
  ],
  "count": 2,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": 2,
  "collected_count": 5,
  "collection_limit": 200,
  "expires_at": 1791154800,
  "truncated": true,
  "partial": false,
  "source_profile": "nasa",
  "requested_collection_size": 5,
  "requested_highlight_limit": 1,
  "content_scope": "saved_highlight_stories",
  "collection_complete": null,
  "partial_reasons": [],
  "media_retained": false
}

Retrieve actual saved Highlight Story IDs, observed owner identities, timestamps and image or video references. A compound Story ID preserves its owner suffix; media_id is the numeric media component. Links may expire and media bytes are not retained. Collection size is 1–200 and highlight_limit is 1–25. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/x/profile/details

JSON request body
{
  "profiles": [
    "nasa"
  ]
}
Illustrative result
{
  "items": [
    {
      "username": "nasa",
      "profile_url": "https://x.com/NASA",
      "display_name": "NASA",
      "target": "nasa",
      "input_index": 0
    }
  ],
  "count": 1,
  "unresolved": []
}

POST /scrapers/x/profile/posts

Request headers
{
  "Idempotency-Key": "x-profile-posts-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/profile/posts" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: x-profile-posts-001" \
  -H "Prefer: respond-async" \
  --data '{"profiles":["nasa"],"limit":2}'
JSON request body
{
  "profiles": [
    "nasa"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "x",
  "operation": "profile/posts",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "target": "nasa",
      "input_index": 0,
      "post_id": "123",
      "url": "https://x.com/nasa/status/123",
      "author": {
        "username": "nasa",
        "profile_url": "https://x.com/nasa"
      },
      "timestamp": "2026-09-01T12:00:00+00:00",
      "text": "Public post example.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "context": []
    }
  ],
  "count": 1,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 1000,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": false,
  "targets": [
    {
      "target": "nasa",
      "input_index": 0,
      "status": "resolved",
      "collected_count": 1,
      "reason": null
    }
  ],
  "unresolved": [],
  "per_target_limit": 50,
  "history_complete": null
}

POST /scrapers/x/people/search

Request headers
{
  "Idempotency-Key": "x-people-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/people/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: x-people-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"José García","locations":["Madrid"],"limit":2}'
JSON request body
{
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "x",
  "operation": "people/search",
  "status": "pending",
  "created_at": 1789905600,
  "expires_at": 1789992000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "query": "José García",
      "rank": 1,
      "status": "candidate",
      "username": "example_person",
      "profile_url": "https://x.com/example_person",
      "profile_id": "123456",
      "display_name": "José García",
      "biography": "Synthetic public biography",
      "image_url": "https://images.example.com/avatar.jpg",
      "missing_fields": [],
      "location": "Madrid"
    }
  ],
  "count": 1,
  "result_id": "rresult_22222222222222222222222222222222",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1789992000,
  "truncated": null,
  "partial": false,
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "effective_query": "José García Madrid",
  "location_handling": "query_hint",
  "filters_enforced": false
}

POST /scrapers/x/profile/followers

Request headers
{
  "Idempotency-Key": "x-followers-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/profile/followers" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: x-followers-001" \
  -H "Prefer: respond-async" \
  --data '{"profile":"nasa","collection_size":100,"limit":2}'
JSON request body
{
  "profile": "nasa",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "x",
  "operation": "profile/followers",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "username": "example_person",
      "profile_url": "https://x.com/example_person",
      "source_profile": "nasa",
      "relation": "followers",
      "relation_evidence": "source_attribution"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "source_profile": "nasa",
  "relation": "followers",
  "requested_collection_size": 100,
  "graph_complete": null,
  "partial_reasons": [
    "missing_optional_fields"
  ]
}

Synthetic attributed edge. Optional profile fields may be missing. Full graph completeness is unknown; null next_page only ends cached pages. There is no upstream graph continuation.

POST /scrapers/x/profile/following

Request headers
{
  "Idempotency-Key": "x-following-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/profile/following" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: x-following-001" \
  -H "Prefer: respond-async" \
  --data '{"profile":"nasa","collection_size":100,"limit":2}'
JSON request body
{
  "profile": "nasa",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "x",
  "operation": "profile/following",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "username": "example_person",
      "profile_url": "https://x.com/example_person",
      "source_profile": "nasa",
      "relation": "following",
      "relation_evidence": "source_attribution"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "source_profile": "nasa",
  "relation": "following",
  "requested_collection_size": 100,
  "graph_complete": null,
  "partial_reasons": [
    "missing_optional_fields"
  ]
}

Synthetic attributed edge. Optional profile fields may be missing. Full graph completeness is unknown; null next_page only ends cached pages. There is no upstream graph continuation. Following is conditional: ambiguous output without qualified source and direction attribution fails unavailable. The source root is excluded.

POST /scrapers/x/posts/search

Request headers
{
  "Idempotency-Key": "x-posts-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/posts/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: x-posts-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"Artemis","from_username":"nasa","since_date":"2020-01-01","until_date":"2021-01-01","collection_size":100,"limit":2}'
JSON request body
{
  "query": "Artemis",
  "from_username": "nasa",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "x",
  "operation": "posts/search",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "456",
      "url": "https://x.com/example_person/status/456",
      "author": {
        "username": "example_person",
        "profile_url": "https://x.com/example_person"
      },
      "timestamp": "2020-06-01T12:00:00+00:00",
      "text": "Synthetic public post example.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "effective_query": "Artemis from:nasa since:2020-01-01 until:2021-01-01",
      "relation_kind": "search_match",
      "relation_evidence": "search_query",
      "relation_target_post_id": null,
      "in_reply_to_post_id": null,
      "conversation_id": null,
      "quoted_post_id": null,
      "is_reply": null,
      "is_quote": null
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "effective_query": "Artemis from:nasa since:2020-01-01 until:2021-01-01",
  "query_scope": "post_search",
  "requested_collection_size": 100,
  "target_post_id": null,
  "search_complete": null,
  "filters_enforced": false,
  "partial_reasons": [
    "rejected_rows_or_unknown_parent"
  ]
}

Search operators are passed upstream; matching and recall are not independently guaranteed. until_date is exclusive. Synthetic result. Cached pagination never expands the collected window; completeness is unknown.

POST /scrapers/x/post/replies

Request headers
{
  "Idempotency-Key": "x-post-replies-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/post/replies" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: x-post-replies-001" \
  -H "Prefer: respond-async" \
  --data '{"post_id":"123","collection_size":100,"limit":2}'
JSON request body
{
  "post_id": "123",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "x",
  "operation": "post/replies",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "456",
      "url": "https://x.com/example_person/status/456",
      "author": {
        "username": "example_person",
        "profile_url": "https://x.com/example_person"
      },
      "timestamp": "2020-06-01T12:00:00+00:00",
      "text": "Synthetic public post example.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "effective_query": "conversation_id:123",
      "relation_kind": "conversation_candidate",
      "relation_evidence": "search_query",
      "relation_target_post_id": "123",
      "in_reply_to_post_id": null,
      "conversation_id": "123",
      "quoted_post_id": null,
      "is_reply": null,
      "is_quote": null
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "effective_query": "conversation_id:123",
  "query_scope": "conversation_candidates",
  "requested_collection_size": 100,
  "target_post_id": "123",
  "search_complete": null,
  "filters_enforced": false,
  "partial_reasons": [
    "rejected_rows_or_unknown_parent"
  ]
}

Conversation matches are candidates. Only observed in_reply_to_post_id values establish direct or nested replies; a null parent stays unknown. This is not a complete reply tree. Synthetic result. Cached pagination never expands the collected window; completeness is unknown.

POST /scrapers/x/post/quotes

Request headers
{
  "Idempotency-Key": "hive-doc-x-post-quotes",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/post/quotes" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-x-post-quotes" \
  -H "Prefer: respond-async" \
  --data '{"limit":2,"post_id":"2065225362544726371","collection_size":50}'
JSON request body
{
  "limit": 2,
  "post_id": "2065225362544726371",
  "collection_size": 50
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "x",
  "operation": "post/quotes",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "2065250261493600416",
      "url": "https://x.com/i/status/2065250261493600416",
      "author": {
        "username": "theo",
        "profile_url": "https://x.com/theo",
        "profile_id": null,
        "display_name": "Theo - t3.gg",
        "image_url": "https://cdn.example.com/image.jpg"
      },
      "timestamp": "2026-06-12T01:50:07+00:00",
      "text": "This might actually be a bit too generous. I am getting suspicious",
      "media": [],
      "media_count_reported": null,
      "media_complete": null,
      "geotag": null,
      "tagged_users": null,
      "mentions": [],
      "likes_count": 4391,
      "comments_count": 190,
      "views_count": 409207,
      "reposts_count": 53,
      "quotes_count": 3,
      "relation_kind": "incoming_quote",
      "relation_evidence": "provider_quote_target",
      "relation_target_post_id": "2065225362544726371",
      "quoted_post_id": "2065225362544726371",
      "in_reply_to_post_id": null,
      "conversation_id": "2065250261493600416",
      "is_reply": null,
      "is_quote": true
    },
    {
      "post_id": "2065257221077020850",
      "url": "https://x.com/i/status/2065257221077020850",
      "author": {
        "username": "minchoi",
        "profile_url": "https://x.com/minchoi",
        "profile_id": null,
        "display_name": "Min Choi",
        "image_url": "https://cdn.example.com/image.jpg"
      },
      "timestamp": "2026-06-12T02:17:46+00:00",
      "text": "Codex now lets you save rate limit resets and use them later.\n\nSmall change, but useful.\nhttps://t.co/aDOQwejJ78",
      "media": [
        {
          "type": "image",
          "url": "https://pbs.twimg.com/media/HKknmxXa8AAjyEC.jpg",
          "role": "original",
          "group_index": 0,
          "mime_type": null,
          "bitrate": null,
          "retention_status": "skipped",
          "retention_reason": "not_requested",
          "asset": null
        }
      ],
      "media_count_reported": null,
      "media_complete": null,
      "geotag": null,
      "tagged_users": null,
      "mentions": [],
      "likes_count": 218,
      "comments_count": 22,
      "views_count": 31149,
      "reposts_count": 7,
      "quotes_count": 1,
      "relation_kind": "incoming_quote",
      "relation_evidence": "provider_quote_target",
      "relation_target_post_id": "2065225362544726371",
      "quoted_post_id": "2065225362544726371",
      "in_reply_to_post_id": null,
      "conversation_id": "2065257221077020850",
      "is_reply": null,
      "is_quote": true
    }
  ],
  "count": 2,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": 2,
  "collected_count": 50,
  "collection_limit": 500,
  "expires_at": 1791154800,
  "truncated": true,
  "partial": false,
  "target_post_id": "2065225362544726371",
  "requested_collection_size": 50,
  "quotes_complete": null,
  "upstream_continuation_available": false,
  "partial_reasons": []
}

Collect incoming quote posts of the requested original, verified through actual quoted-post IDs. Separate quote posts by one author are retained. Missing author IDs or reply parents remain null. Collection size is 1–500; there is no upstream continuation beyond the cached result. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/x/profile/spaces

Request headers
{
  "Idempotency-Key": "hive-doc-x-profile-spaces",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/profile/spaces" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-x-profile-spaces" \
  -H "Prefer: respond-async" \
  --data '{"limit":2,"profile":"barkmeta","collection_size":5}'
JSON request body
{
  "limit": 2,
  "profile": "barkmeta",
  "collection_size": 5
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "x",
  "operation": "profile/spaces",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "2106510350287884296",
      "post_url": "https://x.com/barkmeta/status/2106510350287884296",
      "timestamp": "2026-10-03T22:22:59+00:00",
      "text": "https://t.co/dsIBjaUJT0",
      "shared_by": "barkmeta",
      "space_urls": [
        "https://x.com/i/spaces/1nJOLQwXDwVxR"
      ],
      "relation_evidence": "authored_post_space_link",
      "host_relationship": null
    },
    {
      "post_id": "2106489814304612553",
      "post_url": "https://x.com/barkmeta/status/2106489814304612553",
      "timestamp": "2026-10-03T21:01:23+00:00",
      "text": "https://t.co/DENPHq5NQa",
      "shared_by": "barkmeta",
      "space_urls": [
        "https://x.com/i/spaces/1rGmqprjLZLGy"
      ],
      "relation_evidence": "authored_post_space_link",
      "host_relationship": null
    }
  ],
  "count": 2,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": 2,
  "collected_count": 5,
  "collection_limit": 100,
  "expires_at": 1791154800,
  "truncated": true,
  "partial": false,
  "collection_complete": null,
  "partial_reasons": [],
  "source_profile": "barkmeta",
  "requested_collection_size": 5,
  "effective_query": "from:barkmeta filter:spaces",
  "search_complete": null,
  "discovery_scope": "authored_space_links"
}

Discover Space links in authored posts by a handle. shared_by identifies the posting account; sharing a Space does not prove hosting. Use each explicit Space ID for recording lookup. Collection size is 1–100. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/x/space/recording

Request headers
{
  "Idempotency-Key": "hive-doc-x-space-recording",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/space/recording" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-x-space-recording" \
  -H "Prefer: respond-async" \
  --data '{"limit":1,"space_id":"1MJgNbvMojbGL"}'
JSON request body
{
  "limit": 1,
  "space_id": "1MJgNbvMojbGL"
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "x",
  "operation": "space/recording",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "space_id": "1MJgNbvMojbGL",
      "space_url": "https://x.com/i/spaces/1MJgNbvMojbGL",
      "title": null,
      "state": "ended",
      "started_at": "2026-09-30T21:02:13.104000+00:00",
      "ended_at": "2026-10-01T00:01:58.749000+00:00",
      "creator": {
        "username": "barkmeta",
        "profile_url": "https://x.com/barkmeta",
        "profile_id": "336348053",
        "display_name": "Bark",
        "image_url": null
      },
      "hosts": [
        {
          "username": "barkmeta",
          "profile_url": "https://x.com/barkmeta",
          "profile_id": "336348053",
          "display_name": "Bark",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "shieldmetax",
          "profile_url": "https://x.com/shieldmetax",
          "profile_id": "3588771076",
          "display_name": "Shield",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "cornkry",
          "profile_url": "https://x.com/cornkry",
          "profile_id": "970795858190045187",
          "display_name": "Corn",
          "image_url": "https://cdn.example.com/image.jpg"
        }
      ],
      "speakers": [
        {
          "username": "godsburnt",
          "profile_url": "https://x.com/godsburnt",
          "profile_id": "1547753041092231168",
          "display_name": "Shibo",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "kingfud",
          "profile_url": "https://x.com/kingfud",
          "profile_id": "556601889",
          "display_name": "FUD",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "web3smb",
          "profile_url": "https://x.com/web3smb",
          "profile_id": "832980412452503556",
          "display_name": "Web",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "veemeta",
          "profile_url": "https://x.com/veemeta",
          "profile_id": "32831485",
          "display_name": "Vee",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "hofers",
          "profile_url": "https://x.com/hofers",
          "profile_id": "1053309451",
          "display_name": "Hofer",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "leahbluewater",
          "profile_url": "https://x.com/leahbluewater",
          "profile_id": "555216272",
          "display_name": "Leah",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "doskyart",
          "profile_url": "https://x.com/doskyart",
          "profile_id": "1735567855964344320",
          "display_name": "Dosky🦠",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "anthemhayek",
          "profile_url": "https://x.com/anthemhayek",
          "profile_id": "22429454",
          "display_name": "AnTheM",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "artsymeta",
          "profile_url": "https://x.com/artsymeta",
          "profile_id": "4825824006",
          "display_name": "Artsy",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "cheeksmeta",
          "profile_url": "https://x.com/cheeksmeta",
          "profile_id": "1422223114038087694",
          "display_name": "Cheeks",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "motionmetax",
          "profile_url": "https://x.com/motionmetax",
          "profile_id": "1302437599005609984",
          "display_name": "Motion",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "duvalx",
          "profile_url": "https://x.com/duvalx",
          "profile_id": "1503422852246216710",
          "display_name": "Duval",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "jagobx",
          "profile_url": "https://x.com/jagobx",
          "profile_id": "1246103444378746880",
          "display_name": "Jag",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "vipmetax",
          "profile_url": "https://x.com/vipmetax",
          "profile_id": "2986504127",
          "display_name": "VIP",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "rosemetax",
          "profile_url": "https://x.com/rosemetax",
          "profile_id": "1548699110601105409",
          "display_name": "Rose",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "rafmeta",
          "profile_url": "https://x.com/rafmeta",
          "profile_id": "1528005302657892353",
          "display_name": "Raf",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "brettmetax",
          "profile_url": "https://x.com/brettmetax",
          "profile_id": "1851760662097408000",
          "display_name": "Brett",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "maymetax",
          "profile_url": "https://x.com/maymetax",
          "profile_id": "970695375530151941",
          "display_name": "May",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "barkmediafrica",
          "profile_url": "https://x.com/barkmediafrica",
          "profile_id": "2102882733143789568",
          "display_name": "BMA",
          "image_url": "https://cdn.example.com/image.jpg"
        },
        {
          "username": "notnftnews",
          "profile_url": "https://x.com/notnftnews",
          "profile_id": "128511600",
          "display_name": "Noot",
          "image_url": "https://cdn.example.com/image.jpg"
        }
      ],
      "replay_reported": true,
      "recording_url": "https://prod-fastly-us-east-1.video.pscp.tv/Transcoding/v1/hls/3yv6QtBxgVKMVFGaW0Nje7RnA-LuZ-0KZk-AsMGD6z8Ag3eWZlTEeuEewsOs6uzAU3syr4YJscPzJPtj5uOQuA/non_transcode/us-east-1/periscope-replay-direct-prod-us-east-1-public/audio-space/playlist_16655931155200980624.m3u8?type=replay",
      "recording_status": "accessible",
      "access_evidence": "playlist_and_audio_sample",
      "audio_retained": false,
      "relation_evidence": "observed_space_id"
    }
  ],
  "count": 1,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 1,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 1,
  "expires_at": 1791154800,
  "truncated": null,
  "partial": false,
  "collection_complete": null,
  "partial_reasons": [],
  "space_id": "1MJgNbvMojbGL",
  "audio_retained": false
}

Look up one explicit public Space and its observed creator, hosts, speakers and replay reference. accessible requires an ended Space plus a fetched playlist and recognizable audio sample. This proves a bounded sample, not the whole recording or future URL availability. Audio bytes are not retained. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/x/space/transcript

Request headers
{
  "Idempotency-Key": "hive-doc-x-space-transcript",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/x/space/transcript" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-x-space-transcript" \
  -H "Prefer: respond-async" \
  --data '{"limit":1,"space_id":"1MJgNbvMojbGL","max_minutes":1}'
JSON request body
{
  "limit": 1,
  "space_id": "1MJgNbvMojbGL",
  "max_minutes": 1
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "x",
  "operation": "space/transcript",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "space_id": "1MJgNbvMojbGL",
      "space_url": "https://x.com/i/spaces/1MJgNbvMojbGL",
      "title": "THE GREAT RESET 🚨 BITCOIN WEDNESDAY",
      "transcript_status": "missing_transcript",
      "text": null,
      "language": null,
      "duration_seconds": 60,
      "transcribed_seconds": null,
      "segments": null,
      "prefix_limited": false,
      "transcript_complete": null,
      "method": null,
      "audio_retained": false,
      "relation_evidence": "source_request"
    }
  ],
  "count": 1,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 1,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 1,
  "expires_at": 1791154800,
  "truncated": null,
  "partial": true,
  "collection_complete": null,
  "partial_reasons": [
    "missing_transcript",
    "rejected_rows_or_provider_partial"
  ],
  "space_id": "1MJgNbvMojbGL",
  "requested_max_minutes": 1,
  "audio_retained": false
}

Request a finite speech-to-text prefix of one public Space, max_minutes 1–30. Missing speech is preserved as a partial missing_transcript result and does not qualify operation health. The tested source returned no speech, so transcript extraction remains unqualified. There is no complete-recording transcript guarantee. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/facebook/profile/details

JSON request body
{
  "profiles": [
    "nasa"
  ]
}
Illustrative result
{
  "items": [
    {
      "username": "nasa",
      "profile_url": "https://www.facebook.com/NASA",
      "display_name": "NASA",
      "target": "nasa",
      "input_index": 0
    }
  ],
  "count": 1,
  "unresolved": []
}

POST /scrapers/facebook/profile/posts

Request headers
{
  "Idempotency-Key": "facebook-profile-posts-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/profile/posts" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facebook-profile-posts-001" \
  -H "Prefer: respond-async" \
  --data '{"profiles":["nasa"],"limit":2}'
JSON request body
{
  "profiles": [
    "nasa"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "facebook",
  "operation": "profile/posts",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "target": "nasa",
      "input_index": 0,
      "post_id": "123",
      "url": "https://www.facebook.com/nasa/posts/123",
      "author": {
        "username": "nasa",
        "profile_url": "https://www.facebook.com/nasa"
      },
      "timestamp": "2026-09-01T12:00:00+00:00",
      "text": "Public post example.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "context": []
    }
  ],
  "count": 1,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 500,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": false,
  "targets": [
    {
      "target": "nasa",
      "input_index": 0,
      "status": "resolved",
      "collected_count": 1,
      "reason": null
    }
  ],
  "unresolved": [],
  "per_target_limit": 40,
  "history_complete": null,
  "history_scope": "recent_public_timeline"
}

POST /scrapers/facebook/people/search

Request headers
{
  "Idempotency-Key": "facebook-people-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/people/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facebook-people-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"José García","locations":["Madrid"],"limit":2}'
JSON request body
{
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "facebook",
  "operation": "people/search",
  "status": "pending",
  "created_at": 1789905600,
  "expires_at": 1789992000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "query": "José García",
      "rank": 1,
      "status": "candidate",
      "username": null,
      "profile_url": "https://www.facebook.com/profile.php?id=123456",
      "profile_id": "123456",
      "display_name": "José García",
      "biography": null,
      "image_url": "https://images.example.com/avatar.jpg",
      "missing_fields": [
        "biography"
      ],
      "profile_type": "person",
      "location": "Madrid"
    }
  ],
  "count": 1,
  "result_id": "rresult_22222222222222222222222222222222",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1789992000,
  "truncated": null,
  "partial": true,
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "effective_query": "José García Madrid",
  "location_handling": "query_hint",
  "filters_enforced": false
}

POST /scrapers/facebook/pages/search

Request headers
{
  "Idempotency-Key": "facebook-pages-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/pages/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facebook-pages-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"space science","collection_size":50,"limit":2}'
JSON request body
{
  "query": "space science",
  "collection_size": 50,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "facebook",
  "operation": "pages/search",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "username": "examplepage",
      "profile_url": "https://www.facebook.com/examplepage",
      "profile_id": "123",
      "profile_type": "page",
      "query": "space science",
      "biography": null,
      "location": null,
      "website_url": null,
      "is_verified": null
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 50,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "query": "space science",
  "query_handling": "upstream_hint",
  "filters_enforced": false
}

Synthetic public page candidate, with nullable observed metadata. Only page rows are accepted; a page is neither a group nor a verified identity. The query is an upstream hint and matching/recall are unknown. Cached next_page:null does not establish collection completeness.

POST /scrapers/facebook/posts/search

Request headers
{
  "Idempotency-Key": "facebook-posts-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/posts/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facebook-posts-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"space science","keyword":"science","since_date":"2020-01-01","until_date":"2021-01-01","collection_size":50,"limit":2}'
JSON request body
{
  "query": "space science",
  "keyword": "science",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "collection_size": 50,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "facebook",
  "operation": "posts/search",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "456",
      "url": "https://www.facebook.com/examplepage/posts/456",
      "author": {
        "username": "examplepage",
        "profile_url": "https://www.facebook.com/examplepage"
      },
      "timestamp": "2020-06-01T12:00:00+00:00",
      "text": "Synthetic science post.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "query": "space science",
      "relation_evidence": "search_query"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 50,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "keyword": "science",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "local_filters": [
    "keyword",
    "since_date",
    "until_date"
  ],
  "query": "space science",
  "query_handling": "upstream_hint",
  "group_url": null,
  "post_url": null
}

Synthetic public post search. query is an upstream hint; keyword is a literal case-insensitive substring checked locally, since_date is inclusive and until_date exclusive in UTC. Rows missing fields required by a requested local filter are rejected and the collection remains partial. Search matching and recall remain unknown; no extra query follows filtering.

POST /scrapers/facebook/group/posts

Request headers
{
  "Idempotency-Key": "facebook-group-posts-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/group/posts" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facebook-group-posts-001" \
  -H "Prefer: respond-async" \
  --data '{"group_url":"https://www.facebook.com/groups/publicgroup","keyword":"science","since_date":"2020-01-01","until_date":"2021-01-01","collection_size":50,"limit":2}'
JSON request body
{
  "group_url": "https://www.facebook.com/groups/publicgroup",
  "keyword": "science",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "collection_size": 50,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "facebook",
  "operation": "group/posts",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "post_id": "456",
      "url": "https://www.facebook.com/groups/publicgroup/posts/456",
      "author": {
        "username": "exampleperson",
        "profile_url": "https://www.facebook.com/exampleperson"
      },
      "timestamp": "2020-06-01T12:00:00+00:00",
      "text": "Synthetic science post.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "group_url": "https://www.facebook.com/groups/publicgroup",
      "group_title": null,
      "relation_evidence": "group_permalink"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 50,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "keyword": "science",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "local_filters": [
    "keyword",
    "since_date",
    "until_date"
  ],
  "query": null,
  "query_handling": null,
  "group_url": "https://www.facebook.com/groups/publicgroup",
  "post_url": null
}

Synthetic chronological collection from one explicit public group URL. Matching observed group permalinks establish membership. Keyword/date filters are checked locally with inclusive since_date and exclusive until_date; a lower date can narrow the upstream crawl, while filtering never extends its finite window. Private groups are unsupported and collection completeness is unknown.

POST /scrapers/facebook/post/comments

Request headers
{
  "Idempotency-Key": "facebook-post-comments-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/post/comments" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: facebook-post-comments-001" \
  -H "Prefer: respond-async" \
  --data '{"post_url":"https://www.facebook.com/examplepage/posts/456","keyword":"science","since_date":"2020-01-01","until_date":"2021-01-01","collection_size":100,"limit":2}'
JSON request body
{
  "post_url": "https://www.facebook.com/examplepage/posts/456",
  "keyword": "science",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "collection_size": 100,
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "facebook",
  "operation": "post/comments",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "comment_id": "789",
      "url": "https://www.facebook.com/examplepage/posts/456?comment_id=789",
      "post_url": "https://www.facebook.com/examplepage/posts/456",
      "author": {
        "username": null,
        "profile_url": null,
        "profile_id": "987",
        "display_name": null,
        "image_url": null
      },
      "timestamp": "2020-06-01T13:00:00+00:00",
      "text": "Synthetic science comment.",
      "likes_count": null,
      "parent_comment_id": null,
      "comment_scope": "top_level_candidate",
      "relation_evidence": "comment_permalink"
    }
  ],
  "count": 1,
  "result_id": "rresult_33333333333333333333333333333333",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 100,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": true,
  "requested_collection_size": 100,
  "collection_complete": null,
  "partial_reasons": [
    "rejected_rows"
  ],
  "keyword": "science",
  "since_date": "2020-01-01",
  "until_date": "2021-01-01",
  "local_filters": [
    "keyword",
    "since_date",
    "until_date"
  ],
  "query": null,
  "query_handling": null,
  "group_url": null,
  "post_url": "https://www.facebook.com/examplepage/posts/456"
}

Synthetic comment with an observed parent-post permalink. Unknown depth stays a top-level candidate and parent_comment_id remains null. No nested replies or full comment-thread claim is made. Missing commenter URLs remain null rather than becoming fabricated links. Local literal keyword and UTC date filters use the same rules as posts.

POST /scrapers/facebook/groups/search

Request headers
{
  "Idempotency-Key": "hive-doc-facebook-groups-search",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/facebook/groups/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-facebook-groups-search" \
  -H "Prefer: respond-async" \
  --data '{"limit":2,"query":"Argentinos en Melbourne","collection_size":5}'
JSON request body
{
  "limit": 2,
  "query": "Argentinos en Melbourne",
  "collection_size": 5
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "facebook",
  "operation": "groups/search",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "group_id": "361701421159198",
      "group_url": "https://www.facebook.com/groups/361701421159198",
      "name": "ARGENTINOS EN MELBOURNE -  SIN FILTROS -",
      "vanity": null,
      "privacy": "public",
      "visibility": "visible",
      "members_count": 1390,
      "description": "COMUNIDAD ARGENTOS.\nEste es grupo creado con el propósito de dar un lugar de encuentro a los Argentinos que vivimos en  Australia.\nPostear datos útiles que nos sirvan a todos, cosas curiosas, el que quiera invitar a todos para un encuent...",
      "location": "Melbourne, Victoria, Australia",
      "image_url": "https://cdn.example.com/image.jpg",
      "posts_today": 2,
      "posts_last_month": 34,
      "average_posts_per_day": 1.1,
      "created_at": "2019-07-07T01:32:11+00:00",
      "admin_moderator_count": 4,
      "query": "Argentinos en Melbourne",
      "content_retrieved": false,
      "relation_evidence": "public_group_profile"
    },
    {
      "group_id": "2094162451204965",
      "group_url": "https://www.facebook.com/groups/2094162451204965",
      "name": "Argentinos en Melbourne",
      "vanity": null,
      "privacy": "public",
      "visibility": "visible",
      "members_count": 3,
      "description": "Grupo de Argentinos en Melbourne y también en todo Australia.",
      "location": "Melbourne, Victoria, Australia",
      "image_url": "https://cdn.example.com/image.jpg",
      "posts_today": 0,
      "posts_last_month": 0,
      "average_posts_per_day": 0,
      "created_at": "2026-07-09T21:25:36+00:00",
      "admin_moderator_count": 1,
      "query": "Argentinos en Melbourne",
      "content_retrieved": false,
      "relation_evidence": "public_group_profile"
    }
  ],
  "count": 2,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": 2,
  "collected_count": 5,
  "collection_limit": 200,
  "expires_at": 1791154800,
  "truncated": null,
  "partial": false,
  "query": "Argentinos en Melbourne",
  "requested_collection_size": 5,
  "search_source": "public_web_index",
  "query_handling": "upstream_hint",
  "filters_enforced": false,
  "collection_complete": null,
  "content_retrieved": false,
  "partial_reasons": []
}

Discover actual group IDs, names, public URLs and available metadata through a public web index. Indexed private group names do not grant content access. Matching and exhaustive recall remain unknown. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/tiktok/profile/details

JSON request body
{
  "profiles": [
    "nasa"
  ]
}
Illustrative result
{
  "items": [
    {
      "username": "nasa",
      "profile_url": "https://www.tiktok.com/@nasa",
      "display_name": "NASA",
      "target": "nasa",
      "input_index": 0
    }
  ],
  "count": 1,
  "unresolved": []
}

POST /scrapers/tiktok/profile/posts

Request headers
{
  "Idempotency-Key": "tiktok-profile-posts-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/tiktok/profile/posts" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tiktok-profile-posts-001" \
  -H "Prefer: respond-async" \
  --data '{"profiles":["nasa"],"limit":2}'
JSON request body
{
  "profiles": [
    "nasa"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "tiktok",
  "operation": "profile/posts",
  "status": "pending",
  "created_at": 1789913600,
  "expires_at": 1790000000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "target": "nasa",
      "input_index": 0,
      "post_id": "123",
      "url": "https://www.tiktok.com/@nasa/video/123",
      "author": {
        "username": "nasa",
        "profile_url": "https://www.tiktok.com/@nasa"
      },
      "timestamp": "2026-09-01T12:00:00+00:00",
      "text": "Public post example.",
      "media": [],
      "geotag": null,
      "tagged_users": null,
      "mentions": null,
      "context": []
    }
  ],
  "count": 1,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 500,
  "expires_at": 1790000000,
  "truncated": null,
  "partial": false,
  "targets": [
    {
      "target": "nasa",
      "input_index": 0,
      "status": "resolved",
      "collected_count": 1,
      "reason": null
    }
  ],
  "unresolved": [],
  "per_target_limit": 50,
  "history_complete": null,
  "history_scope": "collected_window"
}

POST /scrapers/tiktok/people/search

Request headers
{
  "Idempotency-Key": "tiktok-people-search-001",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/tiktok/people/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: tiktok-people-search-001" \
  -H "Prefer: respond-async" \
  --data '{"query":"José García","locations":["Madrid"],"limit":2}'
JSON request body
{
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "service": "tiktok",
  "operation": "people/search",
  "status": "pending",
  "created_at": 1789905600,
  "expires_at": 1789992000,
  "poll_url": "/scrapers/jobs/rjob_aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa",
  "result": null,
  "error": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "query": "José García",
      "rank": 1,
      "status": "candidate",
      "username": "example_person",
      "profile_url": "https://www.tiktok.com/@example_person",
      "profile_id": "123456",
      "display_name": "José García",
      "biography": "Synthetic public biography",
      "image_url": "https://images.example.com/avatar.jpg",
      "missing_fields": []
    }
  ],
  "count": 1,
  "result_id": "rresult_22222222222222222222222222222222",
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "collection_limit": 50,
  "expires_at": 1789992000,
  "truncated": null,
  "partial": false,
  "query": "José García",
  "locations": [
    "Madrid"
  ],
  "effective_query": "José García Madrid",
  "location_handling": "query_hint",
  "filters_enforced": false
}

POST /scrapers/sherlock/search

JSON request body
{
  "usernames": [
    "nasa"
  ],
  "limit": 20
}
Illustrative result
{
  "items": [
    {
      "username": "nasa",
      "site": "GitHub",
      "profile_url": "https://github.com/nasa",
      "status": "candidate"
    }
  ],
  "count": 1,
  "truncated": null
}

POST /scrapers/google/search

JSON request body
{
  "queries": [
    "site:nasa.gov Artemis"
  ],
  "limit": 20,
  "pages": 1,
  "country": "us",
  "language": "en"
}
Illustrative result
{
  "items": [
    {
      "query": "site:nasa.gov Artemis",
      "rank": 1,
      "title": "Artemis",
      "url": "https://www.nasa.gov/artemis/",
      "snippet": "Explore the Artemis program."
    }
  ],
  "count": 1
}

POST /scrapers/google_lens/search

JSON request body
{
  "image_url": "https://gpm.nasa.gov/sites/default/files/document_files/NASA-Logo-Large.png",
  "search_types": [
    "visual",
    "exact",
    "text"
  ],
  "limit": 20
}
Illustrative result
{
  "items": [
    {
      "type": "visual",
      "title": "NASA",
      "url": "https://www.nasa.gov/"
    },
    {
      "type": "exact",
      "title": "NASA logo",
      "url": "https://www.nasa.gov/logos/"
    },
    {
      "type": "text",
      "text": "NASA"
    }
  ],
  "count": 3
}

POST /scrapers/reddit/search

Request headers
{
  "Idempotency-Key": "hive-doc-reddit-search",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/reddit/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-reddit-search" \
  -H "Prefer: respond-async" \
  --data '{"limit":2,"query":"python web scraping","content_type":"both","sort":"relevance","time_filter":"all","subreddits":[],"include_nsfw":false,"collection_size":10}'
JSON request body
{
  "limit": 2,
  "query": "python web scraping",
  "content_type": "both",
  "sort": "relevance",
  "time_filter": "all",
  "subreddits": [],
  "include_nsfw": false,
  "collection_size": 10
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "reddit",
  "operation": "search",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "full_id": "t3_a1rp7c",
      "post_id": "a1rp7c",
      "url": "https://www.reddit.com/r/learnprogramming/comments/a1rp7c/i_made_a_python_web_scraping_guide_for_beginners/",
      "subreddit": "learnprogramming",
      "author": {
        "username": "brendanmartin",
        "profile_id": "t2_u94ni",
        "profile_url": "https://www.reddit.com/user/brendanmartin/",
        "image_url": "https://cdn.example.com/image.jpg",
        "is_deleted": null
      },
      "text": "I've been web scraping professionally for a few years and decided to make a series of web scraping tutorials that I wish I had when I started.\n\nThe series will follow a large project I'm building that analyzes political rhetoric in the n...",
      "content_state": "available",
      "timestamp": "2018-11-30T11:39:30+00:00",
      "score": 2193,
      "query": "python web scraping",
      "relation_evidence": "search_result_permalink",
      "kind": "post",
      "title": "I made a Python web scraping guide for beginners",
      "comments_count": 150,
      "outbound_url": null,
      "is_nsfw": false
    },
    {
      "full_id": "t1_pbake4z",
      "post_id": "1wmrzed",
      "url": "https://www.reddit.com/r/dataanalysis/comments/1wmrzed/web_scraping/pbake4z/",
      "subreddit": "dataanalysis",
      "author": {
        "username": "ItsSignalsJerry_",
        "profile_id": "t2_17b10vcuh9",
        "profile_url": "https://www.reddit.com/user/ItsSignalsJerry_/",
        "image_url": "https://cdn.example.com/image.jpg",
        "is_deleted": null
      },
      "text": "The point of CAPTCHA is to prevent bots. A scraper is a bot.\n\nAlso, R is not what you should be using scraping the web when Python makes it much easier.",
      "content_state": "available",
      "timestamp": "2026-09-22T02:48:33+00:00",
      "score": 18,
      "query": "python web scraping",
      "relation_evidence": "search_result_permalink",
      "kind": "comment",
      "comment_id": "pbake4z",
      "parent_id": "t3_1wmrzed",
      "parent_evidence": "post_context",
      "post_title": "Web scraping"
    }
  ],
  "count": 2,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 2,
  "next_page": 2,
  "collected_count": 10,
  "collection_limit": 200,
  "expires_at": 1791154800,
  "truncated": null,
  "partial": false,
  "requested_collection_size": 10,
  "query": "python web scraping",
  "content_type": "both",
  "sort": "relevance",
  "time_filter": "all",
  "subreddits": [],
  "collection_complete": null,
  "query_handling": "upstream_hint",
  "filters_enforced": false,
  "partial_reasons": []
}

Search public Reddit posts, comments or both with optional subreddit, sort and time hints. Returned IDs, permalinks and authors come from source rows. Native query operators are upstream hints and do not guarantee matching. Collection size is 1–200. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/reddit/post/comments

Request headers
{
  "Idempotency-Key": "hive-doc-reddit-post-comments",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/reddit/post/comments" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-reddit-post-comments" \
  -H "Prefer: respond-async" \
  --data '{"limit":3,"post_url":"https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/","sort":"best","collection_size":20}'
JSON request body
{
  "limit": 3,
  "post_url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/",
  "sort": "best",
  "collection_size": 20
}
HTTP 202 · initial job
{
  "job_id": "rjob_22222222222222222222222222222222",
  "service": "reddit",
  "operation": "post/comments",
  "status": "pending",
  "created_at": 1791068400,
  "expires_at": 1791154800,
  "poll_url": "/scrapers/jobs/rjob_22222222222222222222222222222222",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "items": [
    {
      "comment_id": "pbsphka",
      "full_id": "t1_pbsphka",
      "post_id": "1wp6ilq",
      "url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/pbsphka/",
      "post_url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/",
      "subreddit": "stocksandtrading",
      "author": {
        "username": "AutoModerator",
        "profile_id": null,
        "profile_url": "https://www.reddit.com/user/AutoModerator/",
        "image_url": null,
        "is_deleted": false
      },
      "text": "🚀 🌑  -- Join our discord!! https://discord.gg/jcewXNmf6C -- 🚀 🌑\n\nI am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.",
      "content_state": "available",
      "timestamp": "2026-09-24T16:35:25.787000+00:00",
      "edited_at": null,
      "score": null,
      "depth": 0,
      "parent_id": "t3_1wp6ilq",
      "path": [
        "t1_pbsphka"
      ],
      "parent_present_in_collection": true,
      "children_ids": [],
      "direct_reply_count": 0,
      "captured_sequence": 1,
      "relation_evidence": "observed_parent_and_path"
    },
    {
      "comment_id": "pbw6fsu",
      "full_id": "t1_pbw6fsu",
      "post_id": "1wp6ilq",
      "url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/pbw6fsu/",
      "post_url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/",
      "subreddit": "stocksandtrading",
      "author": {
        "username": "therackage",
        "profile_id": null,
        "profile_url": "https://www.reddit.com/user/therackage/",
        "image_url": null,
        "is_deleted": false
      },
      "text": "“And that gap is the whole story”\n\nThanks ChatGPT.",
      "content_state": "available",
      "timestamp": "2026-09-25T02:24:29.917000+00:00",
      "edited_at": null,
      "score": 9,
      "depth": 0,
      "parent_id": "t3_1wp6ilq",
      "path": [
        "t1_pbw6fsu"
      ],
      "parent_present_in_collection": true,
      "children_ids": [
        "t1_pbw6p8d"
      ],
      "direct_reply_count": 1,
      "captured_sequence": 2,
      "relation_evidence": "observed_parent_and_path"
    },
    {
      "comment_id": "pbw6p8d",
      "full_id": "t1_pbw6p8d",
      "post_id": "1wp6ilq",
      "url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/pbw6p8d/",
      "post_url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/",
      "subreddit": "stocksandtrading",
      "author": {
        "username": "GrahamPhisher",
        "profile_id": null,
        "profile_url": "https://www.reddit.com/user/GrahamPhisher/",
        "image_url": null,
        "is_deleted": false
      },
      "text": "Local LLM powered by a B70 after processing an entire night of stock analysis and web scraping from a python program with a 58% win rate*\n\nAnd yes the gap is the story, that's what traders should be looking to leverage.",
      "content_state": "available",
      "timestamp": "2026-09-25T02:25:53.698000+00:00",
      "edited_at": null,
      "score": 3,
      "depth": 1,
      "parent_id": "t1_pbw6fsu",
      "path": [
        "t1_pbw6fsu",
        "t1_pbw6p8d"
      ],
      "parent_present_in_collection": true,
      "children_ids": [],
      "direct_reply_count": 0,
      "captured_sequence": 3,
      "relation_evidence": "observed_parent_and_path"
    }
  ],
  "count": 3,
  "result_id": "rresult_11111111111111111111111111111111",
  "page": 1,
  "limit": 3,
  "next_page": 2,
  "collected_count": 4,
  "collection_limit": 1000,
  "expires_at": 1791154800,
  "truncated": null,
  "partial": false,
  "requested_collection_size": 20,
  "post_url": "https://www.reddit.com/r/stocksandtrading/comments/1wp6ilq/the_bullish_case_for_kraken_robotics_the/",
  "post_id": "1wp6ilq",
  "sort": "best",
  "thread_status": "observed",
  "collection_complete": null,
  "tree_complete": true,
  "max_depth": 100,
  "missing_ancestor_ids": [],
  "nested_replies_requested": true,
  "collapsed_replies_requested": true,
  "partial_reasons": []
}

Expand accessible nested and collapsed replies with actual parent IDs, paths and deleted placeholders. A parent may appear on another cached page. tree_complete refers only to the accessible source tree when supported by its summary; null is unknown. Deleted text and private comments are not recovered. Collection size is 1–1000, with maximum depth 100. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/telegram/groups/search

Request headers
{
  "Idempotency-Key": "hive-doc-telegram-groups-search",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/telegram/groups/search" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-telegram-groups-search" \
  -H "Prefer: respond-async" \
  --data '{"query":"python","language":null,"limit":2,"collection_size":5}'
JSON request body
{
  "query": "python",
  "language": null,
  "limit": 2,
  "collection_size": 5
}
HTTP 202 · initial job
{
  "job_id": "rjob_00000000000000000000000000000001",
  "service": "telegram",
  "operation": "groups/search",
  "status": "pending",
  "created_at": 1791493200,
  "expires_at": 1791579600,
  "poll_url": "/scrapers/jobs/rjob_00000000000000000000000000000001",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "count": 1,
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "expires_at": 1791579600,
  "truncated": null,
  "partial": false,
  "collection_complete": null,
  "partial_reasons": [],
  "items": [
    {
      "group_url": "https://t.me/pythongroup",
      "username": "pythongroup",
      "title": "Python group",
      "description": "Illustrative public group description",
      "query": "python",
      "rank": 1,
      "matched_all_words": true,
      "language": null,
      "observed_at": "2026-10-08T12:00:00+00:00",
      "relation_evidence": "public_group_index"
    }
  ],
  "result_id": "rresult_00000000000000000000000000000001",
  "collection_limit": 200,
  "requested_collection_size": 5,
  "query": "python",
  "language": null,
  "search_source": "public_web_index",
  "query_handling": "upstream_hint",
  "filters_enforced": false,
  "content_retrieved": false
}

Discover public discussion-group candidates by topic or name. Public index results are not exhaustive; discovery collects neither messages nor member lists. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/telegram/group/details

Request headers
{
  "Idempotency-Key": "hive-doc-telegram-group-details",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/telegram/group/details" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-telegram-group-details" \
  -H "Prefer: respond-async" \
  --data '{"group_url":"https://t.me/pythontelegrambotgroup","limit":2}'
JSON request body
{
  "group_url": "https://t.me/pythontelegrambotgroup",
  "limit": 2
}
HTTP 202 · initial job
{
  "job_id": "rjob_00000000000000000000000000000002",
  "service": "telegram",
  "operation": "group/details",
  "status": "pending",
  "created_at": 1791493200,
  "expires_at": 1791579600,
  "poll_url": "/scrapers/jobs/rjob_00000000000000000000000000000002",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "count": 1,
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "expires_at": 1791579600,
  "truncated": null,
  "partial": false,
  "collection_complete": null,
  "partial_reasons": [],
  "items": [
    {
      "group_url": "https://t.me/pythontelegrambotgroup",
      "username": "pythontelegrambotgroup",
      "title": "Python Telegram Bot discussion",
      "description": "Illustrative public group description",
      "members_count": null,
      "online_count": null,
      "is_verified": null,
      "image_url": null,
      "observed_at": "2026-10-08T12:00:00+00:00"
    }
  ],
  "result_id": "rresult_00000000000000000000000000000002",
  "collection_limit": 1,
  "requested_collection_size": 1,
  "group_url": "https://t.me/pythontelegrambotgroup"
}

Read metadata for one public discussion group. Channels and bots are not returned as groups. Hidden previews and source failures do not prove an empty group. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

POST /scrapers/telegram/group/messages

Request headers
{
  "Idempotency-Key": "hive-doc-telegram-group-messages",
  "Prefer": "respond-async"
}
Bash · submit with your Hive key
curl --fail-with-body --max-time 30 -sS "https://hive.arglegal.live/scrapers/telegram/group/messages" \
  -H "Authorization: Bearer $HIVE_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: hive-doc-telegram-group-messages" \
  -H "Prefer: respond-async" \
  --data '{"group_url":"https://t.me/pythontelegrambotgroup","limit":2,"collection_size":5,"keyword":null,"since":null,"until":null}'
JSON request body
{
  "group_url": "https://t.me/pythontelegrambotgroup",
  "limit": 2,
  "collection_size": 5,
  "keyword": null,
  "since": null,
  "until": null
}
HTTP 202 · initial job
{
  "job_id": "rjob_00000000000000000000000000000003",
  "service": "telegram",
  "operation": "group/messages",
  "status": "pending",
  "created_at": 1791493200,
  "expires_at": 1791579600,
  "poll_url": "/scrapers/jobs/rjob_00000000000000000000000000000003",
  "result": null,
  "error": null,
  "retry_after_s": null
}
Completed page · poll result or HTTP 200
{
  "count": 1,
  "page": 1,
  "limit": 2,
  "next_page": null,
  "collected_count": 1,
  "expires_at": 1791579600,
  "truncated": null,
  "partial": false,
  "collection_complete": null,
  "partial_reasons": [],
  "items": [
    {
      "group_url": "https://t.me/pythontelegrambotgroup",
      "message_id": "1",
      "url": "https://t.me/pythontelegrambotgroup/1",
      "text": "Illustrative public discussion message",
      "timestamp": "2026-10-08T12:00:00+00:00",
      "edited": false,
      "sender": {
        "name": null,
        "username": null,
        "profile_url": null
      },
      "reply_to_message_id": null,
      "reply_to_url": null,
      "media": [],
      "relation_evidence": "observed_group_message"
    }
  ],
  "result_id": "rresult_00000000000000000000000000000003",
  "collection_limit": 200,
  "requested_collection_size": 5,
  "group_url": "https://t.me/pythontelegrambotgroup",
  "keyword": null,
  "since": null,
  "until": null,
  "history_scope": "public_accessible_snapshot",
  "history_complete": null,
  "filters_enforced": false
}

Collect up to 200 accessible group messages with observed IDs, UTC timestamps, public senders, reply links and media references. Keyword and time filters are source hints. Full history stays unknown; media references can expire and bytes are not retained. These are illustrative cached snapshots; examples do not establish current operation readiness or promise complete coverage.

Books

Query. Inspect. Download.

Submit POST /books/queries, then poll the opaque query ID. Inspect results through GET /books/items/{book_id}. Submit POST /books/downloads, poll the download, and retrieve its artifact when ready.

Query and download submissions require an Idempotency-Key. Availability is reported separately in /capabilities; configuration alone does not establish a working download.

POST /books/queries · body
{
  "query": "Pride and Prejudice",
  "limit": 10
}

Use a returned book_id and one of its available formats in the download body: {"book_id":"<returned-book-id>","format":"epub"}.

Reference

Manage active sessions

Keep a connection open

Renew before lease_expires_at to keep using the same route and credentials. Browser sessions can run for up to one hour; max_expires_at gives the final deadline. Create a new session once the current one expires.

Change a tunnel’s IP address

Use rotate when you need a new exit IP. Hive checks the replacement against your requested country, region, ASN, and network type. If no matching replacement is available, your current connection stays in place.

Reuse a network identity

Pass the same identity_token to request a sticky allocation. A request that cannot preserve that identity fails rather than assigning a different one. Use mobile: true; non-mobile allocations are currently unavailable.

Connect securely

Browser WebSocket connections use your Hive bearer key. Proxy connections use the separate username and password returned when you create a tunnel. Save those credentials from the creation response.

SOCKS5 and HTTP proxy authentication travels unencrypted. HTTPS CONNECT encrypts traffic to the destination after the proxy handshake; it does not encrypt the proxy credentials.

Errors

Errors and retries

400 invalid request · 401 invalid Hive authentication · 403 forbidden · 404 not found · 409 conflicting or expired state · 429 rate limited · 503 service or constraints unavailable · 504 timeout.

Errors expose stable error codes. Do not blindly retry allocations or timed-out research calls. Reuse an idempotency key where supported, and inspect an existing resource before creating another.

API reference

Requests and responses.

All endpoints use the same base URL. Fields marked required must be supplied.

POST /books/downloads

Submit Download

Parameters

Idempotency-Key header · optional

JSON body

See the lifecycle example for this operation.

Response · 200

Successful Response

Errors

  • HTTP 422 · Validation Error

GET /books/items/{book_id}

Get Item

Parameters

book_id path · required

Response · 200

Successful Response

Errors

  • HTTP 422 · Validation Error

POST /books/queries

Submit Query

Parameters

Idempotency-Key header · optional

JSON body

See the lifecycle example for this operation.

Response · 200

Successful Response

Errors

  • HTTP 422 · Validation Error

POST /browser/capture

Capture public URL text and a full-page PNG

Create an isolated temporary browser, navigate once, return bounded current text and a full-page PNG, and close browser and egress ownership. Defaults to direct internet access through the public-only Relay guard; use_proxy=true opts into managed proxy access. Does not reuse profiles, log in, solve challenges, or guarantee access. navigation_status reports the target HTTP response, including denial. A full-page image exceeding 16 million pixels or 4 MiB fails; text_truncated reports text limits.

JSON body

asn string | integer | null · optional

Default: null.

browser string · optional

Pattern: ^(chrome|firefox)$. Default: "chrome".

country string · optional

Pattern: ^[A-Za-z]{2}$. Default: "US".

identity_token string | null · optional

Default: null. Minimum length: 1. Maximum length: 128.

max_text_chars integer · optional

Minimum: 1. Maximum: 100000. Default: 20000.

mobile boolean · optional

Default: true.

region string | null · optional

Default: null. Minimum length: 1. Maximum length: 100.

timeout_s integer · optional

Minimum: 5. Maximum: 60. Default: 30.

url string · required

Minimum length: 1. Maximum length: 4096.

use_proxy boolean · optional

Default: false.

wait_ms integer · optional

Minimum: 0. Maximum: 5000. Default: 500.

window_size string · optional

Pattern: ^[1-9][0-9]{2,3}x[1-9][0-9]{2,3}$. Default: "1280x720".

Response · 200

Success

blocked_requests integer · required

Minimum: 0.

navigation_status integer | null · required

Minimum: 100. Maximum: 599.

screenshot object · required
screenshot.data_base64 string · required

Minimum length: 1. Maximum length: 5592408.

screenshot.full_page boolean · required
screenshot.height integer · required

Minimum: 1. Maximum: 16384.

screenshot.media_type string · required

Pattern: ^image/png$.

screenshot.sha256 string · required

Pattern: ^[a-f0-9]{64}$.

screenshot.size_bytes integer · required

Minimum: 1. Maximum: 4194304.

screenshot.width integer · required

Minimum: 1. Maximum: 16384.

text string · required

Maximum length: 100000.

text_truncated boolean · required
title string · required

Maximum length: 512.

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 429 · Capture capacity unavailable
  • HTTP 503 · Requested service or constraints unavailable
  • HTTP 504 · Capture deadline exceeded

GET /browser/profiles

List your browser profiles

Parameters

limit query · optional
offset query · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /browser/profiles

Generate a persistent browser identity

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

browser chrome | firefox | null · optional

Default: null.

network object | null · optional

Default: null.

Response · 200

Success

browser chrome | firefox · required
created_at string · required
engine donut | camoufox · required
id string · required
network object | null · optional

Default: null.

runtime_version string | null · required
schema_version integer · required
status string · required
updated_at string · required

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

DELETE /browser/profiles/{profile_id}

Delete an idle browser profile

Parameters

profile_id path · required

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /browser/profiles/{profile_id}

Inspect your browser profile

Parameters

profile_id path · required

Response · 200

Success

browser chrome | firefox · required
created_at string · required
engine donut | camoufox · required
id string · required
network object | null · optional

Default: null.

runtime_version string | null · required
schema_version integer · required
status string · required
updated_at string · required

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

DELETE /browser/session

Close a browser session

Parameters

session_id query · required

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /browser/session

Inspect a browser session

Parameters

session_id query · required

Response · 200

Success

automation_protocol cdp | playwright · required
browser chrome | firefox · optional
cdp_url string · optional
engine donut | camoufox · optional
lease_expires_at string · optional

Format: date-time.

max_expires_at string · optional

Fixed one-hour session deadline. Renewal cannot extend it. Format: date-time.

profile object · optional
profile.browser chrome | firefox · required
profile.created_at string · required
profile.engine donut | camoufox · required
profile.id string · required
profile.network object | null · optional

Default: null.

profile.runtime_version string | null · required
profile.schema_version integer · required
profile.status string · required
profile.updated_at string · required
profile_id string · optional

Pattern: ^[0-9a-f]{32}$.

session_id string · required
status string · optional
viewer_url string · optional
ws_endpoint string · optional

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /browser/session

Create a persistent browser session

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

asn string | integer | null · optional

Default: null.

browser chrome | firefox | null · optional

Default: null.

country string · optional

Proxy exit country when use_proxy is true. Does not change the exit IP of a direct session. Pattern: ^[A-Za-z]{2}$. Default: "US".

identity_token string | null · optional

Default: null. Minimum length: 1. Maximum length: 128.

mobile boolean · optional

Default: true.

profile_id string | null · optional

Default: null. Pattern: ^[0-9a-f]{32}$.

region string | null · optional

Default: null. Minimum length: 1. Maximum length: 100.

use_proxy boolean · optional

Opt into managed proxy access. Omit to inherit a profile's saved network; otherwise access is direct. False conflicts with a network-bound profile. Default: false.

window_size string · optional

Pattern: ^[1-9][0-9]{2,3}x[1-9][0-9]{2,3}$. Default: "1280x720".

Response · 200

Success

automation_protocol cdp | playwright · required
browser chrome | firefox · optional
cdp_url string · optional
engine donut | camoufox · optional
lease_expires_at string · optional

Format: date-time.

max_expires_at string · optional

Fixed one-hour session deadline. Renewal cannot extend it. Format: date-time.

profile object · optional
profile.browser chrome | firefox · required
profile.created_at string · required
profile.engine donut | camoufox · required
profile.id string · required
profile.network object | null · optional

Default: null.

profile.runtime_version string | null · required
profile.schema_version integer · required
profile.status string · required
profile.updated_at string · required
profile_id string · optional

Pattern: ^[0-9a-f]{32}$.

session_id string · required
status string · optional
viewer_url string · optional
ws_endpoint string · optional

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /browser/session/renew

Extend the existing browser lease

Parameters

session_id query · required

Response · 200

Success

automation_protocol cdp | playwright · required
browser chrome | firefox · optional
cdp_url string · optional
engine donut | camoufox · optional
lease_expires_at string · optional

Format: date-time.

max_expires_at string · optional

Fixed one-hour session deadline. Renewal cannot extend it. Format: date-time.

profile object · optional
profile.browser chrome | firefox · required
profile.created_at string · required
profile.engine donut | camoufox · required
profile.id string · required
profile.network object | null · optional

Default: null.

profile.runtime_version string | null · required
profile.schema_version integer · required
profile.status string · required
profile.updated_at string · required
profile_id string · optional

Pattern: ^[0-9a-f]{32}$.

session_id string · required
status string · optional
viewer_url string · optional
ws_endpoint string · optional

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /capabilities

Discover services and supported constraints

Inspect each scraper's operation_status and operation_health. Each health entry has reason (null, rate_limited, timeout, invalid_response, capacity_limited, service_unavailable, not_found, or invalid_request), observed_at (Unix seconds or null), and retry_after_s (remaining minimum recheck delay or null when unknown). Capacity is shared across operations on this deployment. GET /scrapers/capacity reports availability and retry guidance. Hive manages service recovery; clients do not manage its upstream accounts. Google Lens private image jobs require scrapers.google_lens.private_image_ready=true. Discovery never starts research. An elapsed delay does not prove recovery; recover existing jobs with the original idempotency key.

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /exchange/image

Generate image artifacts

JSON body

count integer · optional

Minimum: 1. Maximum: 4.

prompt string · required

Minimum length: 1.

size string · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /exchange/llm

Generate a completion

JSON body

max_tokens integer · optional

Minimum: 1. Maximum: 16384.

messages object[] · required

Minimum items: 1.

messages[].content string · required
messages[].role string · required
response_format object · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /healthz

Read service availability

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /network/pin

Discover an available geographic and ASN pin

Parameters

country query · required
region query · optional
asn query · optional
mobile query · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

DELETE /network/tunnels

Revoke all tunnels owned by this API key

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /network/tunnels

Create a verified network tunnel

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

asn string | integer | null · optional

Default: null.

country string · required

Pattern: ^[A-Za-z]{2}$.

identity_token string | null · optional

Default: null. Minimum length: 1. Maximum length: 128.

idle_timeout_s integer | null · optional

Default: null. Minimum: 1.

max_bytes integer | null · optional

Default: null. Minimum: 1.

max_connections integer | null · optional

Default: null. Minimum: 1.

mobile boolean · optional

Default: true.

region string | null · optional

Default: null. Minimum length: 1. Maximum length: 100.

source_ip string | null · optional

Default: null.

ttl_s integer | null · optional

Default: null. Minimum: 1.

Response · 200

Success

created_at string · optional

Format: date-time.

credentials object · optional
credentials.password string · required
credentials.username string · required
endpoints object · optional
endpoints.http object · optional
endpoints.http.host string · required
endpoints.http.port integer · required
endpoints.http.scheme string · required
endpoints.socks5 object · optional
endpoints.socks5.host string · required
endpoints.socks5.port integer · required
endpoints.socks5.scheme string · required
id string · optional
lease_expires_at string · optional

Format: date-time.

limits object · optional
limits.idle_timeout_s integer · optional
limits.max_bytes integer · optional
limits.max_connections integer · optional
limits.max_ttl_s integer · optional
network object · optional
network.connection_type string · optional
network.exit object · optional
network.exit.asn string · optional
network.exit.country string · optional
network.exit.hosting boolean · optional
network.exit.ip string · optional
network.exit.isp string · optional
network.exit.mobile boolean · optional
network.exit.region string · optional
network.exit.verified_at string · optional

Format: date-time.

network.quality object · optional
network.quality.verified boolean · optional
network.rotatable boolean · optional
network.rotation object · optional
network.rotation.rotated boolean · optional
source_bound boolean · optional
status string · optional
terminal_reason string | null · optional
updated_at string · optional

Format: date-time.

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

DELETE /network/tunnels/{tunnel_id}

Revoke a tunnel

Parameters

tunnel_id path · required

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /network/tunnels/{tunnel_id}

Inspect a tunnel

Parameters

tunnel_id path · required

Response · 200

Success

created_at string · optional

Format: date-time.

endpoints object · optional
endpoints.http object · optional
endpoints.http.host string · required
endpoints.http.port integer · required
endpoints.http.scheme string · required
endpoints.socks5 object · optional
endpoints.socks5.host string · required
endpoints.socks5.port integer · required
endpoints.socks5.scheme string · required
id string · optional
lease_expires_at string · optional

Format: date-time.

limits object · optional
limits.idle_timeout_s integer · optional
limits.max_bytes integer · optional
limits.max_connections integer · optional
limits.max_ttl_s integer · optional
network object · optional
network.connection_type string · optional
network.exit object · optional
network.exit.asn string · optional
network.exit.country string · optional
network.exit.hosting boolean · optional
network.exit.ip string · optional
network.exit.isp string · optional
network.exit.mobile boolean · optional
network.exit.region string · optional
network.exit.verified_at string · optional

Format: date-time.

network.quality object · optional
network.quality.verified boolean · optional
network.rotatable boolean · optional
network.rotation object · optional
network.rotation.rotated boolean · optional
source_bound boolean · optional
status string · optional
terminal_reason string | null · optional
updated_at string · optional

Format: date-time.

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /network/tunnels/{tunnel_id}/credentials

Replace tunnel credentials

Parameters

tunnel_id path · required

Response · 200

Success

credentials object · optional
credentials.password string · required
credentials.username string · required
id string · optional

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /network/tunnels/{tunnel_id}/renew

Extend the lease without changing the exit

Parameters

tunnel_id path · required

JSON body

ttl_s integer · optional

Minimum: 1.

Response · 200

Success

created_at string · optional

Format: date-time.

endpoints object · optional
endpoints.http object · optional
endpoints.http.host string · required
endpoints.http.port integer · required
endpoints.http.scheme string · required
endpoints.socks5 object · optional
endpoints.socks5.host string · required
endpoints.socks5.port integer · required
endpoints.socks5.scheme string · required
id string · optional
lease_expires_at string · optional

Format: date-time.

limits object · optional
limits.idle_timeout_s integer · optional
limits.max_bytes integer · optional
limits.max_connections integer · optional
limits.max_ttl_s integer · optional
network object · optional
network.connection_type string · optional
network.exit object · optional
network.exit.asn string · optional
network.exit.country string · optional
network.exit.hosting boolean · optional
network.exit.ip string · optional
network.exit.isp string · optional
network.exit.mobile boolean · optional
network.exit.region string · optional
network.exit.verified_at string · optional

Format: date-time.

network.quality object · optional
network.quality.verified boolean · optional
network.rotatable boolean · optional
network.rotation object · optional
network.rotation.rotated boolean · optional
source_bound boolean · optional
status string · optional
terminal_reason string | null · optional
updated_at string · optional

Format: date-time.

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /network/tunnels/{tunnel_id}/rotate

Select a verified different exit with the same constraints

Parameters

tunnel_id path · required

Response · 200

Success

created_at string · optional

Format: date-time.

endpoints object · optional
endpoints.http object · optional
endpoints.http.host string · required
endpoints.http.port integer · required
endpoints.http.scheme string · required
endpoints.socks5 object · optional
endpoints.socks5.host string · required
endpoints.socks5.port integer · required
endpoints.socks5.scheme string · required
id string · optional
lease_expires_at string · optional

Format: date-time.

limits object · optional
limits.idle_timeout_s integer · optional
limits.max_bytes integer · optional
limits.max_connections integer · optional
limits.max_ttl_s integer · optional
network object · optional
network.connection_type string · optional
network.exit object · optional
network.exit.asn string · optional
network.exit.country string · optional
network.exit.hosting boolean · optional
network.exit.ip string · optional
network.exit.isp string · optional
network.exit.mobile boolean · optional
network.exit.region string · optional
network.exit.verified_at string · optional

Format: date-time.

network.quality object · optional
network.quality.verified boolean · optional
network.rotatable boolean · optional
network.rotation object · optional
network.rotation.rotated boolean · optional
source_bound boolean · optional
status string · optional
terminal_reason string | null · optional
updated_at string · optional

Format: date-time.

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /readyz

Read service availability

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /resources/captcha

Create a captcha resource

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

Provide captcha_type and the matching site_key/page_url or image_b64, or the equivalent descriptor.

action string · optional
browser_context object · optional

Optional context from the browser submitting the token. Cookie values are transient and are not stored in Hive resource records. Only include cookies needed for this CAPTCHA. Context support depends on solver availability; unsupported combinations fail explicitly.

browser_context.cookies object[] · optional

Maximum items: 100.

browser_context.cookies[].domain string · required
browser_context.cookies[].expires number · optional

Unix timestamp; omit or use -1 for session cookies.

browser_context.cookies[].http_only boolean · optional

Default: false.

browser_context.cookies[].name string · required
browser_context.cookies[].path string · required
browser_context.cookies[].secure boolean · optional

Default: false.

browser_context.cookies[].value string · required
browser_context.user_agent string · optional

Minimum length: 1. Maximum length: 1024.

captcha_type recaptcha_v2 | recaptcha_v2_enterprise | recaptcha_v3 | hcaptcha | turnstile | image · optional
descriptor object · optional
descriptor.action string · optional
descriptor.browser_context object · optional

Optional context from the browser submitting the token. Cookie values are transient and are not stored in Hive resource records. Only include cookies needed for this CAPTCHA. Context support depends on solver availability; unsupported combinations fail explicitly.

descriptor.browser_context.cookies object[] · optional

Maximum items: 100.

descriptor.browser_context.cookies[].domain string · required
descriptor.browser_context.cookies[].expires number · optional

Unix timestamp; omit or use -1 for session cookies.

descriptor.browser_context.cookies[].http_only boolean · optional

Default: false.

descriptor.browser_context.cookies[].name string · required
descriptor.browser_context.cookies[].path string · required
descriptor.browser_context.cookies[].secure boolean · optional

Default: false.

descriptor.browser_context.cookies[].value string · required
descriptor.browser_context.user_agent string · optional

Minimum length: 1. Maximum length: 1024.

descriptor.enterprise boolean · optional
descriptor.image_b64 string · optional
descriptor.min_score number · optional

Minimum: 0. Maximum: 1.

descriptor.page_url string · optional

Format: uri.

descriptor.site_key string · optional
descriptor.type recaptcha_v2 | recaptcha_v2_enterprise | recaptcha_v3 | hcaptcha | turnstile | image · optional
enterprise boolean · optional
image_b64 string · optional
min_score number · optional

Minimum: 0. Maximum: 1.

page_url string · optional

Format: uri.

site_key string · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /resources/mail

Create a mail resource

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

country string · optional

Pattern: ^[A-Za-z]{2}$.

domain string · optional
service string · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /resources/network

Create a network resource

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

asn string | integer | null · optional

Default: null.

country string · required

Pattern: ^[A-Za-z]{2}$.

identity_token string | null · optional

Default: null. Minimum length: 1. Maximum length: 128.

mobile boolean · optional

Default: true.

region string | null · optional

Default: null. Minimum length: 1. Maximum length: 100.

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /resources/phone

Create a phone resource

Parameters

Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

JSON body

country string · optional

Pattern: ^[A-Za-z]{2}$.

service string · optional

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

DELETE /resources/{resource_id}

Release a resource

Parameters

resource_id path · required

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

GET /resources/{resource_id}

Read a resource

Parameters

resource_id path · required

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /resources/{resource_id}/poll

Poll resource progress

Parameters

resource_id path · required
Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /resources/{resource_id}/renew

Extend the resource lease

Parameters

resource_id path · required
Idempotency-Key header · optional

Reuse the same key and request to replay a successful allocation.

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /scrapers/assets

Upload Research Image

Upload private image bytes (10 MiB maximum); retained encrypted for 24 hours.

Response · 201

Successful Response

asset_id string · required
expires_at integer · required
media_type image/jpeg | image/png | image/webp · required
sha256 string · required
size_bytes integer · required
url string · required

Bearer-authenticated, relative Hive download URL.

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

DELETE /scrapers/assets/{asset_id}

Delete Research Image

Delete an image. Already-submitted jobs retain their private input copy until completion.

Parameters

asset_id path · required

Response · 204

Successful Response

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /scrapers/assets/{asset_id}

Download Research Image

Download this bearer's uploaded image, Lens thumbnail or retained post image before its expiry.

Parameters

asset_id path · required

Response · 200

Successful Response

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /scrapers/capacity

Research Capacity

Read service availability and retry guidance without starting work.

Response · 200

Successful Response

available boolean · required

Whether shared research admission is currently available. Individual requests can still be unavailable.

reason capacity_limited | service_unavailable | null · required
retry_after_s integer | null · required

Minimum delay before rechecking, when known. Elapsed time does not prove recovery.

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/group/posts

Group Posts

One chronological collection from an explicit public group URL; observed group permalinks establish membership.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Finite upstream collection bound. limit only pages this immutable snapshot. Minimum: 1. Maximum: 200. Default: 50.

group_url string · required

Minimum length: 1. Maximum length: 512.

keyword string | null · optional

Optional literal case-insensitive substring checked on returned text; not Boolean search syntax. Maximum length: 200.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · optional

Inclusive lower UTC date checked against returned timestamps. Pattern: ^\d{4}-\d{2}-\d{2}$.

until_date string | null · optional

Exclusive upper UTC date checked against returned timestamps. Pattern: ^\d{4}-\d{2}-\d{2}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

group_url string · required
items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].group_title string | null · required
items[].group_url string · required
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].relation_evidence group_permalink · required
items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

keyword string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

local_filters keyword | since_date | until_date[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

post_url null · required
query null · required
query_handling null · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

until_date string | null · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/groups/search

Group Search

Search group names and public profiles; private group content is never retrieved.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 200. Default: 50.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

One group-name query. Public web-index candidates are not an exhaustive Facebook search. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

content_retrieved false · required
count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required
items object[] · required

Maximum items: 50.

items[].admin_moderator_count integer | null · required

Minimum: 0.

items[].average_posts_per_day number | null · required

Minimum: 0.

items[].content_retrieved false · required
items[].created_at string | null · required
items[].description string | null · required
items[].group_id string · required
items[].group_url string · required
items[].image_url string | null · required
items[].location string | null · required
items[].members_count integer | null · required

Minimum: 0.

items[].name string · required
items[].posts_last_month integer | null · required

Minimum: 0.

items[].posts_today integer | null · required

Minimum: 0.

items[].privacy public | private | null · required
items[].query string · required
items[].relation_evidence public_group_profile · required
items[].vanity string | null · required
items[].visibility visible | hidden | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons unavailable_or_invalid_group_rows[] · required
query string · required
query_handling upstream_hint · required
requested_collection_size integer · required

Minimum: 1. Maximum: 200.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

search_source public_web_index · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/pages/search

Page Search

One finite public page search. Page candidates are not groups or verified identities.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Finite upstream collection bound. limit only pages this immutable snapshot. Minimum: 1. Maximum: 200. Default: 50.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required
items object[] · required

Maximum items: 50.

items[].biography string | null · required
items[].display_name string | null · optional
items[].image_url string | null · optional
items[].is_verified boolean | null · required
items[].location string | null · required
items[].profile_id string · required
items[].profile_type page · required
items[].profile_url string · required
items[].query string · required
items[].username string | null · optional
items[].website_url string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

query string · required
query_handling upstream_hint · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/people/search

Facebook People Search

Collect at most 50 native people candidates, including numeric profiles without usernames. Opaque people links retain their returned URL without an inferred numeric ID or username; details/posts still require supported handle or numeric targets. Location texts are appended in caller order as query hints, never enforced city filters. Prefer: respond-async returns promptly; default wait is at most 20 seconds. Cached pages/replays launch no new work.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

locations string[] · optional

Up to five location hints of 1–100 characters, appended in caller order to one native query. Ambiguous names remain text; no city resolution or AND/OR geographic filters are enforced. Maximum items: 5.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

Free-text name, optionally with employer/city. Echoed exactly; matching does not verify identity or enforce filters. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

effective_query string · required

Text sent to native account search, including location hints. Instagram commas become spaces to keep one query.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required

City and employer restrictions are not enforced. Returned accounts are candidates, never verified identities.

items object[] · required

Maximum items: 50.

items[].biography string | null · optional
items[].category string | null · optional
items[].cover_image_url string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].location string | null · optional
items[].missing_fields display_name | biography | image_url[] · required

Unavailable core profile fields. No inferred or purchased enrichment. Maximum items: 3.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_type person | page | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].query string · required
items[].rank integer · required

Absolute one-based position after canonical account deduplication, preserving native order. Minimum: 1. Maximum: 50.

items[].status candidate · required
items[].username string | null · optional
items[].website_url string | null · optional
limit integer · required

Minimum: 1. Maximum: 50.

location_handling not_requested | query_hint · required
locations string[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

query string · required
result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Durable native people search; poll the owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/post/comments

Comments

One top-level public-comment collection. Nested replies and full thread completeness are unavailable.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 200. Default: 100.

keyword string | null · optional

Optional literal case-insensitive substring checked on returned text; not Boolean search syntax. Maximum length: 200.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

post_url string · required

Minimum length: 1. Maximum length: 512.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · optional

Inclusive lower UTC date checked against returned timestamps. Pattern: ^\d{4}-\d{2}-\d{2}$.

until_date string | null · optional

Exclusive upper UTC date checked against returned timestamps. Pattern: ^\d{4}-\d{2}-\d{2}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

group_url null · required
items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · required
items[].author.image_url string | null · required
items[].author.profile_id string | null · required
items[].author.profile_url string | null · required
items[].author.username string | null · required
items[].comment_id string · required
items[].comment_scope top_level | top_level_candidate · required
items[].likes_count integer | null · required

Minimum: 0.

items[].parent_comment_id null · required
items[].post_url string · required
items[].relation_evidence comment_permalink · required
items[].text string | null · required
items[].timestamp string | null · required
items[].url string · required
keyword string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

local_filters keyword | since_date | until_date[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

post_url string · required
query null · required
query_handling null · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

until_date string | null · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/posts/search

Post Search

One public post query; requested literal keyword/date filters are checked locally on returned fields.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Finite upstream collection bound. limit only pages this immutable snapshot. Minimum: 1. Maximum: 200. Default: 50.

keyword string | null · optional

Optional literal case-insensitive substring checked on returned text; not Boolean search syntax. Maximum length: 200.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

One upstream public keyword query; matching and recall are not guaranteed. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · optional

Inclusive lower UTC date checked against returned timestamps. Pattern: ^\d{4}-\d{2}-\d{2}$.

until_date string | null · optional

Exclusive upper UTC date checked against returned timestamps. Pattern: ^\d{4}-\d{2}-\d{2}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

group_url null · required
items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].query string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].relation_evidence search_query · required
items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

keyword string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

local_filters keyword | since_date | until_date[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

post_url null · required
query string · required
query_handling upstream_hint · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

until_date string | null · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/profile/details

Facebook Profile Details

Retrieve public Facebook profile or page details for up to ten targets.

JSON body

One to ten Facebook handles, numeric IDs, or canonical profile/page URLs.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 10.

items object[] · required

Maximum items: 10.

items[].biography string | null · optional
items[].category string | null · optional
items[].cover_image_url string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].location string | null · optional
items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_type person | page | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].target string · required
items[].username string | null · optional
items[].website_url string | null · optional
unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/facebook/profile/posts

Facebook Profile Posts

Return recent public person/page posts, including numeric profiles. At most 40 posts per target; pages cannot recover older history. Prefer: respond-async submits promptly.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

Collect recent posts, dated Page history, or cursor-based personal history.

archive boolean · optional

Default: false.

archive_cursor string | null · optional

Minimum length: 1. Maximum length: 262144.

archive_end_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_kind person | page | null · optional

Required in archive mode: Pages use date windows; personal profiles use a continuation cursor with expanded photo albums.

archive_window_days integer · optional

Minimum: 1. Maximum: 30. Default: 30.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

per_target_limit integer · optional

Maximum posts in one collection: 40 recent by default, up to 500 in archive mode. Personal archives require at least 5 so provider cursors can advance; they default to 100. Minimum: 1. Maximum: 500. Default: 40.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

retain_media boolean · optional

Preserve photos/video thumbnails. Recent jobs allow 20 unique images and 32 MiB over 30 seconds; archive jobs allow 100 images and 96 MiB over 120 seconds. Each image is capped at 10 MiB and 40 million pixels. Shared bearer/global quotas still apply. No video download; assets expire after 24 hours. Cached pages/replay never download again. Default: false.

Response · 200

Successful Response

archive_end_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_start_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

history_complete null · required

Full platform history is unknown. Pages cover only the cached collection.

history_scope collected_window | recent_public_timeline | archive_date_window | archive_cursor_window · optional

Date windows continue through next_archive_end_date; cursor windows continue through next_archive_cursor. Page numbers only traverse the cached result. Default: "collected_window".

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].context object[] · required

Maximum items: 2.

items[].context[].author object · required
items[].context[].author.display_name string | null · optional
items[].context[].author.image_url string | null · optional
items[].context[].author.profile_id string | null · optional
items[].context[].author.profile_url string · required
items[].context[].author.username string | null · optional
items[].context[].comments_count integer | null · optional

Minimum: 0.

items[].context[].geotag object | null · required
items[].context[].kind repost | quote · required
items[].context[].likes_count integer | null · optional

Minimum: 0.

items[].context[].media object[] · required

Maximum items: 500.

items[].context[].media[].asset object | null · optional
items[].context[].media[].bitrate integer | null · optional

Minimum: 0.

items[].context[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].context[].media[].mime_type string | null · optional
items[].context[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].context[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].context[].media[].role original | thumbnail · required
items[].context[].media[].type image | video · required
items[].context[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].context[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].context[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].context[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].context[].post_id string · required
items[].context[].quotes_count integer | null · optional

Minimum: 0.

items[].context[].reposts_count integer | null · optional

Minimum: 0.

items[].context[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].context[].text string | null · required
items[].context[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].context[].url string · required
items[].context[].views_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_pinned boolean | null · optional
items[].is_quote boolean | null · optional
items[].is_repost boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].target string · required
items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_archive_cursor string | null · optional

Submit as archive_cursor in a new Instagram or Facebook personal archive job with the same profile and window settings after consuming cached pages. Null when the provider reached the visible end or could not safely continue. Maximum length: 262144.

next_archive_end_date string | null · optional

Submit a new archive job ending on this date after consuming all cached pages. Null when the date range is exhausted or the current window hit its collection cap; narrow a capped window before continuing. Pattern: ^\d{4}-\d{2}-\d{2}$.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

per_target_limit integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

targets object[] · required

Maximum items: 10.

targets[].collected_count integer · required

Minimum: 0. Maximum: 1000.

targets[].input_index integer · required

Minimum: 0. Maximum: 9.

targets[].reason not_found | private | unavailable | malformed_data | ambiguous | null · required
targets[].status resolved | unresolved · required
targets[].target string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Response · 202

Durable job; poll the returned owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/google/search

Google Search

Search Google for bounded organic results, preserving each query and its reported rank.

JSON body

One to five queries of up to 32 words; search operators are supported.

country string | null · optional

Supported two-letter search country code, such as us or gb. Minimum length: 2. Maximum length: 2.

language string | null · optional

Search interface language, such as en, zh-CN or pt-BR; does not restrict result language. Minimum length: 2. Maximum length: 5.

limit integer · optional

Maximum organic results per distinct query. Minimum: 1. Maximum: 100. Default: 20.

pages integer · optional

Maximum pages per query. Minimum: 1. Maximum: 3. Default: 1.

queries string[] · required

Minimum items: 1. Maximum items: 5.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 500.

items object[] · required

Maximum items: 500.

items[].query string · required
items[].rank integer · required

Minimum: 1.

items[].snippet string | null · optional
items[].title string · required
items[].url string · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/google_lens/jobs

Submit Google Lens Job

Submit once and safely replay the same key and body for seven days. Poll the returned URL; interrupted work is not automatically repeated.

Parameters

Idempotency-Key header · required

JSON body

Search one public HTTPS image or an owner-scoped uploaded image.

asset_id string | null · optional

Private image from POST /scrapers/assets; mutually exclusive with image_url. Pattern: ^rasset_[a-f0-9]{32}$.

image_url string | null · optional

Public HTTPS image URL, or the exact relative URL returned by POST /scrapers/assets. Minimum length: 1. Maximum length: 4096.

limit integer · optional

Maximum total records across all requested types, interleaved in the requested order. Minimum: 1. Maximum: 100. Default: 20.

search_types visual | exact | text[] · optional

Requested Lens modes; defaults to visual matches. Minimum items: 1. Maximum items: 3. Items must be distinct.

Response · 202

Successful Response

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
poll_url string · required
result object | null · optional
retry_after_s integer | null · optional
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /scrapers/google_lens/jobs/{job_id}

Get Google Lens Job

Poll without starting a new search. Results expire after 24 hours.

Parameters

job_id path · required

Response · 200

Successful Response

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
poll_url string · required
result object | null · optional
retry_after_s integer | null · optional
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 410 · Research result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/google_lens/search

Google Lens Search

Find visual or exact image matches and extract text using Google Lens. Results share one total limit; a search can return no matches.

JSON body

Search one public HTTPS image or an owner-scoped uploaded image.

asset_id string | null · optional

Private image from POST /scrapers/assets; mutually exclusive with image_url. Pattern: ^rasset_[a-f0-9]{32}$.

image_url string | null · optional

Public HTTPS image URL, or the exact relative URL returned by POST /scrapers/assets. Minimum length: 1. Maximum length: 4096.

limit integer · optional

Maximum total records across all requested types, interleaved in the requested order. Minimum: 1. Maximum: 100. Default: 20.

search_types visual | exact | text[] · optional

Requested Lens modes; defaults to visual matches. Minimum items: 1. Maximum items: 3. Items must be distinct.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 100.

items object | object[] · required

Maximum items: 100.

items[].image_expires_at integer | null · optional
items[].image_height integer | null · optional

Minimum: 1. Maximum: 100000.

items[].image_kind thumbnail | null · optional
items[].image_sha256 string | null · optional

Pattern: ^[a-f0-9]{64}$.

items[].image_url string | null · optional

Maximum length: 4096.

items[].image_width integer | null · optional

Minimum: 1. Maximum: 100000.

items[].source string | null · optional

Maximum length: 512.

items[].source_icon_url string | null · optional

Maximum length: 4096.

items[].title string · required

Minimum length: 1. Maximum length: 4096.

items[].type visual | exact · required
items[].url string · required

Maximum length: 4096.

items[].text string · required

Minimum length: 1. Maximum length: 20000.

items[].type text · required
partial boolean | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/people/search

Instagram People Search

Collect at most 50 native account candidates. Commas become spaces; locations are query hints. Missing biography/avatar fields are explicit. Prefer: respond-async returns promptly; default wait is at most 20 seconds. Cached pages/replays launch no new work.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

locations string[] · optional

Up to five location hints of 1–100 characters, appended in caller order to one native query. Ambiguous names remain text; no city resolution or AND/OR geographic filters are enforced. Maximum items: 5.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

Free-text name, optionally with employer/city. Echoed exactly; matching does not verify identity or enforce filters. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

effective_query string · required

Text sent to native account search, including location hints. Instagram commas become spaces to keep one query.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required

City and employer restrictions are not enforced. Returned accounts are candidates, never verified identities.

items object[] · required

Maximum items: 50.

items[].biography string | null · optional
items[].category string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].is_business boolean | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional
items[].missing_fields display_name | biography | image_url[] · required

Unavailable core profile fields. No inferred or purchased enrichment. Maximum items: 3.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].query string · required
items[].rank integer · required

Absolute one-based position after canonical account deduplication, preserving native order. Minimum: 1. Maximum: 50.

items[].status candidate · required
items[].username string · required
items[].website_url string | null · optional
limit integer · required

Minimum: 1. Maximum: 50.

location_handling not_requested | query_hint · required
locations string[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

query string · required
result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Durable native account search; poll the owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/details

Instagram Profile Details

Retrieve public Instagram profile details for up to ten handles or URLs.

JSON body

One to ten Instagram handles or canonical public profile URLs.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 10.

items object[] · required

Maximum items: 10.

items[].biography string | null · optional
items[].category string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_business boolean | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional
items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].target string · required
items[].username string · required
items[].website_url string | null · optional
unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/followers

Instagram Followers

One attributed public follower collection, nullable metadata and unknown full graph; cached pages only.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

One finite first collection. Raw continuation tokens are never accepted or returned; cached pages cannot continue the upstream graph. Minimum: 25. Maximum: 1000. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 256.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].biography null · required
items[].created_at null · required
items[].display_name string | null · optional
items[].image_url string | null · optional
items[].is_private boolean | null · required
items[].is_verified boolean | null · required
items[].location null · required
items[].profile_id string · required
items[].profile_url string · required
items[].relation followers | following · required
items[].relation_evidence source_attribution · required
items[].source_profile string · required
items[].username string · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

relation followers | following · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

upstream_continuation_available false · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/following

Instagram Following

One attributed public following collection; native continuation is unavailable through this public API.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

One finite first collection. Raw continuation tokens are never accepted or returned; cached pages cannot continue the upstream graph. Minimum: 25. Maximum: 1000. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 256.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].biography null · required
items[].created_at null · required
items[].display_name string | null · optional
items[].image_url string | null · optional
items[].is_private boolean | null · required
items[].is_verified boolean | null · required
items[].location null · required
items[].profile_id string · required
items[].profile_url string · required
items[].relation followers | following · required
items[].relation_evidence source_attribution · required
items[].source_profile string · required
items[].username string · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

relation followers | following · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

upstream_continuation_available false · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/highlights

Instagram Highlights

Saved public Highlight metadata only; finite summaries and fixed profile outcome, with no Story content or complete-history claim.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Maximum saved Highlight summaries; use the content route for actual Story items. Minimum: 1. Maximum: 100. Default: 10.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 256.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

content_available false · required
count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].content_retrieved false · required
items[].cover_image_url string | null · required
items[].created_at string | null · required
items[].highlight_id string · required
items[].latest_story_at string | null · required
items[].relation_evidence profile_container · required
items[].reported_media_count integer | null · required

Minimum: 0.

items[].source_profile string · required
items[].title string | null · required
items[].updated_at string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

profile_reason private | not_found | rate_limited | fetch_failed | partial_response | null · required
profile_status success | partial | private | not_found | unavailable | rate_limited · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/highlights/content

Content

Actual saved Story items with IDs, dates and media references; media bytes are not retained.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Finite number of actual Story items; limit only pages the saved collection. Minimum: 1. Maximum: 200. Default: 50.

highlight_limit integer · optional

Maximum saved Highlight reels expanded for this profile. Minimum: 1. Maximum: 25. Default: 5.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 512.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

content_scope saved_highlight_stories · required
count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].caption string | null · required
items[].expires_at null · required
items[].hashtags string[] · required

Maximum items: 100.

items[].height integer | null · required

Minimum: 0.

items[].highlight_id string · required

Pattern: ^[1-9][0-9]{0,39}$.

items[].highlight_title string | null · required
items[].highlight_url string · required
items[].links string[] · required

Maximum items: 100.

items[].media_id string · required

Observed numeric media component of the Story ID. Pattern: ^[1-9][0-9]{0,39}$.

items[].media_retained false · required
items[].media_type image | video · required
items[].media_url string · required
items[].mentions string[] · required

Maximum items: 100.

items[].music_artist string | null · required
items[].music_title string | null · required
items[].owner_id string | null · required
items[].position integer · required

Minimum: 1.

items[].relation_evidence observed_owner_and_highlight · required
items[].source_profile string · required
items[].story_id string · required

Observed Story ID, including its validated owner suffix when returned. Pattern: ^[1-9][0-9]{0,39}(?:_[1-9][0-9]{0,39})?$.

items[].thumbnail_url string | null · required
items[].timestamp string · required
items[].width integer | null · required

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

media_retained false · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

requested_collection_size integer · required

Minimum: 1. Maximum: 200.

requested_highlight_limit integer · required

Minimum: 1. Maximum: 25.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/posts

Instagram Profile Posts

Return a cached page or wait up to 20 seconds. Prefer: respond-async submits promptly. Media URLs may expire; full history is unknown.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

Collect recent posts or continue a dated public-profile archive.

archive boolean · optional

Walk a dated public profile window. Continue the same window using next_archive_cursor, then use next_archive_end_date. Default: false.

archive_cursor string | null · optional

Opaque continuation returned as next_archive_cursor by the previous completed archive job for the same profile and date window. Minimum length: 1. Maximum length: 262144.

archive_end_date string | null · optional

Inclusive UTC date at the end of the archive window; required with archive=true. Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_window_days integer · optional

Maximum UTC days in this archive window, up to 365. Post and runtime bounds still apply; follow the returned cursor before moving to an older window. Minimum: 1. Maximum: 365. Default: 30.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

per_target_limit integer · optional

Maximum posts per profile: 50 by default, up to 500 across the collection. Larger windows collect carousel media for more posts. Archive mode defaults to 500 if omitted. Minimum: 1. Maximum: 500. Default: 50.

post_urls string[] | null · optional

Look up up to ten specific public posts for this single profile, including full carousel media. Each post must belong to or explicitly list the profile as a coauthor. Cannot be combined with archive mode. Minimum items: 1. Maximum items: 10.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

retain_media boolean · optional

Preserve photos/video thumbnails. Recent jobs allow 20 unique images and 32 MiB over 30 seconds; archive jobs allow 100 images and 96 MiB over 120 seconds. Each image is capped at 10 MiB and 40 million pixels. Shared bearer/global quotas still apply. No video download; assets expire after 24 hours. Cached pages/replay never download again. Default: false.

Response · 200

Successful Response

archive_end_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_start_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

history_complete null · required

Full platform history is unknown. Pages cover only the cached collection.

history_scope collected_window | recent_public_timeline | archive_date_window | archive_cursor_window · optional

Date windows continue through next_archive_end_date; cursor windows continue through next_archive_cursor. Page numbers only traverse the cached result. Default: "collected_window".

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].context object[] · required

Maximum items: 2.

items[].context[].author object · required
items[].context[].author.display_name string | null · optional
items[].context[].author.image_url string | null · optional
items[].context[].author.profile_id string | null · optional
items[].context[].author.profile_url string · required
items[].context[].author.username string | null · optional
items[].context[].comments_count integer | null · optional

Minimum: 0.

items[].context[].geotag object | null · required
items[].context[].kind repost | quote · required
items[].context[].likes_count integer | null · optional

Minimum: 0.

items[].context[].media object[] · required

Maximum items: 500.

items[].context[].media[].asset object | null · optional
items[].context[].media[].bitrate integer | null · optional

Minimum: 0.

items[].context[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].context[].media[].mime_type string | null · optional
items[].context[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].context[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].context[].media[].role original | thumbnail · required
items[].context[].media[].type image | video · required
items[].context[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].context[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].context[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].context[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].context[].post_id string · required
items[].context[].quotes_count integer | null · optional

Minimum: 0.

items[].context[].reposts_count integer | null · optional

Minimum: 0.

items[].context[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].context[].text string | null · required
items[].context[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].context[].url string · required
items[].context[].views_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_pinned boolean | null · optional
items[].is_quote boolean | null · optional
items[].is_repost boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].target string · required
items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_archive_cursor string | null · optional

Submit as archive_cursor in a new Instagram or Facebook personal archive job with the same profile and window settings after consuming cached pages. Null when the provider reached the visible end or could not safely continue. Maximum length: 262144.

next_archive_end_date string | null · optional

Submit a new archive job ending on this date after consuming all cached pages. Null when the date range is exhausted or the current window hit its collection cap; narrow a capped window before continuing. Pattern: ^\d{4}-\d{2}-\d{2}$.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

per_target_limit integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

targets object[] · required

Maximum items: 10.

targets[].collected_count integer · required

Minimum: 0. Maximum: 1000.

targets[].input_index integer · required

Minimum: 0. Maximum: 9.

targets[].reason not_found | private | unavailable | malformed_data | ambiguous | null · required
targets[].status resolved | unresolved · required
targets[].target string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Response · 202

Durable job; poll the returned owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/instagram/profile/tagged

Instagram Tagged

Collect genuinely tagged incoming posts with their actual source authors; mentions alone never establish a tag.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 200. Default: 50.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 256.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required

Unknown upstream completeness; a final cached page does not prove exhaustion.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].relation_evidence tagged_users · required

Actual returned tag fields establish this relation; a caption mention or echoed target does not.

items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_profile string · required
items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

tagged_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /scrapers/jobs/{job_id}

Get Social Research Job

Wait up to 25 seconds for completion without starting or replaying work.

Parameters

job_id path · required
wait_seconds query · optional

Response · 200

Successful Response

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required
created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/jobs/{job_id}/recover

Recover Social Research Job

Recover accepted terminal X archive data through GET-only reads. No new provider run; original result expiry remains.

Parameters

job_id path · required

Response · 200

Successful Response

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/company/employees

Company Employees

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

companies string[] · required

Minimum items: 1. Maximum items: 20.

company_headcounts string[] · optional

Maximum items: 10.

detail_level basic | full · optional

Default: "basic".

exclude_functions string[] · optional

Maximum items: 30.

exclude_industries string[] · optional

Maximum items: 50.

exclude_locations string[] · optional

Maximum items: 50.

exclude_past_titles string[] · optional

Maximum items: 50.

exclude_seniority_levels string[] · optional

Maximum items: 20.

exclude_titles string[] · optional

Maximum items: 50.

experience_levels string[] · optional

Maximum items: 10.

functions string[] · optional

Maximum items: 30.

industries string[] · optional

Maximum items: 50.

limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

locations string[] · optional

Maximum items: 50.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

past_titles string[] · optional

Maximum items: 50.

query string | null · optional

Minimum length: 1. Maximum length: 300.

recently_changed_jobs boolean | null · optional
seniority_levels string[] · optional

Maximum items: 20.

titles string[] · optional

Maximum items: 50.

years_at_company string[] · optional

Maximum items: 10.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].about string | null · optional
items[].certifications object[] · optional
items[].certifications[].credential_id string | null · optional
items[].certifications[].credential_url string | null · optional
items[].certifications[].expires_at string | null · optional
items[].certifications[].issued_at string | null · optional
items[].certifications[].issuer string | null · optional
items[].certifications[].name string | null · optional
items[].connection_count integer | null · optional
items[].courses object[] · optional
items[].courses[].associated_with string | null · optional
items[].courses[].name string | null · optional
items[].courses[].number string | null · optional
items[].current_positions object[] · optional
items[].current_positions[].company string | null · optional
items[].current_positions[].company_url string | null · optional
items[].current_positions[].date_range object | null · optional
items[].current_positions[].description string | null · optional
items[].current_positions[].duration string | null · optional
items[].current_positions[].employment_type string | null · optional
items[].current_positions[].location string | null · optional
items[].current_positions[].skills string[] · optional
items[].current_positions[].title string | null · optional
items[].current_positions[].workplace_type string | null · optional
items[].education object[] · optional
items[].education[].activities string | null · optional
items[].education[].date_range object | null · optional
items[].education[].degree string | null · optional
items[].education[].description string | null · optional
items[].education[].field_of_study string | null · optional
items[].education[].school string | null · optional
items[].education[].school_url string | null · optional
items[].experience object[] · optional
items[].experience[].company string | null · optional
items[].experience[].company_url string | null · optional
items[].experience[].date_range object | null · optional
items[].experience[].description string | null · optional
items[].experience[].duration string | null · optional
items[].experience[].employment_type string | null · optional
items[].experience[].location string | null · optional
items[].experience[].skills string[] · optional
items[].experience[].title string | null · optional
items[].experience[].workplace_type string | null · optional
items[].first_name string | null · optional
items[].follower_count integer | null · optional
items[].headline string | null · optional
items[].honors object[] · optional
items[].honors[].description string | null · optional
items[].honors[].issued_at string | null · optional
items[].honors[].issuer string | null · optional
items[].honors[].title string | null · optional
items[].image object | null · optional
items[].languages object[] · optional
items[].languages[].name string · required
items[].languages[].proficiency string | null · optional
items[].last_name string | null · optional
items[].location string | null · optional
items[].name string | null · optional
items[].profile_url string | null · optional
items[].projects object[] · optional
items[].projects[].date_range object | null · optional
items[].projects[].description string | null · optional
items[].projects[].name string | null · optional
items[].projects[].url string | null · optional
items[].public_id string | null · optional
items[].publications object[] · optional
items[].publications[].description string | null · optional
items[].publications[].published_at string | null · optional
items[].publications[].publisher string | null · optional
items[].publications[].title string | null · optional
items[].publications[].url string | null · optional
items[].recommendations object[] · optional
items[].recommendations[].author object | null · optional
items[].recommendations[].relationship string | null · optional
items[].recommendations[].text string | null · optional
items[].skills object[] · optional
items[].skills[].endorsements integer | null · optional
items[].skills[].name string · required
items[].volunteering object[] · optional
items[].volunteering[].cause string | null · optional
items[].volunteering[].date_range object | null · optional
items[].volunteering[].description string | null · optional
items[].volunteering[].organization string | null · optional
items[].volunteering[].role string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/people/search

People Search

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

company_headcounts string[] · optional

Maximum items: 10.

company_locations string[] · optional

Maximum items: 70.

current_companies string[] · optional

Maximum items: 50.

current_titles string[] · optional

Maximum items: 50.

detail_level basic | full · optional

Default: "basic".

exclude_company_locations string[] · optional

Maximum items: 70.

exclude_current_companies string[] · optional

Maximum items: 50.

exclude_current_titles string[] · optional

Maximum items: 50.

exclude_functions string[] · optional

Maximum items: 30.

exclude_industries string[] · optional

Maximum items: 50.

exclude_locations string[] · optional

Maximum items: 70.

exclude_past_companies string[] · optional

Maximum items: 50.

exclude_past_titles string[] · optional

Maximum items: 50.

exclude_schools string[] · optional

Maximum items: 50.

exclude_seniority_levels string[] · optional

Maximum items: 20.

experience_levels string[] · optional

Maximum items: 10.

first_names string[] · optional

Maximum items: 50.

functions string[] · optional

Maximum items: 30.

industries string[] · optional

Maximum items: 50.

languages string[] · optional

Maximum items: 20.

last_names string[] · optional

Maximum items: 50.

limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

locations string[] · optional

Maximum items: 70.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

past_companies string[] · optional

Maximum items: 50.

past_titles string[] · optional

Maximum items: 50.

query string | null · optional

Minimum length: 1. Maximum length: 300.

recently_changed_jobs boolean | null · optional
recently_posted boolean | null · optional
schools string[] · optional

Maximum items: 50.

seniority_levels string[] · optional

Maximum items: 20.

years_at_company string[] · optional

Maximum items: 10.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].about string | null · optional
items[].certifications object[] · optional
items[].certifications[].credential_id string | null · optional
items[].certifications[].credential_url string | null · optional
items[].certifications[].expires_at string | null · optional
items[].certifications[].issued_at string | null · optional
items[].certifications[].issuer string | null · optional
items[].certifications[].name string | null · optional
items[].connection_count integer | null · optional
items[].courses object[] · optional
items[].courses[].associated_with string | null · optional
items[].courses[].name string | null · optional
items[].courses[].number string | null · optional
items[].current_positions object[] · optional
items[].current_positions[].company string | null · optional
items[].current_positions[].company_url string | null · optional
items[].current_positions[].date_range object | null · optional
items[].current_positions[].description string | null · optional
items[].current_positions[].duration string | null · optional
items[].current_positions[].employment_type string | null · optional
items[].current_positions[].location string | null · optional
items[].current_positions[].skills string[] · optional
items[].current_positions[].title string | null · optional
items[].current_positions[].workplace_type string | null · optional
items[].education object[] · optional
items[].education[].activities string | null · optional
items[].education[].date_range object | null · optional
items[].education[].degree string | null · optional
items[].education[].description string | null · optional
items[].education[].field_of_study string | null · optional
items[].education[].school string | null · optional
items[].education[].school_url string | null · optional
items[].experience object[] · optional
items[].experience[].company string | null · optional
items[].experience[].company_url string | null · optional
items[].experience[].date_range object | null · optional
items[].experience[].description string | null · optional
items[].experience[].duration string | null · optional
items[].experience[].employment_type string | null · optional
items[].experience[].location string | null · optional
items[].experience[].skills string[] · optional
items[].experience[].title string | null · optional
items[].experience[].workplace_type string | null · optional
items[].first_name string | null · optional
items[].follower_count integer | null · optional
items[].headline string | null · optional
items[].honors object[] · optional
items[].honors[].description string | null · optional
items[].honors[].issued_at string | null · optional
items[].honors[].issuer string | null · optional
items[].honors[].title string | null · optional
items[].image object | null · optional
items[].languages object[] · optional
items[].languages[].name string · required
items[].languages[].proficiency string | null · optional
items[].last_name string | null · optional
items[].location string | null · optional
items[].name string | null · optional
items[].profile_url string | null · optional
items[].projects object[] · optional
items[].projects[].date_range object | null · optional
items[].projects[].description string | null · optional
items[].projects[].name string | null · optional
items[].projects[].url string | null · optional
items[].public_id string | null · optional
items[].publications object[] · optional
items[].publications[].description string | null · optional
items[].publications[].published_at string | null · optional
items[].publications[].publisher string | null · optional
items[].publications[].title string | null · optional
items[].publications[].url string | null · optional
items[].recommendations object[] · optional
items[].recommendations[].author object | null · optional
items[].recommendations[].relationship string | null · optional
items[].recommendations[].text string | null · optional
items[].skills object[] · optional
items[].skills[].endorsements integer | null · optional
items[].skills[].name string · required
items[].volunteering object[] · optional
items[].volunteering[].cause string | null · optional
items[].volunteering[].date_range object | null · optional
items[].volunteering[].description string | null · optional
items[].volunteering[].organization string | null · optional
items[].volunteering[].role string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/post/comments

Post Comments

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

freshness hour | day | week | month | three_months | six_months | year | null · optional
include_replies boolean · optional

Default: true.

limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

posts string[] · required

Minimum items: 1. Maximum items: 20.

replies_limit integer · optional

Minimum: 0. Maximum: 100. Default: 5.

since string | null · optional

Format: date-time.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].author object | null · optional
items[].comment_id string | null · optional
items[].comment_url string | null · optional
items[].created_at string | null · optional
items[].metrics object | null · optional
items[].parent_comment_id string | null · optional
items[].post object | null · optional
items[].replies object[] · optional
items[].replies[].author object | null · optional
items[].replies[].comment_id string | null · optional
items[].replies[].comment_url string | null · optional
items[].replies[].created_at string | null · optional
items[].replies[].metrics object | null · optional
items[].replies[].parent_comment_id string | null · optional
items[].replies[].text string | null · optional
items[].text string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/post/search

Post Search

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

author_companies string[] · optional

Maximum items: 10.

author_industries string[] · optional

Maximum items: 20.

author_keywords string | null · optional

Minimum length: 1. Maximum length: 300.

author_profiles string[] · optional

Maximum items: 10.

authors_employers string[] · optional

Maximum items: 20.

comments_freshness hour | day | week | month | three_months | six_months | year | null · optional
comments_limit integer · optional

Minimum: 0. Maximum: 100. Default: 0.

content_type all | videos | images | jobs | live_videos | documents | collaborative_articles · optional

Default: "all".

freshness hour | day | week | month | three_months | six_months | year | null · optional
limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

mentioned_companies string[] · optional

Maximum items: 10.

mentioned_profiles string[] · optional

Maximum items: 10.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

queries string[] · required

Minimum items: 1. Maximum items: 20.

reactions_limit integer · optional

Minimum: 0. Maximum: 100. Default: 0.

since string | null · optional

Format: date-time.

sort relevance | date · optional

Default: "relevance".

Response · 200

Successful Response

count integer · required
items object[] · required
items[].author object | null · optional
items[].comments object[] · optional
items[].comments[].author object | null · optional
items[].comments[].comment_id string | null · optional
items[].comments[].comment_url string | null · optional
items[].comments[].created_at string | null · optional
items[].comments[].metrics object | null · optional
items[].comments[].parent_comment_id string | null · optional
items[].comments[].post object | null · optional
items[].comments[].replies object[] · optional
items[].comments[].replies[].author object | null · optional
items[].comments[].replies[].comment_id string | null · optional
items[].comments[].replies[].comment_url string | null · optional
items[].comments[].replies[].created_at string | null · optional
items[].comments[].replies[].metrics object | null · optional
items[].comments[].replies[].parent_comment_id string | null · optional
items[].comments[].replies[].text string | null · optional
items[].comments[].text string | null · optional
items[].content_type string | null · optional
items[].media object[] · optional
items[].media[].description string | null · optional
items[].media[].image object | null · optional
items[].media[].kind string | null · optional
items[].media[].title string | null · optional
items[].media[].url string | null · optional
items[].metrics object | null · optional
items[].post_id string | null · optional
items[].post_url string | null · optional
items[].posted_at string | null · optional
items[].reactions object[] · optional
items[].reactions[].author object | null · optional
items[].reactions[].created_at string | null · optional
items[].reactions[].kind string | null · optional
items[].reactions[].post object | null · optional
items[].reactions[].reaction_id string | null · optional
items[].text string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/profile/comments

Profile Comments

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

freshness hour | day | week | month | three_months | six_months | year | null · optional
limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

profiles string[] · required

Minimum items: 1. Maximum items: 20.

since string | null · optional

Format: date-time.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].author object | null · optional
items[].comment_id string | null · optional
items[].comment_url string | null · optional
items[].created_at string | null · optional
items[].metrics object | null · optional
items[].parent_comment_id string | null · optional
items[].post object | null · optional
items[].replies object[] · optional
items[].replies[].author object | null · optional
items[].replies[].comment_id string | null · optional
items[].replies[].comment_url string | null · optional
items[].replies[].created_at string | null · optional
items[].replies[].metrics object | null · optional
items[].replies[].parent_comment_id string | null · optional
items[].replies[].text string | null · optional
items[].text string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/profile/details

Profile Details

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

detail_level basic | full · optional

Default: "full".

profiles string[] · required

Minimum items: 1. Maximum items: 50.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].about string | null · optional
items[].certifications object[] · optional
items[].certifications[].credential_id string | null · optional
items[].certifications[].credential_url string | null · optional
items[].certifications[].expires_at string | null · optional
items[].certifications[].issued_at string | null · optional
items[].certifications[].issuer string | null · optional
items[].certifications[].name string | null · optional
items[].connection_count integer | null · optional
items[].courses object[] · optional
items[].courses[].associated_with string | null · optional
items[].courses[].name string | null · optional
items[].courses[].number string | null · optional
items[].current_positions object[] · optional
items[].current_positions[].company string | null · optional
items[].current_positions[].company_url string | null · optional
items[].current_positions[].date_range object | null · optional
items[].current_positions[].description string | null · optional
items[].current_positions[].duration string | null · optional
items[].current_positions[].employment_type string | null · optional
items[].current_positions[].location string | null · optional
items[].current_positions[].skills string[] · optional
items[].current_positions[].title string | null · optional
items[].current_positions[].workplace_type string | null · optional
items[].education object[] · optional
items[].education[].activities string | null · optional
items[].education[].date_range object | null · optional
items[].education[].degree string | null · optional
items[].education[].description string | null · optional
items[].education[].field_of_study string | null · optional
items[].education[].school string | null · optional
items[].education[].school_url string | null · optional
items[].experience object[] · optional
items[].experience[].company string | null · optional
items[].experience[].company_url string | null · optional
items[].experience[].date_range object | null · optional
items[].experience[].description string | null · optional
items[].experience[].duration string | null · optional
items[].experience[].employment_type string | null · optional
items[].experience[].location string | null · optional
items[].experience[].skills string[] · optional
items[].experience[].title string | null · optional
items[].experience[].workplace_type string | null · optional
items[].first_name string | null · optional
items[].follower_count integer | null · optional
items[].headline string | null · optional
items[].honors object[] · optional
items[].honors[].description string | null · optional
items[].honors[].issued_at string | null · optional
items[].honors[].issuer string | null · optional
items[].honors[].title string | null · optional
items[].image object | null · optional
items[].languages object[] · optional
items[].languages[].name string · required
items[].languages[].proficiency string | null · optional
items[].last_name string | null · optional
items[].location string | null · optional
items[].name string | null · optional
items[].profile_url string | null · optional
items[].projects object[] · optional
items[].projects[].date_range object | null · optional
items[].projects[].description string | null · optional
items[].projects[].name string | null · optional
items[].projects[].url string | null · optional
items[].public_id string | null · optional
items[].publications object[] · optional
items[].publications[].description string | null · optional
items[].publications[].published_at string | null · optional
items[].publications[].publisher string | null · optional
items[].publications[].title string | null · optional
items[].publications[].url string | null · optional
items[].recommendations object[] · optional
items[].recommendations[].author object | null · optional
items[].recommendations[].relationship string | null · optional
items[].recommendations[].text string | null · optional
items[].skills object[] · optional
items[].skills[].endorsements integer | null · optional
items[].skills[].name string · required
items[].volunteering object[] · optional
items[].volunteering[].cause string | null · optional
items[].volunteering[].date_range object | null · optional
items[].volunteering[].description string | null · optional
items[].volunteering[].organization string | null · optional
items[].volunteering[].role string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/profile/posts

Profile Posts

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

comments_limit integer · optional

Minimum: 0. Maximum: 100. Default: 0.

freshness hour | day | week | month | three_months | six_months | year | null · optional
include_quotes boolean · optional

Default: true.

include_reposts boolean · optional

Default: true.

limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

profiles string[] · required

Minimum items: 1. Maximum items: 20.

reactions_limit integer · optional

Minimum: 0. Maximum: 100. Default: 0.

since string | null · optional

Format: date-time.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].author object | null · optional
items[].comments object[] · optional
items[].comments[].author object | null · optional
items[].comments[].comment_id string | null · optional
items[].comments[].comment_url string | null · optional
items[].comments[].created_at string | null · optional
items[].comments[].metrics object | null · optional
items[].comments[].parent_comment_id string | null · optional
items[].comments[].post object | null · optional
items[].comments[].replies object[] · optional
items[].comments[].replies[].author object | null · optional
items[].comments[].replies[].comment_id string | null · optional
items[].comments[].replies[].comment_url string | null · optional
items[].comments[].replies[].created_at string | null · optional
items[].comments[].replies[].metrics object | null · optional
items[].comments[].replies[].parent_comment_id string | null · optional
items[].comments[].replies[].text string | null · optional
items[].comments[].text string | null · optional
items[].content_type string | null · optional
items[].media object[] · optional
items[].media[].description string | null · optional
items[].media[].image object | null · optional
items[].media[].kind string | null · optional
items[].media[].title string | null · optional
items[].media[].url string | null · optional
items[].metrics object | null · optional
items[].post_id string | null · optional
items[].post_url string | null · optional
items[].posted_at string | null · optional
items[].reactions object[] · optional
items[].reactions[].author object | null · optional
items[].reactions[].created_at string | null · optional
items[].reactions[].kind string | null · optional
items[].reactions[].post object | null · optional
items[].reactions[].reaction_id string | null · optional
items[].text string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/linkedin/profile/reactions

Profile Reactions

Retrieve public LinkedIn data using the documented filters and pagination fields.

JSON body

freshness hour | day | week | month | three_months | six_months | year | null · optional
limit integer · optional

Minimum: 1. Maximum: 1000. Default: 100.

page integer · optional

Minimum: 1. Maximum: 100. Default: 1.

pages integer · optional

Minimum: 1. Maximum: 20. Default: 1.

profiles string[] · required

Minimum items: 1. Maximum items: 20.

since string | null · optional

Format: date-time.

Response · 200

Successful Response

count integer · required
items object[] · required
items[].author object | null · optional
items[].created_at string | null · optional
items[].kind string | null · optional
items[].post object | null · optional
items[].reaction_id string | null · optional
next_page integer | null · optional
page integer | null · optional
total integer | null · optional

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/reddit/post/comments

Comments

Expand accessible nested and collapsed replies while preserving observed parents, paths and deleted placeholders.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Finite total comment bound, including all accessible nested replies and deleted/removed placeholders. Minimum: 1. Maximum: 1000. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

post_url string · required

One explicit public Reddit post permalink. Comment-target, share and redirect URLs are not accepted. Minimum length: 1. Maximum length: 512.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

sort best | top | new | controversial | old | qa · optional

Default: "best".

Response · 200

Successful Response

collapsed_replies_requested true · required
collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.image_url string | null · required
items[].author.is_deleted boolean | null · required
items[].author.profile_id string | null · required
items[].author.profile_url string | null · required
items[].author.username string | null · required
items[].captured_sequence integer · required

Minimum: 0.

items[].children_ids string[] · required

Actual direct children in this collection, using Reddit fullnames. Absence does not imply no further replies exist. Maximum items: 1000.

items[].comment_id string · required
items[].content_state available | deleted | removed | deleted_or_removed | unknown · required
items[].depth integer · required

Minimum: 0. Maximum: 100.

items[].direct_reply_count integer | null · required

Minimum: 0.

items[].edited_at string | null · required
items[].full_id string · required
items[].parent_id string · required
items[].parent_present_in_collection boolean · required

Whether this parent is the requested post or appears anywhere in the cached collection, including other pages.

items[].path string[] · required

Minimum items: 1. Maximum items: 101.

items[].post_id string · required
items[].post_url string · required
items[].relation_evidence observed_parent_and_path · required
items[].score integer | null · required
items[].subreddit string · required
items[].text string | null · required
items[].timestamp string | null · required
items[].url string · required
limit integer · required

Minimum: 1. Maximum: 50.

max_depth 100 · required
missing_ancestor_ids string[] · required

Maximum items: 100000.

nested_replies_requested true · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons unavailable_or_invalid_rows | missing_ancestors | empty_unverified | collection_limit | depth_limit | upstream_incomplete | source_unavailable | summary_unavailable[] · required
post_id string · required
post_url string · required
requested_collection_size integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

sort best | top | new | controversial | old | qa · required
thread_status observed | empty | private | not_found | unavailable | unknown · required
tree_complete boolean | null · required

Accessible upstream tree completeness, when proven by a source summary; null is unknown. Deleted text and private comments are never recovered.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/reddit/search

Search

One finite public keyword search over posts, comments or both, with actual source IDs.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 200. Default: 50.

content_type posts | comments | both · optional

Default: "posts".

include_nsfw boolean · optional

Default: false.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

One Reddit query with optional native operators. Returned rows do not guarantee matching or exhaustive recall. Minimum length: 1. Maximum length: 700.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

sort relevance | new | top | hot | comments · optional

Default: "relevance".

subreddits string[] · optional

At most five explicit public communities. Omit for all-Reddit search; related-community discovery is disabled. Maximum items: 5.

time_filter all | hour | day | week | month | year · optional

Default: "all".

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

content_type posts | comments | both · required
count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required
items object | object[] · required

Maximum items: 50.

items[].author object · required
items[].author.image_url string | null · required
items[].author.is_deleted boolean | null · required
items[].author.profile_id string | null · required
items[].author.profile_url string | null · required
items[].author.username string | null · required
items[].comments_count integer | null · required

Minimum: 0.

items[].content_state available | deleted | removed | deleted_or_removed | unknown · required
items[].full_id string · required
items[].is_nsfw boolean | null · required
items[].kind post · required
items[].outbound_url string | null · required
items[].post_id string · required
items[].query string · required
items[].relation_evidence search_result_permalink · required
items[].score integer | null · required
items[].subreddit string · required
items[].text string | null · required
items[].timestamp string | null · required
items[].title string · required
items[].url string · required
items[].author object · required
items[].author.image_url string | null · required
items[].author.is_deleted boolean | null · required
items[].author.profile_id string | null · required
items[].author.profile_url string | null · required
items[].author.username string | null · required
items[].comment_id string · required
items[].content_state available | deleted | removed | deleted_or_removed | unknown · required
items[].full_id string · required
items[].kind comment · required
items[].parent_evidence post_context | parent_id | unknown · required
items[].parent_id string | null · required
items[].post_id string · required
items[].post_title string | null · required
items[].query string · required
items[].relation_evidence search_result_permalink · required
items[].score integer | null · required
items[].subreddit string · required
items[].text string | null · required
items[].timestamp string | null · required
items[].url string · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons unavailable_or_invalid_rows[] · required
query string · required
query_handling upstream_hint · required
requested_collection_size integer · required

Minimum: 1. Maximum: 200.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

sort relevance | new | top | hot | comments · required
subreddits string[] · required
time_filter all | hour | day | week | month | year · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/sherlock/search

Sherlock Search

Discover public username candidates with Sherlock. Matching handles do not establish identity; completeness may be unknown.

JSON body

Discover candidate accounts for one to five public usernames.

limit integer · optional

Maximum candidates per distinct username. Minimum: 1. Maximum: 100. Default: 20.

usernames string[] · required

Minimum items: 1. Maximum items: 5.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 500.

items object[] · required

Maximum items: 500.

items[].profile_url string · required
items[].site string · required
items[].status candidate · required
items[].username string · required
truncated true | null · required

True when a work or result limit omitted results; null means completeness is unknown. Candidate matches do not establish identity.

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/telegram/group/details

Group Details

Read observed public discussion-group metadata; channels and bots do not become groups.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay, polling and cached pages start no additional work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 1. Default: 1.

group_url string · required

One public group handle, @username or https://t.me/username URL. Invite links, private chats and message links are rejected. Minimum length: 1. Maximum length: 256.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

group_url string · required
items object[] · required

Maximum items: 50.

items[].description string | null · required
items[].group_url string · required
items[].image_url string | null · required
items[].is_verified boolean | null · required
items[].members_count integer | null · required

Minimum: 0.

items[].observed_at string | null · required
items[].online_count integer | null · required

Minimum: 0.

items[].title string · required
items[].username string · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons unavailable_or_invalid_rows[] · required
requested_collection_size integer · required

Minimum: 1. Maximum: 200.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/telegram/group/messages

Group Messages

Collect up to 200 accessible group messages with sender, timestamp, reply links and media references.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay, polling and cached pages start no additional work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 200. Default: 50.

group_url string · required

One public discussion group. Broadcast channels are not accepted as group message data. Minimum length: 1. Maximum length: 256.

keyword string | null · optional

Optional upstream keyword hint in the selected group. Whole-word matching, * endings and + clauses depend on the source. Minimum length: 1. Maximum length: 200.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

since string | null · optional

Inclusive timezone-aware timestamp hint. Results remain a bounded snapshot, not full history. Maximum length: 40.

until string | null · optional

Exclusive timezone-aware timestamp hint, after since when both are present. Maximum length: 40.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required
group_url string · required
history_complete null · required
history_scope public_accessible_snapshot · required
items object[] · required

Maximum items: 50.

items[].edited boolean | null · required
items[].group_url string · required
items[].media object[] · required

Maximum items: 100.

items[].media[].kind photo | video | document | sticker | voice | audio | animation | unknown · required
items[].media[].url string | null · required

Public source reference, which may expire; no media bytes are retained by this operation.

items[].message_id string · required
items[].relation_evidence observed_group_message · required
items[].reply_to_message_id string | null · required
items[].reply_to_url string | null · required
items[].sender object · required
items[].sender.name string | null · required
items[].sender.profile_url string | null · required
items[].sender.username string | null · required
items[].text string | null · required
items[].timestamp string · required
items[].url string · required
keyword string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons unavailable_or_invalid_rows[] · required
requested_collection_size integer · required

Minimum: 1. Maximum: 200.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

since string | null · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

until string | null · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/telegram/groups/search

Group Search

Discover public group candidates by topic or name; no member list or group content is collected.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay, polling and cached pages start no additional work.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 200. Default: 50.

language string | null · optional

Optional two-letter language hint passed to the public index. Pattern: ^[a-z]{2}$.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

One topic or group-name query over a public Telegram index. Results are group candidates, not exhaustive discovery. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

content_retrieved false · required
count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required
items object[] · required

Maximum items: 50.

items[].description string | null · required
items[].group_url string · required
items[].language string | null · required
items[].matched_all_words boolean | null · required
items[].observed_at string | null · required
items[].query string · required
items[].rank integer · required

Minimum: 1.

items[].relation_evidence public_group_index · required
items[].title string · required
items[].username string · required
language string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons unavailable_or_invalid_rows[] · required
query string · required
query_handling upstream_hint · required
requested_collection_size integer · required

Minimum: 1. Maximum: 200.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

search_source public_web_index · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/tiktok/people/search

Tiktok People Search

Collect at most 50 native user candidates with available profile fields. Location texts are appended in caller order as query hints, never enforced city filters. Prefer: respond-async returns promptly; default wait is at most 20 seconds. Cached pages/replays launch no new work.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

locations string[] · optional

Up to five location hints of 1–100 characters, appended in caller order to one native query. Ambiguous names remain text; no city resolution or AND/OR geographic filters are enforced. Maximum items: 5.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

Free-text name, optionally with employer/city. Echoed exactly; matching does not verify identity or enforce filters. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

effective_query string · required

Text sent to native account search, including location hints. Instagram commas become spaces to keep one query.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required

City and employer restrictions are not enforced. Returned accounts are candidates, never verified identities.

items object[] · required

Maximum items: 50.

items[].biography string | null · optional
items[].created_at string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].friends_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].is_organization boolean | null · optional
items[].is_private boolean | null · optional
items[].is_seller boolean | null · optional
items[].is_verified boolean | null · optional
items[].language string | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].missing_fields display_name | biography | image_url[] · required

Unavailable core profile fields. No inferred or purchased enrichment. Maximum items: 3.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].query string · required
items[].rank integer · required

Absolute one-based position after canonical account deduplication, preserving native order. Minimum: 1. Maximum: 50.

items[].status candidate · required
items[].username string · required
items[].website_url string | null · optional
limit integer · required

Minimum: 1. Maximum: 50.

location_handling not_requested | query_hint · required
locations string[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

query string · required
result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Durable native user search; poll the owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/tiktok/profile/details

Tiktok Profile Details

Retrieve public TikTok profile details for up to ten usernames or profile URLs.

JSON body

One to ten TikTok usernames or canonical public @username URLs.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 10.

items object[] · required

Maximum items: 10.

items[].biography string | null · optional
items[].created_at string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].friends_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_organization boolean | null · optional
items[].is_private boolean | null · optional
items[].is_seller boolean | null · optional
items[].is_verified boolean | null · optional
items[].language string | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].target string · required
items[].username string · required
items[].website_url string | null · optional
unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/tiktok/profile/posts

Tiktok Profile Posts

Return public videos and ordered slideshow images, with separate cover thumbnails and source authors. Prefer: respond-async submits promptly.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

Collect recent videos or a dated window of public videos and slideshows.

archive boolean · optional

Collect one dated profile window. Follow next_archive_end_date with a new Idempotency-Key to continue backward. Default: false.

archive_end_date string | null · optional

Inclusive UTC end date; required with archive=true. Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_window_days integer · optional

Inclusive days in the archive window. Narrow a window that reaches the collection cap. Minimum: 1. Maximum: 30. Default: 30.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

per_target_limit integer · optional

Maximum videos per profile: 50 by default or up to 1,000 in one dated archive window. Archive mode defaults to 1,000 if omitted. Minimum: 1. Maximum: 1000. Default: 50.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

retain_media boolean · optional

Preserve photos/video thumbnails. Recent jobs allow 20 unique images and 32 MiB over 30 seconds; archive jobs allow 100 images and 96 MiB over 120 seconds. Each image is capped at 10 MiB and 40 million pixels. Shared bearer/global quotas still apply. No video download; assets expire after 24 hours. Cached pages/replay never download again. Default: false.

Response · 200

Successful Response

archive_end_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_start_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

history_complete null · required

Full platform history is unknown. Pages cover only the cached collection.

history_scope collected_window | recent_public_timeline | archive_date_window | archive_cursor_window · optional

Date windows continue through next_archive_end_date; cursor windows continue through next_archive_cursor. Page numbers only traverse the cached result. Default: "collected_window".

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].context object[] · required

Maximum items: 2.

items[].context[].author object · required
items[].context[].author.display_name string | null · optional
items[].context[].author.image_url string | null · optional
items[].context[].author.profile_id string | null · optional
items[].context[].author.profile_url string · required
items[].context[].author.username string | null · optional
items[].context[].comments_count integer | null · optional

Minimum: 0.

items[].context[].geotag object | null · required
items[].context[].kind repost | quote · required
items[].context[].likes_count integer | null · optional

Minimum: 0.

items[].context[].media object[] · required

Maximum items: 500.

items[].context[].media[].asset object | null · optional
items[].context[].media[].bitrate integer | null · optional

Minimum: 0.

items[].context[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].context[].media[].mime_type string | null · optional
items[].context[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].context[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].context[].media[].role original | thumbnail · required
items[].context[].media[].type image | video · required
items[].context[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].context[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].context[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].context[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].context[].post_id string · required
items[].context[].quotes_count integer | null · optional

Minimum: 0.

items[].context[].reposts_count integer | null · optional

Minimum: 0.

items[].context[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].context[].text string | null · required
items[].context[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].context[].url string · required
items[].context[].views_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_pinned boolean | null · optional
items[].is_quote boolean | null · optional
items[].is_repost boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].target string · required
items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_archive_cursor string | null · optional

Submit as archive_cursor in a new Instagram or Facebook personal archive job with the same profile and window settings after consuming cached pages. Null when the provider reached the visible end or could not safely continue. Maximum length: 262144.

next_archive_end_date string | null · optional

Submit a new archive job ending on this date after consuming all cached pages. Null when the date range is exhausted or the current window hit its collection cap; narrow a capped window before continuing. Pattern: ^\d{4}-\d{2}-\d{2}$.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

per_target_limit integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

targets object[] · required

Maximum items: 10.

targets[].collected_count integer · required

Minimum: 0. Maximum: 1000.

targets[].input_index integer · required

Minimum: 0. Maximum: 9.

targets[].reason not_found | private | unavailable | malformed_data | ambiguous | null · required
targets[].status resolved | unresolved · required
targets[].target string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Response · 202

Durable job; poll the returned owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/people/search

X People Search

Collect at most 50 native account candidates with available profile fields. Locations are query hints, not filters. Prefer: respond-async returns promptly; default wait is at most 20 seconds. Cached pages/replays launch no new work.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

locations string[] · optional

Up to five location hints of 1–100 characters, appended in caller order to one native query. Ambiguous names remain text; no city resolution or AND/OR geographic filters are enforced. Maximum items: 5.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · required

Free-text name, optionally with employer/city. Echoed exactly; matching does not verify identity or enforce filters. Minimum length: 1. Maximum length: 200.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

effective_query string · required

Text sent to native account search, including location hints. Instagram commas become spaces to keep one query.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required

City and employer restrictions are not enforced. Returned accounts are candidates, never verified identities.

items object[] · required

Maximum items: 50.

items[].biography string | null · optional
items[].cover_image_url string | null · optional
items[].created_at string | null · optional

Account creation timestamp in ISO 8601 UTC.

items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].is_blue_verified boolean | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional

Verification observation; absent or null when the source supplies only an ambiguous untyped flag.

items[].likes_count integer | null · optional

Minimum: 0.

items[].listed_count integer | null · optional

Minimum: 0.

items[].location string | null · optional
items[].media_count integer | null · optional

Minimum: 0.

items[].missing_fields display_name | biography | image_url[] · required

Unavailable core profile fields. No inferred or purchased enrichment. Maximum items: 3.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].query string · required
items[].rank integer · required

Absolute one-based position after canonical account deduplication, preserving native order. Minimum: 1. Maximum: 50.

items[].status candidate · required
items[].username string · required
items[].verification_type string | null · optional
items[].website_url string | null · optional
limit integer · required

Minimum: 1. Maximum: 50.

location_handling not_requested | query_hint · required
locations string[] · required
next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

query string · required
result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Durable native account search; poll the owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/post/quotes

Quotes

Collect incoming quotes whose returned original-post ID matches the requested target. One bounded collection, with unknown full coverage.

Parameters

Idempotency-Key header · optional

Required for initial collection; result_id pages require no key.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 500. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

post_id string · required

Pattern: ^[1-9][0-9]{0,24}$.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].conversation_id string | null · required
items[].geotag object | null · required
items[].in_reply_to_post_id string | null · required

Actual reply parent when returned; a quote-provider target is never reused as a reply parent.

items[].is_quote boolean | null · required
items[].is_reply boolean | null · required
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quoted_post_id string · required

Actual original-post ID returned with this quote. It must match the requested target. Pattern: ^[1-9][0-9]{0,24}$.

items[].quotes_count integer | null · optional

Minimum: 0.

items[].relation_evidence provider_quote_target · required

The returned row explicitly names the original post in the reviewed quote-provider output; request echo and search membership alone cannot establish this relationship.

items[].relation_kind incoming_quote · required
items[].relation_target_post_id string · required

Pattern: ^[1-9][0-9]{0,24}$.

items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons rejected_rows[] · required
quotes_complete null · required

Full incoming quote coverage is unknown, including when cached next_page is null.

requested_collection_size integer · required

Minimum: 1. Maximum: 500.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

target_post_id string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

upstream_continuation_available false · required

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/post/replies

Replies

Search conversation candidates by post ID. Only returned parent IDs establish direct or nested replies; this is not a complete reply tree.

Parameters

Idempotency-Key header · optional

Required for initial collection; result_id pages require no key.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 800. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

post_id string · required

Pattern: ^[1-9][0-9]{0,24}$.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

effective_query string · required
expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required

Search operators are sent upstream; matching and recall are not independently guaranteed.

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].conversation_id string | null · required
items[].effective_query string · required
items[].geotag object | null · required
items[].in_reply_to_post_id string | null · required

Actual returned parent ID; null is unknown and never inferred from conversation membership.

items[].is_quote boolean | null · required
items[].is_reply boolean | null · required
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quoted_post_id string | null · required

Actual returned outgoing quote target; this does not enumerate incoming quotes.

items[].quotes_count integer | null · optional

Minimum: 0.

items[].relation_evidence search_query | parent_id | conversation_and_parent_id · required
items[].relation_kind search_match | conversation_candidate | direct_reply | nested_reply · required
items[].relation_target_post_id string | null · required
items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons rejected_rows_or_unknown_parent[] · required
query_scope post_search | conversation_candidates · required
requested_collection_size integer · required

Minimum: 1. Maximum: 800.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

search_complete null · required
target_post_id string | null · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/posts/search

Post Search

One bounded X web query with preserved source authors; operators and recall remain upstream-dependent.

Parameters

Idempotency-Key header · optional

Required for initial collection; result_id pages require no key.

prefer header · optional

JSON body

collection_size integer · optional

Minimum: 1. Maximum: 800. Default: 100.

from_username string | null · optional

Maximum length: 256.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

mention_username string | null · optional

Maximum length: 256.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

query string · optional

One X web-search query. from:, to:, @mention, since: and until: may be supplied here or through the structured fields. Search recall is unknown. Maximum length: 512. Default: "".

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

since_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

to_username string | null · optional

Maximum length: 256.

until_date string | null · optional

Exclusive upper UTC date in the effective search query. Pattern: ^\d{4}-\d{2}-\d{2}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

effective_query string · required
expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

filters_enforced false · required

Search operators are sent upstream; matching and recall are not independently guaranteed.

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].conversation_id string | null · required
items[].effective_query string · required
items[].geotag object | null · required
items[].in_reply_to_post_id string | null · required

Actual returned parent ID; null is unknown and never inferred from conversation membership.

items[].is_quote boolean | null · required
items[].is_reply boolean | null · required
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quoted_post_id string | null · required

Actual returned outgoing quote target; this does not enumerate incoming quotes.

items[].quotes_count integer | null · optional

Minimum: 0.

items[].relation_evidence search_query | parent_id | conversation_and_parent_id · required
items[].relation_kind search_match | conversation_candidate | direct_reply | nested_reply · required
items[].relation_target_post_id string | null · required
items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons rejected_rows_or_unknown_parent[] · required
query_scope post_search | conversation_candidates · required
requested_collection_size integer · required

Minimum: 1. Maximum: 800.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

search_complete null · required
target_post_id string | null · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/details

X Profile Details

Retrieve public X profile details for up to ten handles or profile URLs.

JSON body

One to ten X handles or canonical x.com or twitter.com profile URLs.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

Response · 200

Successful Response

count integer · required

Minimum: 0. Maximum: 10.

items object[] · required

Maximum items: 10.

items[].biography string | null · optional
items[].cover_image_url string | null · optional
items[].created_at string | null · optional

Account creation timestamp in ISO 8601 UTC.

items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_blue_verified boolean | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional

Verification observation; absent or null when the source supplies only an ambiguous untyped flag.

items[].likes_count integer | null · optional

Minimum: 0.

items[].listed_count integer | null · optional

Minimum: 0.

items[].location string | null · optional
items[].media_count integer | null · optional

Minimum: 0.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].target string · required
items[].username string · required
items[].verification_type string | null · optional
items[].website_url string | null · optional
unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/dumps

Submit X Dump

Start a resumable signed-in browser collection. A finished timeline does not prove complete account history.

Parameters

Idempotency-Key header · optional

JSON body

Collect signed-in browser timelines for one public X profile.

profiles string[] · required

Minimum items: 1. Maximum items: 1.

tabs tweets | legacy_tweets | with_replies | reposts | media | photos | legacy_media[] · optional

Minimum items: 1. Maximum items: 7.

Response · 200

Successful Response

created_at number · required
error_code string | null · required
history_coverage unverified | below_profile_count | matches_profile_count | above_profile_count · required

A count comparison, never a guarantee of full history.

job_id string · required
max_runs integer · optional

Default: 32.

media_count integer · required
media_post_count integer · required
next_run_at number | null · optional
post_count integer · required
profile string · required
profile_reported_posts integer | null · required
revision integer · optional

Default: 1.

run_count integer · optional

Default: 0.

status pending | running | partial | timelines_exhausted · required
tabs object[] · required
tabs[].end_page integer | null · optional
tabs[].last_page integer · required
tabs[].no_new_pages integer · required
tabs[].status string · required
tabs[].tab tweets | legacy_tweets | with_replies | reposts | media | photos | legacy_media · required
unreconciled_post_count integer | null · required
updated_at number · required

Response · 202

Accepted

created_at number · required
error_code string | null · required
history_coverage unverified | below_profile_count | matches_profile_count | above_profile_count · required

A count comparison, never a guarantee of full history.

job_id string · required
max_runs integer · optional

Default: 32.

media_count integer · required
media_post_count integer · required
next_run_at number | null · optional
post_count integer · required
profile string · required
profile_reported_posts integer | null · required
revision integer · optional

Default: 1.

run_count integer · optional

Default: 0.

status pending | running | partial | timelines_exhausted · required
tabs object[] · required
tabs[].end_page integer | null · optional
tabs[].last_page integer · required
tabs[].no_new_pages integer · required
tabs[].status string · required
tabs[].tab tweets | legacy_tweets | with_replies | reposts | media | photos | legacy_media · required
unreconciled_post_count integer | null · required
updated_at number · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /scrapers/x/profile/dumps/{job_id}

Get X Dump

Read progress. wait_seconds waits for completion; after_revision waits for a progress change instead.

Parameters

job_id path · required
wait_seconds query · optional
after_revision query · optional

Response · 200

Successful Response

created_at number · required
error_code string | null · required
history_coverage unverified | below_profile_count | matches_profile_count | above_profile_count · required

A count comparison, never a guarantee of full history.

job_id string · required
max_runs integer · optional

Default: 32.

media_count integer · required
media_post_count integer · required
next_run_at number | null · optional
post_count integer · required
profile string · required
profile_reported_posts integer | null · required
revision integer · optional

Default: 1.

run_count integer · optional

Default: 0.

status pending | running | partial | timelines_exhausted · required
tabs object[] · required
tabs[].end_page integer | null · optional
tabs[].last_page integer · required
tabs[].no_new_pages integer · required
tabs[].status string · required
tabs[].tab tweets | legacy_tweets | with_replies | reposts | media | photos | legacy_media · required
unreconciled_post_count integer | null · required
updated_at number · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /scrapers/x/profile/dumps/{job_id}/posts

Get X Dump Posts

Page deduplicated posts and media source URLs; start at after=0 and follow next_after.

Parameters

job_id path · required
after query · optional
limit query · optional

Response · 200

Successful Response

has_more boolean · required

More posts are cached now; poll the job while it is still running.

job_id string · required
next_after integer · required
posts object[] · required
posts[].author string · required
posts[].created_at string | null · required
posts[].in_reply_to_post_id string | null · optional
posts[].in_reply_to_user_id string | null · optional
posts[].is_quote boolean · required
posts[].is_reply boolean · required
posts[].is_repost boolean · required
posts[].media object[] · required
posts[].media[].media_id string | null · optional
posts[].media[].type photo | video | animated_gif · required
posts[].media[].url string · required
posts[].media[].video_variants object[] | null · optional
posts[].post_id string · required
posts[].text string | null · required
posts[].url string · required
status pending | running | partial | timelines_exhausted · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/dumps/{job_id}/resume

Resume X Dump

Continue selected paused tabs at saved cursors. Optional page bounds cannot skip or rewind upstream pages.

Parameters

job_id path · required
retry_stalled query · optional

JSON body

end_page integer | null · optional

Inclusive stop page; reaching this bound leaves the job partial. Minimum: 1. Maximum: 1000000.

start_page integer | null · optional

Must equal last_page + 1 for every selected tab; cursors cannot skip pages. Minimum: 1. Maximum: 1000000.

tabs tweets | legacy_tweets | with_replies | reposts | media | photos | legacy_media[] | null · optional

Minimum items: 1. Maximum items: 7.

See the lifecycle example for this operation.

Response · 200

Successful Response

created_at number · required
error_code string | null · required
history_coverage unverified | below_profile_count | matches_profile_count | above_profile_count · required

A count comparison, never a guarantee of full history.

job_id string · required
max_runs integer · optional

Default: 32.

media_count integer · required
media_post_count integer · required
next_run_at number | null · optional
post_count integer · required
profile string · required
profile_reported_posts integer | null · required
revision integer · optional

Default: 1.

run_count integer · optional

Default: 0.

status pending | running | partial | timelines_exhausted · required
tabs object[] · required
tabs[].end_page integer | null · optional
tabs[].last_page integer · required
tabs[].no_new_pages integer · required
tabs[].status string · required
tabs[].tab tweets | legacy_tweets | with_replies | reposts | media | photos | legacy_media · required
unreconciled_post_count integer | null · required
updated_at number · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/followers

Followers

Collect attributed public followers; nullable profile fields and unknown full graph. Pages read the cached collection.

Parameters

Idempotency-Key header · optional

Required for initial collection; result_id pages require no key.

prefer header · optional

JSON body

collection_size integer · optional

Maximum rows in one public graph collection; page limit only pages that immutable result. No upstream graph continuation is available. Minimum: 5. Maximum: 2000. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 256.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

graph_complete null · required

Full public graph completeness is unknown, including when next_page is null.

items object[] · required

Maximum items: 50.

items[].biography string | null · optional
items[].cover_image_url string | null · optional
items[].created_at string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].is_blue_verified boolean | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].listed_count integer | null · optional

Minimum: 0.

items[].location string | null · optional
items[].media_count integer | null · optional

Minimum: 0.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].relation followers | following · required
items[].relation_evidence source_attribution | provider_expansion · required

Source attribution is explicit in the row; provider expansion requires a privately reviewed single-source/direction contract and fixture. The source root is excluded.

items[].source_profile string · required
items[].username string · required
items[].verification_type string | null · optional
items[].website_url string | null · optional
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons rejected_rows | missing_optional_fields[] · required
relation followers | following · required
requested_collection_size integer · required

Minimum: 5. Maximum: 2000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/following

Following

Conditional public following collection. Unqualified ambiguous output fails; the source root cannot become an edge.

Parameters

Idempotency-Key header · optional

Required for initial collection; result_id pages require no key.

prefer header · optional

JSON body

collection_size integer · optional

Maximum rows in one public graph collection; page limit only pages that immutable result. No upstream graph continuation is available. Minimum: 5. Maximum: 2000. Default: 100.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 256.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

graph_complete null · required

Full public graph completeness is unknown, including when next_page is null.

items object[] · required

Maximum items: 50.

items[].biography string | null · optional
items[].cover_image_url string | null · optional
items[].created_at string | null · optional
items[].display_name string | null · optional
items[].followers_count integer | null · optional

Minimum: 0.

items[].following_count integer | null · optional

Minimum: 0.

items[].image_url string | null · optional
items[].is_blue_verified boolean | null · optional
items[].is_private boolean | null · optional
items[].is_verified boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].listed_count integer | null · optional

Minimum: 0.

items[].location string | null · optional
items[].media_count integer | null · optional

Minimum: 0.

items[].posts_count integer | null · optional

Minimum: 0.

items[].profile_id string | null · optional
items[].profile_url string · required
items[].public_links object[] | null · optional
items[].relation followers | following · required
items[].relation_evidence source_attribution | provider_expansion · required

Source attribution is explicit in the row; provider expansion requires a privately reviewed single-source/direction contract and fixture. The source root is excluded.

items[].source_profile string · required
items[].username string · required
items[].verification_type string | null · optional
items[].website_url string | null · optional
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons rejected_rows | missing_optional_fields[] · required
relation followers | following · required
requested_collection_size integer · required

Minimum: 5. Maximum: 2000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/posts

X Profile Posts

Return recent posts or a dated X archive window. Cached pages read one immutable result; next_archive_end_date starts the next older job. Prefer: respond-async submits promptly.

Parameters

Idempotency-Key header · optional

Required for initial collection; cached result_id pages need no key.

prefer header · optional

JSON body

Collect recent X posts or a dated archive window, then page the cached result.

archive boolean · optional

Search one profile by date instead of collecting only its recent timeline. Use next_archive_end_date from the result to request the preceding window with a new Idempotency-Key. Default: false.

archive_end_date string | null · optional

Inclusive UTC date at the end of this archive window. Required with archive=true. Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_window_days integer · optional

Number of inclusive days in one archive search. Use a smaller window if truncated is true; cached page numbers do not extend history. Minimum: 1. Maximum: 30. Default: 30.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

media_filter any | images | videos · optional

Archive search filter. Use images to find older photo posts without scrolling through text replies. Default: "any".

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

per_target_limit integer · optional

Maximum posts per profile: 50 by default, up to 100 for recent mode and 1,000 for one dated archive window. Archive mode defaults to 1,000 if omitted. Minimum: 1. Maximum: 1000. Default: 50.

profiles string[] · required

Minimum items: 1. Maximum items: 10.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

retain_media boolean · optional

Preserve photos/video thumbnails. Recent jobs allow 20 unique images and 32 MiB over 30 seconds; archive jobs allow 100 images and 96 MiB over 120 seconds. Each image is capped at 10 MiB and 40 million pixels. Shared bearer/global quotas still apply. No video download; assets expire after 24 hours. Cached pages/replay never download again. Default: false.

Response · 200

Successful Response

archive_end_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

archive_start_date string | null · optional

Pattern: ^\d{4}-\d{2}-\d{2}$.

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

history_complete null · required

Full platform history is unknown. Pages cover only the cached collection.

history_scope collected_window | recent_public_timeline | archive_date_window | archive_cursor_window · optional

Date windows continue through next_archive_end_date; cursor windows continue through next_archive_cursor. Page numbers only traverse the cached result. Default: "collected_window".

items object[] · required

Maximum items: 50.

items[].author object · required
items[].author.display_name string | null · optional
items[].author.image_url string | null · optional
items[].author.profile_id string | null · optional
items[].author.profile_url string · required
items[].author.username string | null · optional
items[].comments_count integer | null · optional

Minimum: 0.

items[].context object[] · required

Maximum items: 2.

items[].context[].author object · required
items[].context[].author.display_name string | null · optional
items[].context[].author.image_url string | null · optional
items[].context[].author.profile_id string | null · optional
items[].context[].author.profile_url string · required
items[].context[].author.username string | null · optional
items[].context[].comments_count integer | null · optional

Minimum: 0.

items[].context[].geotag object | null · required
items[].context[].kind repost | quote · required
items[].context[].likes_count integer | null · optional

Minimum: 0.

items[].context[].media object[] · required

Maximum items: 500.

items[].context[].media[].asset object | null · optional
items[].context[].media[].bitrate integer | null · optional

Minimum: 0.

items[].context[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].context[].media[].mime_type string | null · optional
items[].context[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].context[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].context[].media[].role original | thumbnail · required
items[].context[].media[].type image | video · required
items[].context[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].context[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].context[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].context[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].context[].post_id string · required
items[].context[].quotes_count integer | null · optional

Minimum: 0.

items[].context[].reposts_count integer | null · optional

Minimum: 0.

items[].context[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].context[].text string | null · required
items[].context[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].context[].url string · required
items[].context[].views_count integer | null · optional

Minimum: 0.

items[].geotag object | null · required
items[].input_index integer · required

Minimum: 0. Maximum: 9.

items[].is_pinned boolean | null · optional
items[].is_quote boolean | null · optional
items[].is_repost boolean | null · optional
items[].likes_count integer | null · optional

Minimum: 0.

items[].media object[] · required

Maximum items: 500.

items[].media[].asset object | null · optional
items[].media[].bitrate integer | null · optional

Minimum: 0.

items[].media[].group_index integer · required

Zero-based position in the source carousel or album; variants share a group. Minimum: 0. Maximum: 499.

items[].media[].mime_type string | null · optional
items[].media[].retention_reason not_requested | video_not_retained | count_limit | job_byte_limit | destination_rejected | source_unavailable | unsupported_format | invalid_image | too_large | quota_exceeded | timeout | download_failed | storage_unavailable | null · optional

Default: "not_requested".

items[].media[].retention_status retained | skipped | failed · optional

Default: "skipped".

items[].media[].role original | thumbnail · required
items[].media[].type image | video · required
items[].media[].url string · required

Original public source reference; may expire. Separate from the authenticated retained asset URL.

items[].media_complete boolean | null · optional

False when the provider reports media missing from this post; null when completeness cannot be established.

items[].media_count_reported integer | null · optional

Media count reported by the source provider, when available. Minimum: 0.

items[].mentions object[] | null · required

Caption/text mentions, separate from actual tags. Maximum items: 100.

items[].post_id string · required
items[].quotes_count integer | null · optional

Minimum: 0.

items[].reposts_count integer | null · optional

Minimum: 0.

items[].tagged_users object[] | null · required

Actual source tags; null means unavailable. Maximum items: 100.

items[].target string · required
items[].text string | null · required
items[].timestamp string | null · required

UTC ISO 8601 time; null means unavailable, never inferred epoch zero.

items[].url string · required
items[].views_count integer | null · optional

Minimum: 0.

limit integer · required

Minimum: 1. Maximum: 50.

next_archive_cursor string | null · optional

Submit as archive_cursor in a new Instagram or Facebook personal archive job with the same profile and window settings after consuming cached pages. Null when the provider reached the visible end or could not safely continue. Maximum length: 262144.

next_archive_end_date string | null · optional

Submit a new archive job ending on this date after consuming all cached pages. Null when the date range is exhausted or the current window hit its collection cap; narrow a capped window before continuing. Pattern: ^\d{4}-\d{2}-\d{2}$.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

per_target_limit integer · required

Minimum: 1. Maximum: 1000.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

targets object[] · required

Maximum items: 10.

targets[].collected_count integer · required

Minimum: 0. Maximum: 1000.

targets[].input_index integer · required

Minimum: 0. Maximum: 9.

targets[].reason not_found | private | unavailable | malformed_data | ambiguous | null · required
targets[].status resolved | unresolved · required
targets[].target string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

unresolved object[] · required

Maximum items: 10.

unresolved[].input_index integer · required

Minimum: 0. Maximum: 9.

unresolved[].reason not_found | private | unavailable | malformed_data | ambiguous · required
unresolved[].target string · required

Response · 202

Durable job; poll the returned owner-scoped URL.

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/profile/spaces

Discovery

Find actual Space links in posts by a handle; sharing alone does not establish hosting.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

collection_size integer · optional

Finite authored-post search window. limit only pages the saved collection. Minimum: 1. Maximum: 100. Default: 20.

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

profile string · required

Minimum length: 1. Maximum length: 512.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

Response · 200

Successful Response

collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

discovery_scope authored_space_links · required
effective_query string · required
expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].host_relationship null · required
items[].post_id string · required
items[].post_url string · required
items[].relation_evidence authored_post_space_link · required
items[].shared_by string · required
items[].space_urls string[] · required

Minimum items: 1. Maximum items: 20.

items[].text string | null · required
items[].timestamp string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

requested_collection_size integer · required

Minimum: 1. Maximum: 100.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

search_complete null · required
source_profile string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/space/recording

Recording

Resolve one observed replay reference and check a bounded audio sample; bytes are not retained.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

space_id string · required

One explicit public Space ID or canonical x.com/i/spaces URL. Minimum length: 8. Maximum length: 512.

Response · 200

Successful Response

audio_retained false · required
collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].access_evidence playlist_and_audio_sample | null · required
items[].audio_retained false · required
items[].creator object | null · required
items[].ended_at string | null · required
items[].hosts object[] · required

Maximum items: 100.

items[].hosts[].display_name string | null · required
items[].hosts[].image_url string | null · required
items[].hosts[].profile_id string | null · required
items[].hosts[].profile_url string · required
items[].hosts[].username string · required
items[].recording_status accessible | expired | inaccessible | missing_reference | unverified · required
items[].recording_url string | null · required
items[].relation_evidence observed_space_id · required
items[].replay_reported boolean | null · required
items[].space_id string · required
items[].space_url string · required
items[].speakers object[] · required

Maximum items: 100.

items[].speakers[].display_name string | null · required
items[].speakers[].image_url string | null · required
items[].speakers[].profile_id string | null · required
items[].speakers[].profile_url string · required
items[].speakers[].username string · required
items[].started_at string | null · required
items[].state ended | live | scheduled | unknown · required
items[].title string | null · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

space_id string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

POST /scrapers/x/space/transcript

Transcript

Transcribe a finite recorded Space prefix with real returned text and optional timed segments.

Parameters

Idempotency-Key header · optional

Required for initial collection; replay and cached pages start no new work.

prefer header · optional

JSON body

limit integer · optional

Minimum: 1. Maximum: 50. Default: 20.

max_minutes integer · optional

Finite prefix transcription bound; longer recordings are clipped. No full-recording guarantee. Minimum: 1. Maximum: 30. Default: 5.

page integer · optional

One-based cached page. Initial collection requires page 1; later pages require result_id and cannot extend the collected window. Minimum: 1. Maximum: 500. Default: 1.

result_id string | null · optional

Owner- and operation-bound immutable collection. Keep the same business request, ordered inputs and limit. Cached pages require no Idempotency-Key and launch no additional work. Pattern: ^rresult_[a-f0-9]{32}$.

space_id string · required

One explicit public Space ID or canonical x.com/i/spaces URL. Minimum length: 8. Maximum length: 512.

Response · 200

Successful Response

audio_retained false · required
collected_count integer · required

Minimum: 0. Maximum: 2000.

collection_complete null · required
collection_limit integer · required

Minimum: 1. Maximum: 2000.

count integer · required

Minimum: 0. Maximum: 50.

expires_at integer · required

Unix seconds; the collection expires 24 hours after submission. Polling, replay and cached pages do not renew it.

items object[] · required

Maximum items: 50.

items[].audio_retained false · required
items[].duration_seconds number | null · required

Minimum: 0.

items[].language string | null · required
items[].method speech_to_text | null · required
items[].prefix_limited boolean · required
items[].relation_evidence source_url | source_request · required
items[].segments object[] | null · required

Maximum items: 2000.

items[].space_id string · required
items[].space_url string · required
items[].text string | null · required

Maximum length: 200000.

items[].title string | null · required
items[].transcribed_seconds number | null · required
items[].transcript_complete null · required
items[].transcript_status available | unavailable | missing_transcript · required
limit integer · required

Minimum: 1. Maximum: 50.

next_page integer | null · required

Minimum: 2. Maximum: 500.

page integer · required

Minimum: 1. Maximum: 500.

partial boolean · required

Collection-wide flag repeated on every cached page: missing candidate core fields, unresolved post targets, or image retention failures/limits can set it, even when absent from this page.

partial_reasons string[] · required

Maximum items: 10.

requested_max_minutes integer · required

Minimum: 1. Maximum: 30.

result_id string · required

Pattern: ^rresult_[a-f0-9]{32}$.

space_id string · required
truncated true | null · required

True when collection omitted results; null means completeness is unknown.

Response · 202

Accepted

created_at integer · required
error string | null · optional
expires_at integer · required
job_id string · required
operation string · required
poll_url string · required

Relative Hive URL. Poll with the same bearer; wait_seconds up to 25 waits for completion. GET returns this envelope (200), with the completed first page under result. Terminal jobs never launch another paid run on replay; eligible saved archive data has an explicit recover endpoint.

result object | null · optional
retry_after_s integer | null · optional
service string · required
status pending | running | completed | failed | outcome_unknown · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 404 · Public profile not found
  • HTTP 409 · Conflicting immutable request or unavailable recovery
  • HTTP 410 · Research job or result expired
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out

GET /statusz

Read service availability

Response · 200

Success

Errors

  • HTTP 400 · Invalid request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 409 · Conflicting state or idempotency key
  • HTTP 410 · Result expired
  • HTTP 503 · Requested service or constraints unavailable

POST /web/wayback

Get Wayback

Resolve one existing HTML capture within a year or date. Blocked or throttled archive access never means no capture exists.

JSON body

date string | null · optional

YYYY-MM-DD; provide exactly one year or date. Pattern: ^\d{4}-\d{2}-\d{2}$.

url string · required

Exact public HTTP/HTTPS URL to look up; it is never fetched from its live host. Minimum length: 1. Maximum length: 2048.

year integer | null · optional

Minimum: 1996. Maximum: 2100.

Response · 200

Successful Response

archive_url string · required
cache_expires_at number · required
cache_hit boolean · required
capture_datetime string · required
capture_timestamp string · required
capture_validation cdx_and_replay_url | cdx_replay_and_memento_datetime · required
cdx_digest string · required
content_type string · required
decoding_lossy boolean · required
encoding string · required
history_complete null · optional

Unknown; this lookup does not enumerate all historical captures.

html string · required

Decoded identity replay HTML, with no local rewriting or script execution.

html_base64 string · required

Exact decoded HTTP entity bytes; decode base64 before comparing sha256.

replay_mode id_ · required
requested_period string · required
selected_capture_timestamp string · required
selection earliest_returned_capture | archive_redirect_to_indexed_capture · required
sha256 string · required
size_bytes integer · required
text string · required

Text extracted from HTML without executing JavaScript or loading assets.

text_truncated boolean · required
url string · required

Errors

  • HTTP 400 · Invalid research request
  • HTTP 401 · Invalid Hive authentication
  • HTTP 403 · Archive access blocked
  • HTTP 404 · Public profile not found
  • HTTP 429 · Research rate limited
  • HTTP 503 · Research unavailable
  • HTTP 504 · Research timed out