From d11dffc779cc3938282455aa86ef6927a96b95bb Mon Sep 17 00:00:00 2001 From: aadithyan rajesh Date: Mon, 28 Sep 2026 01:09:22 +0530 Subject: [PATCH 1/2] feat(plugin): align workflows with current MCP tools --- .cursor-plugin/plugin.json | 2 +- README.md | 21 +++++++++++---------- commands/extract-web-data.md | 2 +- commands/scrape-url.md | 2 +- rules/prefer-context-dev-mcp.mdc | 13 +++++++------ scripts/validate-plugin.mjs | 19 +++++++++++-------- skills/connect-context-dev/SKILL.md | 2 +- skills/context-brand/SKILL.md | 5 ++--- skills/context-crawl/SKILL.md | 6 +++--- skills/context-extract/SKILL.md | 6 +++--- skills/context-scrape/SKILL.md | 13 +++++++------ skills/context-search/SKILL.md | 19 +++++++++++-------- 12 files changed, 59 insertions(+), 51 deletions(-) diff --git a/.cursor-plugin/plugin.json b/.cursor-plugin/plugin.json index e939ae7..fbdeaf5 100644 --- a/.cursor-plugin/plugin.json +++ b/.cursor-plugin/plugin.json @@ -1,7 +1,7 @@ { "name": "context-dev", "displayName": "Context.dev", - "version": "2.1.0", + "version": "2.2.0", "description": "Search, scrape, crawl, extract, parse, monitor, and process the live web with Context.dev.", "author": { "name": "Context.dev", diff --git a/README.md b/README.md index 83d3381..cc6b715 100644 --- a/README.md +++ b/README.md @@ -1,6 +1,6 @@ # Context.dev for Cursor -The official [Context.dev](https://context.dev) plugin for Cursor. Give Cursor reliable access to the live web: search, scraping, crawling, structured extraction, document parsing, brand intelligence, screenshots, recurring monitors, and large asynchronous batches. +The official [Context.dev](https://context.dev) plugin for Cursor. Give Cursor reliable access to the live web: search, company news, scraping, crawling, structured extraction, document parsing, brand intelligence, screenshots, recurring monitors, and large asynchronous batches. ## Install and connect @@ -34,7 +34,7 @@ Cursor automatically selects the appropriate Context.dev tool. Tool calls requir | Component | What it provides | | --- | --- | -| MCP server | The production Context.dev MCP with 34 direct, typed tools and OAuth | +| MCP server | The production Context.dev MCP with 40 direct, typed tools and OAuth | | Skills | Focused MCP workflows, direct API guidance, Cursor connection help, and Logo Link integration | | Commands | `/brand-colors`, `/scrape-url`, `/search-web`, and `/extract-web-data` | | Rules | Routes live-web tasks to the right Context.dev tool and keeps credentials out of client code | @@ -44,9 +44,9 @@ Cursor automatically selects the appropriate Context.dev tool. Tool calls requir | Skill | When Cursor uses it | | --- | --- | | `context-search` | Live web research and source discovery | -| `context-scrape` | Markdown, HTML, images, or screenshots from one known URL | -| `context-crawl` | Sitemap discovery and focused multi-page crawling | -| `context-extract` | Schema-shaped JSON from websites | +| `context-scrape` | Markdown, HTML, images, screenshots, or structured JSON from one known URL | +| `context-crawl` | URL discovery and focused multi-page crawling | +| `context-extract` | Schema-shaped JSON from a known webpage | | `context-parse` | PDFs, Office files, images, and other local document bytes | | `context-brand` | Brand profiles, design systems, fonts, and industry codes | | `context-monitor` | Recurring website-change detection and history | @@ -60,13 +60,14 @@ Cursor automatically selects the appropriate Context.dev tool. Tool calls requir | Group | Tools | | --- | --- | | Parse | `parse-document` | -| Scrape and crawl | `web-scrape-html`, `web-scrape-markdown`, `web-scrape-images`, `web-scrape-sitemap`, `web-crawl`, `web-screenshot` | -| Extract and search | `web-extract`, `web-search`, `web-naics`, `web-sic` | -| Brand and design | `get-brand`, `brand-retrieve-unified`, `web-styleguide`, `web-fonts` | +| Web | `web-scrape`, `web-map`, `web-crawl`, `web-search`, `web-answers` | +| Company and people | `get-news-search`, `get-brand`, `brand-retrieve-unified`, `brand-search`, `people-enrich`, `web-styleguide` | | Batches | `submit-batch`, `list-batches`, `get-batch`, `get-batch-results`, `cancel-batch`, `delete-batch` | -| Monitors | `create-monitor`, `list-monitors`, `get-monitor`, `update-monitor`, `delete-monitor`, `run-monitor-now`, `list-monitor-runs`, `get-monitor-run`, `list-monitor-changes`, `list-account-runs`, `list-monitor-credit-usage`, `list-changes`, `get-change` | +| Monitors | `create-monitor`, `list-monitors`, `get-monitor`, `update-monitor`, `delete-monitor`, `get-monitor-limits`, `list-monitor-credit-usage`, `run-monitor-now`, `list-monitor-runs`, `get-monitor-run`, `list-account-runs`, `list-monitor-changes`, `get-change`, `list-changes`, `rotate-monitor-webhook-secret` | +| Webhooks | `list-webhook-deliveries`, `get-webhook-delivery`, `list-webhook-delivery-attempts`, `retry-webhook-delivery` | +| Account and feedback | `list-logs`, `get-log`, `submit-feedback` | -For a known page, use `web-scrape-markdown`. Use `web-search` when the URL is unknown, `web-crawl` for a focused multi-page request, and `submit-batch` for up to 25,000 URLs or a large asynchronous crawl. +For a known page, use `web-scrape` with `formats: { markdown: true }`. Use `web-search` when the URL is unknown, `web-crawl` for a focused multi-page request, and `submit-batch` for up to 25,000 URLs or a large asynchronous crawl. ## Building with the Context.dev API diff --git a/commands/extract-web-data.md b/commands/extract-web-data.md index 12a030a..5826ce1 100644 --- a/commands/extract-web-data.md +++ b/commands/extract-web-data.md @@ -8,5 +8,5 @@ description: Extract structured JSON from a website with Context.dev using a cle 1. Obtain the starting URL and the exact fields the user needs. 2. Build the smallest JSON Schema that represents those fields. Mark only genuinely required fields as required and describe ambiguous fields. 3. Confirm the `context` MCP server is enabled and authenticated. If authentication is required, use the `connect-context-dev` skill. -4. Call `web-extract` with the URL, schema, and concise instructions. Enable fact checking when every value must be stated on the source pages. +4. Call `web-scrape` with the URL, `formats: { json: true }`, and the schema in `jsonParams.schema`. 5. Return the structured result without silently filling missing fields. Explain unsupported or empty values plainly. diff --git a/commands/scrape-url.md b/commands/scrape-url.md index d8c8e12..5c090e9 100644 --- a/commands/scrape-url.md +++ b/commands/scrape-url.md @@ -7,6 +7,6 @@ description: Scrape a URL to markdown via Context.dev MCP for live page content 1. Obtain the full **URL** from the user (must include scheme, e.g. `https://example.com/docs`). 2. Confirm the `context` MCP server is enabled and authenticated. If authentication is required, use the `connect-context-dev` skill. -3. Call `web-scrape-markdown` with the URL. +3. Call `web-scrape` with the URL and `formats: { markdown: true }`. 4. Return the relevant Markdown or answer the user's question from it. Preserve source links when useful. 5. Do not fabricate page content. If the scrape fails, report the actual error and suggest a narrower selector, a longer timeout, or a retry only when appropriate. diff --git a/rules/prefer-context-dev-mcp.mdc b/rules/prefer-context-dev-mcp.mdc index 592bae7..149dd52 100644 --- a/rules/prefer-context-dev-mcp.mdc +++ b/rules/prefer-context-dev-mcp.mdc @@ -1,5 +1,5 @@ --- -description: Use the Context.dev MCP for live web search, scraping, extraction, parsing, brand data, monitoring, and batches +description: Use the Context.dev MCP for live web search, company news, scraping, extraction, parsing, brand data, monitoring, and batches alwaysApply: false --- @@ -8,13 +8,14 @@ Use the `context` MCP server instead of memory when the user needs current publi Choose the narrowest direct tool: - Unknown source or current topic: `web-search` -- One known page: `web-scrape-markdown`; raw DOM only: `web-scrape-html` -- Several linked pages: `web-crawl`; URL inventory: `web-scrape-sitemap` -- Typed fields from a site: `web-extract` +- News about one company by name, domain, ticker, or ISIN: `get-news-search` +- One known page: `web-scrape`; set only the needed `formats` fields +- Several linked pages: `web-crawl`; URL inventory: `web-map` +- Typed fields from a known page: `web-scrape` with `formats: { json: true }` and `jsonParams.schema` +- Sourced JSON research when the URL is unknown: `web-answers` - A local file: `parse-document` - Visual brand profile by domain: `get-brand`; raw or non-domain brand lookup: `brand-retrieve-unified` -- Design system, fonts, or rendered page: `web-styleguide`, `web-fonts`, or `web-screenshot` -- Industry classification: `web-naics` or `web-sic` +- Design system and fonts: `web-styleguide`; rendered page: `web-scrape` with `formats: { screenshot: true }` - Large asynchronous work: `submit-batch`, then `get-batch` and `get-batch-results` - Recurring change detection: the monitor tools diff --git a/scripts/validate-plugin.mjs b/scripts/validate-plugin.mjs index 2f8aa5a..f0a2e9b 100644 --- a/scripts/validate-plugin.mjs +++ b/scripts/validate-plugin.mjs @@ -18,6 +18,7 @@ const staleInstructionPatterns = [ ["retired code-mode search tool", /\bsearch_docs\b/], ["retired code-mode SDK execution", /client\.(?:brand\.retrieveSimplified|web\.webScrapeMd)/], ["retired MCP API-key header", /x-context-dev-api-key/i], + ["retired MCP tool name", /\b(?:web-scrape-html|web-scrape-markdown|web-scrape-images|web-scrape-sitemap|web-screenshot|web-extract|web-fonts|web-naics|web-sic)\b/], ]; function addError(message) { @@ -249,7 +250,9 @@ async function validateInstructionsAreCurrent() { "README.md", ...((await walkFiles(path.join(repoRoot, "commands"))).map((file) => path.relative(repoRoot, file))), ...((await walkFiles(path.join(repoRoot, "rules"))).map((file) => path.relative(repoRoot, file))), - "skills/connect-context-dev/SKILL.md", + ...((await walkFiles(path.join(repoRoot, "skills"))) + .filter((file) => path.basename(file) === "SKILL.md" && !file.endsWith(path.join("context-dev", "SKILL.md"))) + .map((file) => path.relative(repoRoot, file))), ]; for (const relativeFile of files) { @@ -263,9 +266,9 @@ async function validateInstructionsAreCurrent() { const commandToolRequirements = new Map([ ["commands/brand-colors.md", "get-brand"], - ["commands/scrape-url.md", "web-scrape-markdown"], + ["commands/scrape-url.md", "web-scrape"], ["commands/search-web.md", "web-search"], - ["commands/extract-web-data.md", "web-extract"], + ["commands/extract-web-data.md", "web-scrape"], ]); for (const [relativeFile, toolName] of commandToolRequirements) { const content = await fs.readFile(path.join(repoRoot, relativeFile), "utf8"); @@ -275,12 +278,12 @@ async function validateInstructionsAreCurrent() { } const skillToolRequirements = new Map([ - ["skills/context-search/SKILL.md", ["web-search"]], - ["skills/context-scrape/SKILL.md", ["web-scrape-markdown", "web-scrape-html", "web-scrape-images", "web-screenshot"]], - ["skills/context-crawl/SKILL.md", ["web-scrape-sitemap", "web-crawl"]], - ["skills/context-extract/SKILL.md", ["web-extract"]], + ["skills/context-search/SKILL.md", ["web-search", "get-news-search", "web-scrape"]], + ["skills/context-scrape/SKILL.md", ["web-scrape"]], + ["skills/context-crawl/SKILL.md", ["web-map", "web-crawl", "web-scrape"]], + ["skills/context-extract/SKILL.md", ["web-scrape", "web-map"]], ["skills/context-parse/SKILL.md", ["parse-document"]], - ["skills/context-brand/SKILL.md", ["get-brand", "brand-retrieve-unified", "web-styleguide", "web-fonts", "web-naics", "web-sic"]], + ["skills/context-brand/SKILL.md", ["get-brand", "brand-retrieve-unified", "brand-search", "web-styleguide", "people-enrich"]], ["skills/context-monitor/SKILL.md", ["create-monitor", "update-monitor", "delete-monitor", "run-monitor-now"]], ["skills/context-batches/SKILL.md", ["submit-batch", "get-batch", "get-batch-results", "cancel-batch", "delete-batch"]], ]); diff --git a/skills/connect-context-dev/SKILL.md b/skills/connect-context-dev/SKILL.md index 4b1f580..d6c69de 100644 --- a/skills/connect-context-dev/SKILL.md +++ b/skills/connect-context-dev/SKILL.md @@ -23,7 +23,7 @@ Use a read-only request: Use Context.dev to scrape https://www.context.dev and return the page title. ``` -The agent should call `web-scrape-markdown`. A successful tool call confirms both the MCP connection and the authenticated Context account. +The agent should call `web-scrape` with `formats: { markdown: true }`. A successful tool call confirms both the MCP connection and the authenticated Context account. ## Troubleshooting diff --git a/skills/context-brand/SKILL.md b/skills/context-brand/SKILL.md index 76542c1..562c955 100644 --- a/skills/context-brand/SKILL.md +++ b/skills/context-brand/SKILL.md @@ -11,10 +11,9 @@ Choose the Context.dev MCP tool by output and identifier: | --- | --- | | Visual brand profile for a domain | `get-brand` | | Raw structured brand data or a non-domain lookup | `brand-retrieve-unified` | +| Lightweight company-name or domain search | `brand-search` | | Website design system and component styling | `web-styleguide` | -| Website font inventory | `web-fonts` | -| NAICS classification | `web-naics` | -| SIC classification | `web-sic` | +| Person enrichment from identity clues | `people-enrich` | ## Workflow diff --git a/skills/context-crawl/SKILL.md b/skills/context-crawl/SKILL.md index bd91d31..3663f07 100644 --- a/skills/context-crawl/SKILL.md +++ b/skills/context-crawl/SKILL.md @@ -7,15 +7,15 @@ description: Discover or read multiple pages from a website with Context.dev. Us Choose between URL discovery and content collection: -- Use `web-scrape-sitemap` to discover and rank URLs without reading every page. +- Use `web-map` to discover and filter URLs without reading every page. - Use `web-crawl` to retrieve content from a bounded set of linked pages. ## Workflow 1. Confirm the target domain or starting URL and the section the user cares about. -2. Use sitemap search when the user wants particular pages rather than the whole site. +2. Use map search when the user wants particular pages rather than the whole site. 3. Apply path, subdomain, and link limits that match the request. 4. Keep synchronous crawls focused; do not expand scope beyond the requested site or section. 5. Return page URLs alongside the relevant content so results remain traceable. -Use `web-scrape-markdown` for one known page. For a large crawl or thousands of URLs, use `submit-batch` instead of forcing the work through a synchronous crawl. +Use `web-scrape` with `formats: { markdown: true }` for one known page. For a large crawl or thousands of URLs, use `submit-batch` instead of forcing the work through a synchronous crawl. diff --git a/skills/context-extract/SKILL.md b/skills/context-extract/SKILL.md index 9ae78be..a216197 100644 --- a/skills/context-extract/SKILL.md +++ b/skills/context-extract/SKILL.md @@ -5,16 +5,16 @@ description: Extract schema-shaped JSON from websites with Context.dev. Use when # Extract structured web data -Use the Context.dev `web-extract` MCP tool when the output must have a predictable JSON shape rather than free-form page text. +Use the Context.dev `web-scrape` MCP tool with `formats: { json: true }` and `jsonParams.schema` when a known webpage must return a predictable JSON shape rather than free-form text. ## Workflow 1. Define the smallest JSON Schema that contains only the fields the user needs. 2. Give every field a clear description that distinguishes similar values. 3. Make fields optional or nullable when the source may omit them. -4. Pass the relevant URL or URLs and describe the records to extract. +4. Pass the relevant URL and request only the JSON format. 5. Validate that the response follows the requested schema before presenting it. Preserve source semantics. Tied ranks such as `=19`, ranges such as `101-150`, and unavailable scores are legitimate source values; do not silently convert or invent them. -Extraction does not remove site pagination. Discover additional pages with `web-scrape-sitemap`, then use a batch when many page URLs must be processed. Use `web-scrape-markdown` when the user only needs readable text. +Extraction does not remove site pagination. Discover additional pages with `web-map`, then use a batch when many page URLs must be processed. Use `web-scrape` with `formats: { markdown: true }` when the user only needs readable text. diff --git a/skills/context-scrape/SKILL.md b/skills/context-scrape/SKILL.md index 9877374..748ac2d 100644 --- a/skills/context-scrape/SKILL.md +++ b/skills/context-scrape/SKILL.md @@ -5,14 +5,15 @@ description: Scrape or capture one known webpage with Context.dev. Use when the # Scrape one page -Choose the narrowest Context.dev MCP tool for the requested output: +Use the Context.dev `web-scrape` MCP tool and request only the formats the user needs: -| Need | Tool | +| Need | Format | | --- | --- | -| Clean readable content for analysis or RAG | `web-scrape-markdown` | -| Raw DOM or HTML for code-level inspection | `web-scrape-html` | -| Images and their source metadata | `web-scrape-images` | -| Visual rendering of the page | `web-screenshot` | +| Clean readable content for analysis or RAG | `formats: { markdown: true }` | +| Raw DOM or HTML for code-level inspection | `formats: { html: true }` | +| Images and their source metadata | `formats: { images: true }` | +| Visual rendering of the page | `formats: { screenshot: true }` | +| Schema-shaped JSON | `formats: { json: true }` with `jsonParams.schema` | ## Workflow diff --git a/skills/context-search/SKILL.md b/skills/context-search/SKILL.md index 280f29e..c756c81 100644 --- a/skills/context-search/SKILL.md +++ b/skills/context-search/SKILL.md @@ -1,20 +1,23 @@ --- name: context-search -description: Search the live web with Context.dev and return current, cited sources. Use when the user asks to search, research, look something up, find recent announcements or articles, compare information across sites, or answer a current question without providing a known URL. +description: Search the live web or company news with Context.dev and return current, cited sources. Use when the user asks to search, research, look something up, find recent company announcements or articles, compare information across sites, or answer a current question without providing a known URL. --- # Search the live web -Use the Context.dev `web-search` MCP tool when the user needs source discovery or current information and does not already have the exact page URL. +Use the Context.dev `get-news-search` MCP tool for live or historical news about one company identified by name, +domain, ticker, or ISIN. Use `web-search` for broader source discovery or current information when the user does not +already have the exact page URL. ## Workflow -1. Turn the request into a concise search query that preserves named entities, dates, and constraints. -2. Set freshness or country only when the request calls for it. -3. Use the smallest result count that can answer the question. -4. Prefer authoritative or primary sources when the user asks for official information. -5. Summarize the relevant findings and cite the returned source URLs. +1. Choose `get-news-search` when the request is specifically about one company's coverage; otherwise use `web-search`. +2. Preserve named entities, dates, and constraints in the query or company identifier. +3. Set filters, freshness, or country only when the request calls for them. +4. Use the smallest result count that can answer the question. +5. Prefer authoritative or primary sources when the user asks for official information. +6. Summarize the relevant findings and cite the returned source URLs. -If the user provides a known URL, use `web-scrape-markdown` instead. If a search result must be read in full, scrape only the selected result rather than every result. +If the user provides a known URL, use `web-scrape` with `formats: { markdown: true }` instead. If a search result must be read in full, scrape only the selected result rather than every result. Avoid repeating an identical search unless the first call failed or the query materially changed. From 201bbb66da467fd4303115d35fc8d27576e5daa0 Mon Sep 17 00:00:00 2001 From: aadithyan rajesh Date: Mon, 28 Sep 2026 01:09:27 +0530 Subject: [PATCH 2/2] docs(skill): sync Context API guide --- skills/context-dev/SKILL.md | 563 +++++++++++------------------------- 1 file changed, 166 insertions(+), 397 deletions(-) diff --git a/skills/context-dev/SKILL.md b/skills/context-dev/SKILL.md index f430133..c76c230 100644 --- a/skills/context-dev/SKILL.md +++ b/skills/context-dev/SKILL.md @@ -4,432 +4,201 @@ description: Build application code directly against the Context.dev REST API or license: MIT metadata: author: context.dev - version: "3.1" + version: "5.2" + last_verified: "2026-09-26" --- -# Context.dev +# Context.dev integration guide -Context.dev turns any domain or URL into structured, typed JSON: brand profiles, design systems, scraped content, extracted products, industry codes, and merchant identity. Base URL `https://api.context.dev/v1`. Full docs: [docs.context.dev](https://docs.context.dev) · machine index: [docs.context.dev/llms.txt](https://docs.context.dev/llms.txt). +Context.dev provides web scraping, URL mapping, crawling, web search and research, document parsing, and brand, people, and news data through `https://api.context.dev/v1`. One Scrape request can return any mix of nine outputs: Markdown, rendered HTML, a screenshot, images, original bytes, CSS-selected fields, highlights, JSON, and product data. Use the public OpenAPI document at [docs.context.dev/openapi.json](https://docs.context.dev/openapi.json) as the authority for paths, methods, parameters, and response fields. -Every endpoint takes one bearer token and returns typed JSON. No HTML parsing, no Open Graph scraping. This file tells you which endpoint answers which question, exactly what each takes, and exactly what each gives back. +## Get access -## Setup +- To call Context.dev from this agent session without writing code, connect the MCP server `https://mcp.context.dev/mcp`. It signs in with OAuth in the user's browser and does not use an API key. Setup for each client: https://docs.context.dev/install-mcp.md +- Application code and the CLI read an API key from `CONTEXT_DEV_API_KEY`. If none is configured, follow https://www.context.dev/auth.md: register with the user's email, deliver the returned setup link and code to the user before polling, and store the key in an ignored environment file or secret manager without displaying it. Never ask the user to paste a key into chat. +- Complete setup instructions: https://docs.context.dev/agent-quickstart.md -Authenticate with a bearer token read from `CONTEXT_DEV_API_KEY`. Never hardcode it; never ship it to client-side code (use a backend proxy). +## Keep credentials server-side + +Read the bearer token from `CONTEXT_DEV_API_KEY`. Never print it, commit it, include it in browser code, or forward it to a target website. ```bash export CONTEXT_DEV_API_KEY="ctxt_secret_..." -``` - -Install an SDK, or call REST directly with `curl`: - -| Language | Install | Import | -| ---------- | -------------------------------------------------- | --------------------------------------------------------------- | -| TypeScript | `npm install context.dev` | `import ContextDev from "context.dev"` | -| Python | `pip install context.dev` | `from context.dev import ContextDev` | -| Ruby | `gem install context.dev` | `require "context_dev"` | -| Go | `go get github.com/context-dot-dev/context-go-sdk` | `import contextdev "github.com/context-dot-dev/context-go-sdk"` | -| PHP | `composer require context-dev/context-dev-php` | `use ContextDev\Client;` | -```typescript -import ContextDev from "context.dev"; -const client = new ContextDev({ apiKey: process.env.CONTEXT_DEV_API_KEY }); -const { brand } = await client.brand.retrieve({ type: "by_domain", domain: "stripe.com" }); +curl https://api.context.dev/v1/web/scrape \ + -H "Authorization: Bearer $CONTEXT_DEV_API_KEY" \ + -H "Content-Type: application/json" \ + -d '{ + "url": "https://example.com", + "formats": { "markdown": true }, + "sharedParams": { "mainContentOnly": true } + }' ``` -**SDK naming.** Methods below are shown in TypeScript camelCase (`client.brand.retrieve`). Python and Ruby use snake_case (`retrieve`); PHP uses the same camelCase method names as TypeScript with named parameters (`$client->brand->retrieve(type: 'by_domain', domain: 'stripe.com')`); Go uses PascalCase and renames a few (`client.Brand.Get`, `client.Industry.GetNaics`). Methods are grouped under five namespaces that don't always match the URL path: `brand.*`, `web.*`, `ai.*`, `industry.*` (NAICS/SIC, despite `/web/` paths), `utility.*` (prefetch). The Python SDK currently lags the others: a few methods live under different namespaces (`client.style.*` for styleguide/fonts) or are missing; if an SDK method is missing or unavailable, call the REST path directly. See [best practices](https://docs.context.dev/optimization/best-practices). - -## Choosing an endpoint - -Pick the narrowest endpoint that answers the question. Start from what you already have: - -| You have | Use | Path | -| --------------------------------- | ------------------------------------------------ | ----------------------------------------------------------------- | -| A domain, want everything | Retrieve Brand | `POST /brand/retrieve` with `type: "by_domain"` | -| A domain, only need logo + colors | Retrieve Simplified (same price, smaller/faster) | `GET /brand/retrieve-simplified` | -| A company name | Retrieve by Name | `POST /brand/retrieve` with `type: "by_name"` | -| A work email | Retrieve by Email | `POST /brand/retrieve` with `type: "by_email"` | -| A stock ticker | Retrieve by Ticker | `POST /brand/retrieve` with `type: "by_ticker"` | -| An ISIN | Retrieve by ISIN | `GET /brand/retrieve-by-isin` | -| A card/bank descriptor | Brand transaction lookup | `POST /brand/retrieve` with `type: "by_transaction"` | -| A person's email, name, or social profile URL | Enrich Person (beta) | `POST /people/enrich` | -| A specific page URL you want scraped for brand fields | Retrieve by direct URL | `POST /brand/retrieve` with `type: "by_direct_url"` | -| A URL → clean text for an LLM | Scrape Markdown | `GET /web/scrape/markdown` | -| A whole site → text for RAG | Crawl | `POST /web/crawl` | -| A design system to copy/theme | Styleguide | `GET /web/styleguide` | -| A product page → structured data | Extract Product | `POST /brand/ai/product` | -| A custom schema for a site | Structured Extract | `POST /web/extract` | -| An industry code (NAICS/SIC) | Classify | `GET /web/naics` · `/web/sic` | +## Route from input to operation + +Choose the narrowest operation that directly returns the needed result. + +| Input and desired result | Operation | Task guide | +| --- | --- | --- | +| One URL to Markdown or rendered HTML | `POST /web/scrape` with `formats.markdown` or `formats.html` | [Scrape](https://docs.context.dev/scrape/overview) | +| Known CSS selectors on one page to JSON | `POST /web/scrape` with `formats.parse` and `parseParams.rules` | [Parse fields](https://docs.context.dev/scrape/parse-fields) | +| One URL and a question to the most relevant passages | `POST /web/scrape` with `formats.highlights` and `highlightsParams.query` | [Highlights](https://docs.context.dev/scrape/highlights) | +| One URL and a JSON Schema to a structured object | `POST /web/scrape` with `formats.json` and `jsonParams.schema` | [JSON](https://docs.context.dev/scrape/json) | +| Exact URL to an inline PNG, JPEG, or WebP screenshot | `POST /web/scrape` with `formats.screenshot` | [Screenshot](https://docs.context.dev/scrape/screenshot) | +| URL to image assets | `POST /web/scrape` with `formats.images` | [Images](https://docs.context.dev/scrape/images) | +| Resource URL to complete base64 bytes | `POST /web/scrape` with `formats.bytes` | [Bytes](https://docs.context.dev/scrape/bytes) | +| Product page to a product record with price, availability, images, and variants | `POST /web/scrape` with `formats.product` | [Product](https://docs.context.dev/scrape/product) | +| YouTube video URL to metadata and a timestamped transcript | `POST /web/scrape` with `formats.markdown` | [YouTube](https://docs.context.dev/scrape/youtube) | +| Page that needs clicks, scrolls, or waits before capture | `POST /web/scrape` with `sharedParams.actions` | [Browser actions](https://docs.context.dev/scrape/browser-actions) | +| Domain to its indexed URLs with page titles and descriptions | `GET /web/urls` | [Map URLs](https://docs.context.dev/map/overview) | +| Starting URL to linked pages in one response, up to 500 pages | `POST /web/crawl` | [Crawl](https://docs.context.dev/crawl/overview) | +| Starting URL or sitemap to a background crawl, up to 25,000 pages | `POST /batch/submit` with `input.mode: "crawl"` | [Async crawls](https://docs.context.dev/crawl/async) | +| URL list to Markdown or HTML in the background, up to 25,000 URLs | `POST /batch/submit` | [Batches](https://docs.context.dev/batches/overview) | +| Search query to ranked web results, optionally with Markdown | `POST /web/search` | [Search](https://docs.context.dev/search/overview) | +| Research task to structured JSON and source URLs | `POST /web/answers` | [Answers](https://docs.context.dev/answers/overview) | +| Uploaded document up to 50 MiB to Markdown | `POST /parse` | [Parse](https://docs.context.dev/parse/overview) | +| Domain, name, work email, ticker, direct URL, or transaction descriptor to company profile | `POST /brand/retrieve` | [Brand](https://docs.context.dev/brand/overview) | +| Partial name or domain to matching indexed brands | `GET /brand/search` | [Brand search](https://docs.context.dev/brand/search) | +| Domain or exact URL to design styles, font families, and available font files | `GET /web/styleguide`; read `styleguide.typography` and `styleguide.fontLinks` for fonts | [Styleguide](https://docs.context.dev/brand/styleguide) | +| Domain or work email known before a Brand or Styleguide request | `POST /utility/prefetch` | [Prefetch](https://docs.context.dev/brand/prefetching) | +| Person email, profile URL, or name plus company, school, or location to a person profile | `POST /people/enrich` (beta, paid plans) | [People](https://docs.context.dev/people/overview) | +| Company name, domain, ticker, or ISIN to news articles | `POST /news/search` | [News](https://docs.context.dev/news/overview) | +| URL or site to recurring change detection | `/monitors` operations | [Monitors](https://docs.context.dev/monitors/overview) | +| Context.dev bug, docs mismatch, or friction you hit while integrating | `POST /feedback` | [Agent feedback](https://docs.context.dev/optimization/agent-feedback) | + +Do not use a general scrape when a purpose-built Brand or monitor operation already returns the required shape. + +## Verify the installed SDK before using it + +| Language | Package | Client setup | Scrape and Map URLs | +| --- | --- | --- | --- | +| TypeScript | `context.dev` | `new ContextDev({ apiKey: process.env.CONTEXT_DEV_API_KEY })` | `client.web.scrape`, `client.web.mapUrls` | +| Python | `context.dev` | `ContextDev(api_key=os.environ["CONTEXT_DEV_API_KEY"])` | `client.web.scrape`, `client.web.map_urls` | +| Ruby | `context.dev` | `ContextDev::Client.new(api_key: ENV.fetch("CONTEXT_DEV_API_KEY"))` | `client.web.scrape`, `client.web.map_urls` | +| Go | `github.com/context-dot-dev/context-go-sdk/v2` | Use the `/v2` module; inspect its generated request model. | `client.Web.Scrape`, `client.Web.MapURLs` | +| PHP | `context-dev/context-dev-php` | Inspect the installed package version and generated method signature. | `$client->web->scrape`, `$client->web->mapUrls` | + +- Go needs the `/v2` module path for `Web.Scrape`, `Web.MapURLs`, and the current Brand request body. +- The PHP Brand helper can't express a single lookup type; use the SDK's low-level request method for Brand calls. +- Published packages can lag the API. If an installed scrape method doesn't accept `highlightsParams`, `jsonParams`, or `productParams`, send the request with the SDK's low-level request method. Do not invent generated types or pass fields the installed signature does not accept. +- SDKs time out after 60 seconds per attempt (Go has no default timeout) and retry twice on connection errors, `408`, `409`, `429`, and `5xx`. Raise the SDK timeout above `timeoutOpts.milliseconds` for long requests. + +For installation, runnable examples, and low-level request methods, read [SDKs](https://docs.context.dev/sdks). For a first request and a trimmed response, read the [Quickstart](https://docs.context.dev/quickstart). + +## Scrape controls + +- `formats`: enable at least one of `html`, `markdown`, `screenshot`, `images`, `bytes`, `parse`, `highlights`, `json`, or `product`. Outputs share one page visit. Each output has `requested`, `success`, and `data`; `success` is `true` when retrieved, `false` when retrieval fails, and `null` when not requested. Failed outputs have `data: null` and do not discard successful outputs. Cost depends on the formats; see the cost table below. +- `maxAgeMs`: chooses acceptable cache age. Scrape and both scrape/crawl batches default to three days (`259200000` ms) and accept up to 365 days (`31536000000` ms). Synchronous Crawl defaults to one day and accepts up to 30 days. Set `0` to fetch fresh and refresh the requested outputs. Each Scrape output has its own cache key, so one response can combine cached outputs from different visits. Brand and Styleguide default to three months, accept `0` for a hard refresh, and clamp values above one year. The direct-URL and transaction Brand variants do not accept this control. Custom headers, actions, and `zdr` bypass cache reads and writes; otherwise inspect `cache_metadata.status` (`hit`, `miss`, or `zdr`) and `age_ms`. +- `timeoutOpts`: set `milliseconds` up to 300,000 and choose `behavior: "fail"` or `"return-partial"` where supported. Reaching the overall deadline with `"fail"` returns `408`. Scrape requires at least 5,000 ms for `return-partial` and marks partial captures or failed outputs with `isPartial: true`; individual outputs can fail under either behavior while others succeed. Crawl and Map URLs return `partial: true`. Prefetch is fail-only and Parse does not accept `timeoutOpts`. See [timeouts](https://docs.context.dev/optimization/timeouts). +- `maxPages`, `maxDepth`, `urlRegex`, and `stopAfterMs`: bound crawl coverage, time, and credit exposure. Crawl caps `maxPages` at 500. +- `sharedParams.mainContentOnly`, `includeSelectors`, `excludeSelectors`, and `includeFrames`: control page content before conversion. Exclusions win. Filters apply to HTML, Markdown, images, and parsed fields, never to `bytes`. +- `formats.parse` with `parseParams.rules`: map each field name to a CSS selector string or `{selector, type: "item" | "list", output: "text" | "html" | "@attribute" | nested rules}`. Results arrive in `parsed.data`; missing items return `null` and missing lists return `[]`. Runs without an LLM. +- `sharedParams.waitFor` (milliseconds or a CSS selector; default 500 ms), `settleAnimations`, and browser actions (`type: "perform" | "scroll" | "wait" | "waitFor"`, 1 to 5 per request): handle dynamic page state. Actions require a paid plan and bypass the cache. Action or selector failures can leave affected outputs with `success: false`; verify the page state from successful captured outputs. +- `formats.bytes`: returns the original HTTP body as `{contentType, base64}` after decompression, up to 20 MiB decoded. Waiting, actions, and content filters never change it; request `formats.html` for rendered HTML. +- `sharedParams.headers` go to the target origin only and are separate from the API bearer key. `sharedParams.country` is a two-letter code for the network exit of the page fetch and image downloads; an unsupported code returns `400` before the scrape starts. See [location and headers](https://docs.context.dev/scrape/location-and-headers). +- `screenshot.data` is an inline data URL (`data:image/png;base64,...`) when `screenshot.success` is `true`; otherwise it is `null`. Choose the capture area with `screenshotParams.area` (`viewport`, `fullPage`, `{selector}`, or a rectangle) up to 40 megapixels. +- `zdr: "enabled"`: bypasses shared caches and retained content logs only when Zero Data Retention is enabled for the organization; otherwise the request fails with `403`. + +Browser rendering and proxy routing improve access but do not guarantee that a page can be read. Preserve a fallback for `WEBSITE_BLOCKED`, login walls, missing content, and partial crawl results. + +## Map a website's URLs + +`GET /web/urls` lists a site's URLs from Context.dev's index and adds stored `title`, `description`, `keywords`, and `language` where available. Pass `domain` without a protocol; narrow with `urlRegex`, `maxLinks` (default 10,000), `includeSubdomains`, or `search` (relevance-ordered). The SDK methods are `client.web.mapUrls` (TypeScript) and `client.web.map_urls` (Python, Ruby). + +URLs without stored metadata return with only `url` and are queued for background enrichment, so a later request can include their metadata. Requests with `zdr=enabled` or credential-bearing target headers return URLs only and do not queue enrichment. `partial: true` means the deadline stopped URL mapping. Feed the result into Scrape, Crawl, or a batch. + +## Retrieve brand data + +Send exactly one lookup variant: + +| `type` | Required field | Important constraint | +| --- | --- | --- | +| `by_domain` | `domain` | Prefer a bare company domain. | +| `by_name` | `name` | 3 to 30 characters; use `country_gl` as an ambiguity hint. | +| `by_email` | `email` | Free and disposable providers return `422`. | +| `by_ticker` | `ticker` | Add `ticker_exchange` when known. | +| `by_direct_url` | `direct_url` | Reads only that page; no wider resolver or cross-source enrichment. | +| `by_transaction` | `transaction_info` | Add MCC, city, country, or phone hints when available. | -**Prefer bare domains:** use `stripe.com` instead of `https://stripe.com` or `www.stripe.com`. The API normalizes protocol and `www.`, but a string with no TLD fails validation. - ---- - -## Brand intelligence - -Brand lookups share the same response envelope: `{ status, code, brand }`. They cost **10 credits** each and only bill on a successful resolution (a 400 `NOT_FOUND` "no brand" response is free). Guide: [Get brand data](https://docs.context.dev/guides/get-brand-data). - -### The `brand` object (shared response shape) - -This is what `brand` contains on full Brand responses from `POST /brand/retrieve`, including domain, name, email, ticker, ISIN, and transaction lookups. Any field may be `null`/absent, so always provide fallbacks. - -| Field | Type | Notes | -| ----------------------------- | ------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `brand.domain` | string | Canonical domain after normalization. | -| `brand.title` | string | Company name. | -| `brand.description` | string | One-paragraph description. | -| `brand.slogan` | string | Tagline. | -| `brand.colors[]` | array | `{ hex, name }`, ordered by prominence. `name` is generated, not official; key off `hex`. | -| `brand.logos[]` | array | `{ url, mode, type, resolution{width,height,aspect_ratio}, colors[] }`. `mode` ∈ `light` / `dark` / `has_opaque_background`; `type` ∈ `icon` (square) / `logo` (horizontal). **Filter by `mode`+`type`; don't assume `logos[0]`.** | -| `brand.backdrops[]` | array | Hero imagery: `{ url, colors[], resolution }`. | -| `brand.socials[]` | array | `{ type, url }`. `type` ∈ x, facebook, instagram, linkedin, youtube, tiktok, github, +24 more. | -| `brand.address` | object | `street, city, state_province, state_code, country, country_code, postal_code`. | -| `brand.stock` | object\|null | `{ ticker, exchange }`. `null` for private companies. | -| `brand.employees` | object\|null | `{ range, exact }`. `range` is a bucket from `1 to 10` through `10001+`; `exact` is the precise headcount when known. `null` when unknown. | -| `brand.industries.eic[]` | array | `{ industry, subindustry }`: Context's own taxonomy ([EIC](https://docs.context.dev/guides/classification/EIC)), inline on every full response. For NAICS/SIC use the dedicated endpoints. | -| `brand.links` | object | `careers, blog, pricing, contact, terms, privacy`, each nullable. | -| `brand.email` / `brand.phone` | string | Public contact info, when found. | -| `brand.primary_language` | string\|null | Detected site language. | -| `brand.is_nsfw` | boolean | Safe-content flag. | - -```json -{ - "status": "ok", - "code": 200, - "brand": { - "domain": "stripe.com", - "title": "Stripe", - "colors": [{ "hex": "#543cfb", "name": "Meteor Shower" }], - "logos": [ - { - "url": "https://media.brand.dev/….svg", - "mode": "light", - "type": "logo", - "resolution": { "width": 150, "height": 48, "aspect_ratio": 3.13 } - } - ], - "industries": { - "eic": [ - { "industry": "Finance", "subindustry": "Payments & Money Movement" } - ] - }, - "stock": null - } -} +```typescript +const response = await client.brand.retrieve({ type: "by_domain", domain: "stripe.com" }); +const title = response.brand?.title ?? "Unknown company"; ``` -### Retrieve brand by domain - -`POST /brand/retrieve` with `type: "by_domain"` · 10 credits · SDK `client.brand.retrieve` (Go `Brand.Get`) -[Guide](https://docs.context.dev/guides/get-brand-data#get-a-brand-by-domain) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/retrieve-brand-data-by-domain) - -- **When:** you have the company's website domain and want the full profile (logos, colors, socials, address, industry, stock). -- **Takes (JSON body):** `type: "by_domain"` (**required** discriminator), `domain` (string, **required**, bare domain). Optional: `maxSpeed` (bool; skip slow steps for a faster, lighter answer), `force_language` ([SupportedLanguage](https://docs.context.dev/guides/get-brand-data) enum), `maxAgeMs` (int, default `7776000000` ≈ 90d, clamped 1d–1y), `timeoutMS` (int, max `300000`). -- **Gives:** the shared `{ status, code, brand }` envelope above. - -### Retrieve simplified - -`GET /brand/retrieve-simplified` · 10 credits · SDK `client.brand.retrieveSimplified` (Go `Brand.GetSimplified`) -[Guide](https://docs.context.dev/guides/get-brand-data#get-a-brand-by-domain) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/retrieve-simplified-brand-data-by-domain) - -- **When:** you have a domain and only need lightweight visual assets. Fastest payload for logo walls, signup pre-fill, or theming. -- **Takes:** `domain` (string, **required**), `theme` (`light`|`dark`; selects matching assets), `maxAgeMs`, `timeoutMS`. No name/email/ticker/`maxSpeed`/`force_language`. -- **Gives:** `{ status, code, brand }` where `brand` is **stripped to `domain`, `title`, `colors[]`, `logos[]`, `backdrops[]` only**: no description, socials, address, stock, industries, or links. Same 10-credit price as the full retrieve, just less data. - -### Retrieve by company name - -`POST /brand/retrieve` with `type: "by_name"` · 10 credits · SDK `client.brand.retrieve` (Go `Brand.Get`) -[Guide](https://docs.context.dev/guides/get-brand-data#look-up-by-company-name) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/brand) - -- **When:** you only know the company name and need Context to resolve it to a domain + full profile. -- **Takes (JSON body):** `type: "by_name"` (**required**), `name` (string, **required**, 3–30 chars). Optional: `country_gl` (ISO 3166-1 alpha-2 hint to disambiguate, e.g. `us`), `maxSpeed`, `force_language`, `maxAgeMs`, `timeoutMS`. -- **Gives:** the full shared `brand` envelope. - -### Retrieve by work email - -`POST /brand/retrieve` with `type: "by_email"` · 10 credits · SDK `client.brand.retrieve` (Go `Brand.Get`) -[Guide](https://docs.context.dev/guides/get-brand-data#look-up-by-work-email) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/brand) - -- **When:** lead/onboarding enrichment from a work email; the domain is extracted automatically. -- **Takes (JSON body):** `type: "by_email"` (**required**), `email` (string, **required**). Optional: `maxSpeed`, `force_language`, `maxAgeMs`, `timeoutMS`. -- **Gives:** the full shared `brand` envelope. -- **Note:** free providers (gmail, outlook…) and disposable addresses return **HTTP 422** (`FREE_EMAIL_DETECTED` / `DISPOSABLE_EMAIL_DETECTED`); handle 422 as "skip enrichment", not a hard error. - -### Retrieve by ticker - -`POST /brand/retrieve` with `type: "by_ticker"` · 10 credits · SDK `client.brand.retrieve` (Go `Brand.Get`) -[Guide](https://docs.context.dev/guides/get-brand-data#look-up-by-stock-ticker) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/brand) - -- **When:** investor/finance flows keyed on a listed ticker. -- **Takes (JSON body):** `type: "by_ticker"` (**required**), `ticker` (string, **required**, 1–15 chars, e.g. `AAPL`, `BRK.A`). Optional `ticker_exchange` (**defaults to NASDAQ**; set it for non-NASDAQ listings), plus `maxSpeed`/`force_language`/`maxAgeMs`/`timeoutMS`. -- **Gives:** the full shared `brand` envelope, with `brand.stock` populated. +An uncached domain or email lookup with `behavior: "fail"` and `timeoutOpts.milliseconds` under 10,000 returns `422 COLD_DOMAIN_TIMEOUT_TOO_LOW`; allow 60,000 ms or prefetch. Brand arrays and nested fields can be missing or empty. Select logos by `type`, `mode`, and resolution; do not assume `logos[0]` is appropriate. Treat address, contacts, employee count, classifications, and stock data as discovered data rather than verified legal records. -### Retrieve by ISIN +## Permissions and other operation notes -`GET /brand/retrieve-by-isin` · 10 credits · SDK `client.brand.retrieveByIsin` (Go `Brand.GetByIsin`) -[Guide](https://docs.context.dev/guides/get-brand-data#look-up-by-isin) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/retrieve-brand-data-by-isin) +- Restricted API keys need the relevant [scope](https://docs.context.dev/account/api-keys): `data:execute` for direct data calls, Read/Manage scopes for monitors and batches, and `logs:read` for request logs. Any key can submit Agent Feedback. Dashboard [team roles](https://docs.context.dev/account/team) are independent; Members can use shared credentials and incur charges. +- Answers defaults to `ultra` mode; `fast` is cheaper and quicker. `json_format` is an example object, and source URLs are not per-field citations. +- Scrape, Map URLs, Crawl, Search, Answers, Parse, People Enrich, and Styleguide honor `zdr` when the organization is entitled. +- Page monitors accept `include_selectors`/`exclude_selectors`; changing them establishes a new baseline. -- **When:** you're joining against financial feeds keyed on ISIN. Ticker lookups are the common path; ISIN is the niche/hidden one. -- **Takes:** `isin` (string, **required**, exactly 12 chars `^[A-Z]{2}[A-Z0-9]{9}[0-9]$`, e.g. `US0378331005`), plus `maxSpeed`/`force_language`/`maxAgeMs`/`timeoutMS`. -- **Gives:** the full shared `brand` envelope, with `brand.stock` populated. +## Budget requests -### Transaction enrichment +| Operation | Credits | +| --- | --- | +| `POST /web/scrape` | 1, or 2 with `sharedParams.actions`. Highlights +3 when passages are returned; JSON +4 on success; product +1 on success or a missing page; product AI fallback +6 when used; PDF OCR +1 per recovered page on a fresh fetch. 0 when every output fails, except a target 404, which charges the base (+1 with product). | +| `GET /web/urls` | 1, or 2 with `search` | +| `POST /web/crawl` | 1 per page; rate-limit weight 10 | +| `POST /web/search` | 1 per 10 results, with or without page content | +| `POST /web/answers` | 10 (`fast`) or 100 (`ultra`, the default), charged only on success | +| `POST /parse` | 1, plus 1 per OCR-recovered page | +| `POST /brand/retrieve`, `GET /web/styleguide` | 10 | +| `GET /brand/search` | 1; free on Pro, Growth, and Scale | +| `POST /people/enrich` | 20 per match; paid plans | +| `POST /news/search` | 1 per 10 results | +| `POST /utility/prefetch` | 0; paid plans | +| `POST /batch/submit` | 1 per page: 1 reserved per accepted URL (a crawl reserves `maxUrls`), refunded for pages that don't succeed, plus 1 per OCR-recovered PDF page | +| `POST /monitors`, `POST /monitors/{id}/run` | Per run: 1 for a page or sitemap monitor, 10 for an extract monitor. Baseline runs, at creation or after a re-baseline, are charged; failed and skipped runs are free. | +| Other batch and monitor operations, webhook deliveries, request logs, `POST /feedback` | 0 | -`POST /brand/retrieve` with `type: "by_transaction"` · 10 credits · SDK `client.brand.retrieve` -[Guide](https://docs.context.dev/guides/enrich-transaction-codes) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/brand) +Processed `404` results on priced data APIs and successful partial results are charged. Validation, authentication, rate-limit, timeout, and server errors are not. Record `key_metadata.credits_consumed` or the `X-Credits-Used` header instead of inferring cost from status. Plan allowances are on [Credits and pricing](https://docs.context.dev/account/credits). -- **When:** you have a messy card/ACH descriptor (`AMZN MKTP US`, `SQ *COFFEE BAR`) and need the real merchant brand for spend analytics or categorization. -- **Takes (JSON body):** `type: "by_transaction"` (**required**), `transaction_info` (string, **required**). Optional disambiguators sharply improve accuracy: `mcc` (4-digit category code), `city`, `country_gl` (ISO alpha-2), `phone` (number), `high_confidence_only` (bool, default false; set true for fewer false matches), `maxSpeed`, `force_language`, `timeoutMS`. -- **Gives:** the full shared `brand` envelope (the identified merchant). -- **Note:** this is the only brand lookup that does **not** accept `maxAgeMs`. - -### Retrieve by direct URL - -`POST /brand/retrieve` with `type: "by_direct_url"` · 10 credits · SDK `client.brand.retrieve` (Go `Brand.Get`) -[Guide](https://docs.context.dev/guides/get-brand-data#look-up-by-direct-url) · [API reference](https://docs.context.dev/api-reference/brand-intelligence/brand) - -- **When:** you want brand fields scraped from **one specific page** — a subpath, landing page, preview URL, or a domain the database doesn't cover yet. Fetches only that URL; no domain resolution, DB lookup, or cross-source enrichment. -- **Takes (JSON body):** `type: "by_direct_url"` (**required**), `direct_url` (string, **required**, full http(s) URL). Optional: `timeoutMS`. Rejects `maxSpeed`, `force_language`, and `maxAgeMs` — those would be silently ignored by the single-page scrape flow. -- **Gives:** the shared `brand` envelope, but only fields extractable from the page: `domain`, `title`, `description`, `logos` (URLs only), `socials`, `email`, `phone`, `links`. Enrichment-only fields (`colors`, `backdrops`, `industries`, `stock`, `address`) are omitted. -- **Note:** unreachable URLs return `400` `WEBSITE_ACCESS_ERROR` (or `WEBSITE_NOT_FOUND` when the domain doesn't resolve at all). - ---- - -## Web scraping - -Render, crawl, and search the live web. Bot-detection bypass and proxy escalation are automatic. Guide: [Scrape websites](https://docs.context.dev/guides/scrape-websites-to-markdown). - -### Scrape Markdown - -`GET /web/scrape/markdown` · 1 credit · SDK `client.web.webScrapeMd` (Go `Web.WebScrapeMd`) -[Guide](https://docs.context.dev/guides/scrape-websites-to-markdown#scrape-a-single-page-to-markdown) · [API reference](https://docs.context.dev/api-reference/web-scraping/scrape-markdown) - -- **When:** turn one page into clean, LLM-ready GitHub-Flavored Markdown (nav/ads stripped). The default for feeding pages to a model. -- **Takes:** `url` (string URI, **required**). Optional: `includeLinks` (default true), `includeImages` (default false), `shortenBase64Images` (default true), `useMainContentOnly` (default false), `includeHTML` (default false; adds the source HTML beside the Markdown), `includeFrames` (default false), `includeSelectors[]`/`excludeSelectors[]` (CSS selectors to keep/remove before conversion; exclusion wins), `pdf` object (`shouldParse` default true, `start`/`end` page range, `ocr` default false — OCRs scanned pages at **1 credit per recovered page**; a scan with `ocr` off returns `400 PDF_IMAGES_ONLY`), `maxAgeMs` (default 1d, 0–30d), `waitForMs` (0–30000 render wait), `settleAnimations`, `country`, `actions[]` (up to 5 ordered `{do: "wait", timeMs}` / `{do: "perform", action}` steps; **paid plan only, 2 credits, bypasses cache**), `headers` object (forwarded to the target URL; bypasses cache), `timeoutMS`. -- **Gives:** `{ success: true, markdown, html?, contentLength, url, metadata }`. `metadata.headings[]` contains `{ level, text }` entries in document order when headings are present. - -### Scrape HTML - -`GET /web/scrape/html` · **1 credit (2 with `actions`)** · SDK `client.web.webScrapeHTML` (Go `Web.WebScrapeHTML`) -[Guide](https://docs.context.dev/guides/scrape-websites-to-markdown#scrape-a-single-page-to-markdown) · [API reference](https://docs.context.dev/api-reference/web-scraping/scrape-html) - -- **When:** you need the fully-rendered raw HTML (to parse the DOM, attributes, or scripts) instead of cleaned text. -- **Takes:** `url` (**required**), `pdf`, `includeFrames`, `useMainContentOnly`, `includeSelectors[]`/`excludeSelectors[]`, `maxAgeMs`, `waitForMs`, `settleAnimations`, `country`, `actions[]` (paid plan only, 2 credits, bypasses cache; same shape as Scrape Markdown), `headers`, `timeoutMS` (same shapes as Scrape Markdown). -- **Gives:** `{ success: true, html, url, type, metadata }`. `metadata.headings[]` is present when headings were found. - -### Scrape Markdown / HTML with actions - -Attach `actions` to `/web/scrape/markdown` or `/web/scrape/html` when the page only reveals its content after interaction (cookie banners, "Load more" buttons, tab switches). Up to **5 ordered actions** run after page load and before capture. Each is a discriminated object: `{do: "wait", timeMs: <0–30000>}` or `{do: "perform", action: ""}`. **Paid plan required** — free-tier keys get `403 PAID_PLAN_REQUIRED`. Any action bumps the call to **2 credits** (5 for enriched images) and bypasses the scrape cache and HTML/PDF fast paths. - -### Scrape Images - -`GET /web/scrape/images` · **1 credit · 2 with `actions` · 5 if any enrichment flag is set** · SDK `client.web.webScrapeImages` (Go `Web.WebScrapeImages`) -[Guide](https://docs.context.dev/guides/scrape-websites-to-markdown#extract-every-image-on-a-page) · [API reference](https://docs.context.dev/api-reference/web-scraping/scrape-images) - -- **When:** enumerate every image on a page (`img`, inline SVG, CSS backgrounds, video posters, data URIs) and optionally measure / classify / CDN-host them. -- **Takes:** `url` (**required**), `maxAgeMs`, `dedupe` (perceptually remove near-duplicates), `waitForMs`, `actions[]` (paid plan only, 2 credits, bypasses cache; same shape as Scrape Markdown — enriched calls stay at 5 credits), `headers` (forwarded to the target URL; bypasses cache), `timeoutMS`, and an `enrichment` object: `resolution` (bool), `hostedUrl` (bool), `classification` (bool), `maxTimePerMs` (int). **Enabling any enrichment flag makes the whole call cost 5 credits** (not per-image). -- **Gives:** `{ success, images[], url }`. Each image: `{ src, element (img|svg|css|background|…), type (url|html|base64), alt|null, enrichment{ width, height, mimetype, url, type(photography|illustration|logo|wordmark|icon|…) } }`. The `enrichment` sub-fields populate only for the flags you requested. - -### Crawl Sitemap - -`GET /web/scrape/sitemap` · **1 credit (2 with `search`)** · SDK `client.web.webScrapeSitemap` (Go `Web.WebScrapeSitemap`) -[Guide](https://docs.context.dev/guides/scrape-websites-to-markdown#get-all-urls-of-a-domain) · [API reference](https://docs.context.dev/api-reference/web-scraping/crawl-sitemap) - -- **When:** discover the URL inventory of a domain (cheap, no page content) before deciding what to scrape or crawl. -- **Takes:** `domain` (string, **required**, bare domain, not a URL). Optional: `maxLinks` (default 10000, 1–100000), `urlRegex` (RE2, ≤256 chars, filters which URLs return), `search` (string, 2–200 chars; filters the crawled sitemap to the pages about that phrase, e.g. `pricing and plans`, most relevant first — bumps the call to **2 credits**), `headers` (forwarded to the target URL; bypasses cache), `timeoutMS`. -- **Gives:** `{ success, domain, urls[], meta{ sitemapsDiscovered, sitemapsFetched, sitemapsSkipped, errors } }`. Returns URLs only; it does not fetch page content despite the "Crawl" name. With `search` set, `urls[]` contains only the matching pages, most relevant first. - -### Crawl Website - -`POST /web/crawl` · **1 credit per page** · SDK `client.web.webCrawlMd` (Go `Web.WebCrawlMd`) -[Guide](https://docs.context.dev/guides/scrape-websites-to-markdown#crawl-a-whole-site) · [API reference](https://docs.context.dev/api-reference/web-scraping/crawl-website-&-scrape-markdown) - -- **When:** traverse a site from a seed URL and collect every page's Markdown in one call, e.g. ingesting a docs site or blog into RAG. -- **Takes (JSON body):** `url` (**required**). Optional: `maxPages` (default 100, **hard cap 500**), `maxDepth`, `urlRegex` (limit which links to follow), `followSubdomains` (default false), `includeLinks`/`includeImages`/`shortenBase64Images`/`useMainContentOnly`, `includeSelectors[]`/`excludeSelectors[]`, `pdf`, `includeFrames`, `maxAgeMs`, `waitForMs`, `settleAnimations`, `country`, `stopAfterMs` (soft budget, default 80000, range 10000–110000, returns partial results early), `timeoutMS` (hard abort). -- **Gives:** `{ results[], metadata }`. Each result: `{ markdown, metadata{ url, title, crawlDepth, statusCode, success } }`. Top-level `metadata`: `{ numUrls, maxCrawlDepth, numSucceeded, numFailed, numSkipped }`. Failed pages appear with empty `markdown` and `success: false`; skipped URLs are counted but omitted. **Billed per page crawled**, so set `maxPages` conservatively. - -### Web Search - -`POST /web/search` · **1 credit per 10 results** · SDK `client.web.search` (Go `Web.Search`) -[Guide](https://docs.context.dev/guides/scrape-websites-to-markdown) · [API reference](https://docs.context.dev/api-reference/web-scraping/web-search) - -- **When:** find relevant pages for a natural-language query across the web (optionally scraping each hit to Markdown in the same round-trip) and you don't already have a URL. -- **Takes (JSON body):** `query` (string, **required**, 1–500 chars). Optional: `includeDomains[]`, `excludeDomains[]`, `freshness` (`last_24_hours`|`last_week`|`last_month`|`last_year`), `queryFanout` (bool), `markdownOptions` (off by default; set `enabled: true` to scrape each result, with the same markdown sub-options as Scrape Markdown), `timeoutMS`. -- **Gives:** `{ results[], query }`. Each result: `{ url, title, description, relevance (high|medium|low), markdown{ markdown|null, code } }`. **Always check `markdown.code` first**: `NOT_REQUESTED` (scraping off), `SUCCESS`, `TIMEOUT`, `CONTENT_TOO_LARGE` (result page over the 20 MB cap), `WEBSITE_ACCESS_ERROR`, `ERROR`; only `SUCCESS` guarantees non-null markdown. - ---- +## Handle errors by category -## Design system +Inspect both the HTTP status and `error_code` for request errors. Scrape also reports output failures inside `200` responses: inspect each output's `success` and preserve `isPartial: true`. Use raw HTTPS or the SDK's raw response if an older generated model does not expose these fields. -Extract a site's visual system to reproduce or theme on-brand. XOR rule: pass **exactly one** of `domain` or `directUrl` (omitting both → 400). Guide: [Extract a design system](https://docs.context.dev/guides/extract-design-system-from-website). +| Status | Treat as | Action | +| --- | --- | --- | +| `400` | Invalid request, inaccessible or blocked target on operations that fetch one directly (`WEBSITE_ACCESS_ERROR`, `WEBSITE_BLOCKED`), image-only PDF on Parse (`PDF_IMAGES_ONLY`), skipped PDF (`PDF_SKIPPED`), or no match, depending on `error_code` | Fix the input or options. For `PDF_IMAGES_ONLY`, retry Parse with `ocr=true`. Use a fallback for a no-match or inaccessible site. | +| `401` | Missing or unknown key (`NOT_FOUND`), disabled key (`DISABLED`), or not enough credits (`USAGE_EXCEEDED`) | Fix the key or account state; do not retry unchanged. | +| `403` | Missing key scope (`INSUFFICIENT_PERMISSIONS`, with `required_permission`), paid-plan feature (`PAID_PLAN_REQUIRED`), too many active batches (`BATCH_LIMIT_EXCEEDED`), or ZDR not enabled (`ZDR_NOT_ENABLED`) | Do not retry unchanged; inspect `error_code`. | +| `404` | Target or entity not found on operations that define it | Treat as an expected empty outcome where appropriate. | +| `408` | The request reached `timeoutOpts.milliseconds` with `behavior: "fail"` (`REQUEST_TIMEOUT`) | Retry outside a user-facing path, raise the budget, use `return-partial` where supported, or prefetch supported Brand and Styleguide requests. | +| `200` with a failed Scrape output | Retrieval, parsing, actions, selectors, or output limits prevented that output from completing | Preserve successful outputs and fix the target or options before retrying the failed output. Oversized outputs have `success: false` and `data: null`. | +| `413`, `415` | Content too large (such as a Parse upload over 50 MiB) or unsupported, on operations that define these errors | Use a smaller or supported input. Scrape marks such outputs as failed instead. | +| `422` | An operation-specific input restriction, such as a free email domain or a Brand timeout that is too low | Change the input or options; do not retry unchanged. | +| `429` | Rate limit for this API key | Honor `Retry-After`; retry with jittered, bounded backoff. | +| `500`, `502`, `503` | Transient failure: a service error, an incomplete browser capture, or no browser capacity | Retry with jittered, bounded backoff, then surface a fallback. | -### Styleguide +Do not retry validation, permission, no-match, content-size, or unsupported-media failures unchanged. The SDKs already retry twice; account for that before adding another retry layer. See [Troubleshooting](https://docs.context.dev/optimization/troubleshooting) and [Rate limits](https://docs.context.dev/optimization/rate-limits) for operation-specific behavior. -`GET /web/styleguide` · 10 credits · SDK `client.web.extractStyleguide` (Go `Web.ExtractStyleguide`; Python `client.style.extract_styleguide`) -[Guide](https://docs.context.dev/guides/extract-design-system-from-website#extract-the-full-styleguide) · [API reference](https://docs.context.dev/api-reference/web-extraction/scrape-styleguide) +## Report problems with Agent Feedback -- **When:** you need the full design system: palette, type scale, spacing, shadows, and paste-ready button/card CSS. -- **Takes:** `domain` **or** `directUrl` (one required), `colorScheme` (`light`|`dark`; emulates `prefers-color-scheme` and is included in the cache key), `maxAgeMs` (default `7776000000` ≈ 90d, clamped 1d–1y), `timeoutMS`. -- **Gives:** `{ status, domain, code, styleguide }` where `styleguide` = - - `mode` (`light`|`dark`) - - `colors` `{ accent, background, text }` (hex) - - `typography.headings.{h1..h4}` and `typography.p`, each `{ fontFamily, fontFallbacks[], fontSize, fontWeight, lineHeight, letterSpacing }` - - `elementSpacing` `{ xs, sm, md, lg, xl }` - - `shadows` `{ sm, md, lg, xl, inner }` - - `components.button.{primary,secondary,link}` and `components.card`; each carries a precomputed **`css`** string you can paste directly (note card uses `textColor`, buttons use `color`) - - `fontLinks`: map of family → `{ type (google|custom), files{ "400": url, "700": url, … }, category, displayName }`; may be `{}` when no families resolve to downloadable URLs. +When an endpoint, docs page, SDK, or CLI behaves differently than documented, offer to report it with [Agent Feedback](https://docs.context.dev/optimization/agent-feedback): `POST https://api.context.dev/v1/feedback` with the same bearer key. It works with any API key and has its own rate limit. Send one report per problem with the affected `request_id` so the team can see the exact request, and keep secrets and personal data out of the note. -### Fonts - -`GET /web/fonts` · 5 credits · SDK `client.web.extractFonts` (Go `Web.ExtractFonts`; Python `client.style.extract_fonts`) -[Guide](https://docs.context.dev/guides/extract-design-system-from-website#extract-just-the-fonts) · [API reference](https://docs.context.dev/api-reference/web-extraction/scrape-fonts) - -- **When:** you only need typography: which families a site uses, fallbacks, dominance, and downloadable file URLs. -- **Takes:** `domain` **or** `directUrl` (one required), `maxAgeMs` (default `7776000000` ≈ 90d, clamped 1d–1y), `timeoutMS`. -- **Gives:** `{ status, domain, code, fonts[], fontLinks? }`. Each font: `{ font, uses[], fallbacks[], num_elements, num_words, percent_elements, percent_words }` (word % = text volume, element % = DOM coverage; a mono font can dominate elements but carry few words). `fontLinks` is omitted entirely when nothing resolves. - ---- - -## Screenshots - -### Capture Screenshot - -`GET /web/screenshot` · 1 credit · SDK `client.web.screenshot` (Go `Web.Screenshot`) -[Guide](https://docs.context.dev/guides/take-webpage-screenshot) · [API reference](https://docs.context.dev/api-reference/web-scraping/scrape-screenshot) - -- **When:** a rendered PNG of a page, for link previews, share cards, or visual archives. -- **Takes:** `domain` **or** `directUrl` (one required, XOR). Optional: `fullScreenshot` (**string** `"true"`/`"false"`, not a JSON bool), `handleCookiePopup` (boolean, default `false`), `colorScheme` (`light`|`dark`), `viewport` object `{ width 240–7680 default 1920, height 240–4320 default 1080 }`, `scrollOffset` (integer 0–100000; takes precedence over `fullScreenshot`), `page` (enum `login|signup|blog|careers|pricing|terms|privacy|contact`; auto-finds that page type; **only works with `domain`**, ignored with `directUrl`), `country` (ISO 3166-1 alpha-2), `maxAgeMs`, `waitForMs` (default 3000), `timeoutMS`. -- **Gives:** `{ status, domain, screenshot, screenshotType (viewport|fullPage), width, height, code }`. `screenshot` is a hosted public image URL, not inline bytes. - ---- - -## AI extraction - -LLM-backed structured extraction. All 10 credits, `POST` with JSON body. SDK namespace is `client.ai.*` for the product endpoints and `client.web.*` for structured web extraction (`/web/extract`). - -### Extract a single product - -`POST /brand/ai/product` · 10 credits · SDK `client.ai.extractProduct` (Go `AI.ExtractProduct`) -[Guide](https://docs.context.dev/guides/extract-product-from-websites) · [API reference](https://docs.context.dev/api-reference/web-extraction/extract-a-single-product-from-a-url) - -- **When:** you have one product-page URL and want its structured details. -- **Takes:** `url` (string URI, **required**), `maxAgeMs` (default 7d, 0–30d), `timeoutMS`. -- **Gives:** `{ is_product_page, platform (amazon|tiktok_shop|etsy|generic|null), product|null }`. `product`: `{ name, description, price|null, currency|null, billing_frequency(monthly|yearly|one_time|usage_based)|null, pricing_model(per_seat|flat|tiered|freemium|custom)|null, url, category|null, features[], target_audience[], tags[], image_url|null, images[], sku|null }`. -- **Note:** `product` and `platform` can be `null` even when `is_product_page` is `true`, so **branch on both**. Pages we couldn't reach (blocked, WAF, timeout) and listing/category grids return `400 WEBSITE_ACCESS_ERROR` — not a silent `is_product_page: false`. - -### Extract products from a site - -`POST /brand/ai/products` · 10 credits · **Beta** · SDK `client.ai.extractProducts` (Go `AI.ExtractProducts`) -[Guide](https://docs.context.dev/guides/extract-product-from-websites) · [API reference](https://docs.context.dev/api-reference/web-extraction/extract-products-from-a-brands-website) - -- **When:** discover and extract a brand's product list from its site. -- **Takes (JSON body):** `domain` **or** `directUrl` (exactly one, XOR; both or neither → 400). Optional `maxProducts` (1–12), `maxAgeMs`, `timeoutMS`. (Ruby: `extract_products(body: { domain: … })`; Go: `OfByDomain`.) -- **Gives:** `{ products[] }`, each product the same shape as `/brand/ai/product`'s `product`. `images[]` may be omitted for non-physical products (SaaS). - -### Structured web extraction - -`POST /web/extract` · 10 credits · SDK `client.web.extract` (Go `Web.Extract`) -[Guide](https://docs.context.dev/guides/extract-structured-data-from-websites) · [API reference](https://docs.context.dev/api-reference/web-extraction/query-website-data-using-ai) - -- **When:** extract arbitrary caller-defined data from a site with a JSON Schema. Replaces a scrape-many-pages-then-LLM pipeline. -- **Takes (JSON body):** `url` (**required**) and `schema` (**required**). Optional `instructions`, `maxPages` (1–50, default 5), `maxDepth`, `factCheck`, `followSubdomains`, `pdf`, `includeFrames`, `maxAgeMs` (default 7d), `waitForMs`, `stopAfterMs` (soft crawl budget, default 80000), and `timeoutMS`. -- **Gives:** `{ data, urls_analyzed[], metadata }`, where `data` matches the schema you sent. -- **Legacy:** `POST /brand/ai/query` remains available for older integrations, but prefer `/web/extract` for new structured extraction work. - ---- - -## Industry classification - -Dedicated code lookups. **SDK namespace is `client.industry.*`, not `web.*`**, despite the `/web/` path. Both 10 credits, `GET`. Overview: [Classification](https://docs.context.dev/guides/classification/overview). - -### NAICS - -`GET /web/naics` · 10 credits · SDK `client.industry.retrieveNaics` (Go `Industry.GetNaics`) -[Guide](https://docs.context.dev/guides/classification/NAICS) · [API reference](https://docs.context.dev/api-reference/web-extraction/classify-naics-industries) - -- **When:** you need 2022 NAICS codes for regulatory reporting or TAM segmentation. -- **Takes:** `input` (string, **required**; a domain is preferred, a free-text name also works), `minResults` (1–10, default 1), `maxResults` (1–10, default 5), `timeoutMS`. -- **Gives:** `{ status, domain, type, codes[] }`, each code `{ code, name, confidence (high|medium|low) }`. - -### SIC - -`GET /web/sic` · 10 credits · SDK `client.industry.retrieveSic` (Go `Industry.GetSic`) -[Guide](https://docs.context.dev/guides/classification/SIC) · [API reference](https://docs.context.dev/api-reference/web-extraction/classify-sic-industries) - -- **When:** SEC filings, tax/accounting, or legacy systems that still speak SIC. -- **Takes:** `input` (**required**), `type` (`original_sic` default | `latest_sec`; picks the dataset), `minResults`/`maxResults` (1–10), `timeoutMS`. -- **Gives:** `{ status, domain, type, classification, codes[] }`. `original_sic` codes carry `majorGroup` + `majorGroupName`; `latest_sec` codes carry `office`. Each code also has `code`, `name`, `confidence`. - ---- - -## People enrichment - -### Enrich Person - -`POST /people/enrich` · 20 credits · **Beta, paid plans only** · SDK `client.people.enrich` -[API reference](https://docs.context.dev/api-reference/people/enrich) - -- **When:** you have identity clues for a person — a social profile URL, a work email, or a name plus company/education/location — and want one normalized profile with a confidence score. -- **Takes (JSON body):** any combination of `social_urls[]` (1–20 profile URLs), `name` (`{ first, last }`), `email`, `company` (`{ name, domain }`), `education[]` (`{ institution{name,domain}, degree, field_of_study, graduation_year }`), `location` (`{ city, region, country }`), plus `timeoutMS` and `tags`. All supplied clues are considered together — more clues, better match. -- **Gives:** `{ match, key_metadata }`. `match` is a discriminated union on `status`: `{ status: "candidate", score (0–100), person{ name{full,first,last}, email, avatar_url, bio, location, social_urls[], website_urls[], current_role{title, organization{name,domain}, location} } }` or `{ status: "not_found", score: null, person: null }`. **Branch on `match.status`** and treat low `score` values as weak matches. -- **Note:** free-provider and disposable emails return **422** (`FREE_EMAIL_DETECTED` / `DISPOSABLE_EMAIL_DETECTED`) before any credits are charged — handle as "skip enrichment". Free-tier keys get 403; every paid plan has access. - ---- - -## Prefetch (cache warming) - -Free, no rate limit, **paid-subscriber only** (403 `FORBIDDEN` otherwise). Fire-and-forget: the 200 only confirms that work was queued; it returns no brand or styleguide payload. Guide: [Prefetching](https://docs.context.dev/optimization/prefetching). - -### Prefetch brand data or a styleguide - -`POST /utility/prefetch` · 0 credits · SDK `client.utility.prefetch` (Go `Utility.Prefetch`) -[API reference](https://docs.context.dev/api-reference/utility/prefetch) - -- **When:** you know a domain or work email ahead of when you'll need brand data or a website styleguide and want the later read to land on a warm cache. -- **Takes:** `type: "brand" | "styleguide"` (**required**), `identifier.domain` or `identifier.email` (exactly one required; both or neither returns **400**), `timeoutMS`. Use `brand` before `/brand/retrieve` and `styleguide` before `/web/styleguide`. **Gives:** `{ status, message, type, domain, key_metadata }`. -- **Note:** email identifiers extract the domain automatically; free/disposable emails return **422** (`FREE_EMAIL_DETECTED` / `DISPOSABLE_EMAIL_DETECTED`). - -### Legacy prefetch endpoints - -`POST /brand/prefetch` and `POST /brand/prefetch-by-email` still work unchanged for existing integrations. New integrations should use `POST /utility/prefetch`. - -- **Domain takes:** `domain` (string, **required**), `timeoutMS`. **Gives:** `{ status, message, domain }`. -- **Email takes:** `email` (string, **required**), `timeoutMS`. **Gives:** `{ status, message, domain }` (the extracted domain). -- **Note:** legacy email prefetch returns **422** (`FREE_EMAIL_DETECTED` / `DISPOSABLE_EMAIL_DETECTED`) for free or disposable providers. - ---- - -## Latency - -Cached brand lookups return in **under 1 second** (~60% of calls hit the cache). A cold lookup runs a full crawl: **p50 ≈ 7s, p90 ≈ 18s, p99 ≈ 1 min**. If you know the domain/email ahead of time, [prefetch](https://docs.context.dev/optimization/prefetching) it (free) so the later retrieve lands warm; otherwise set `timeoutMS` generously (up to `300000`). Brand prefetch warms `/brand/retrieve`; styleguide prefetch warms `/web/styleguide`. Neither affects other `/web/*` or `/brand/ai/*` calls. Details: [rate limits](https://docs.context.dev/optimization/rate-limits) · [best practices](https://docs.context.dev/optimization/best-practices). - -## Errors - -Errors carry an `error_code`; through the SDKs they surface as typed exceptions with the status code (the SDKs do **not** auto-retry). Common cases: - -| Status | Meaning | Recovery | -| ------ | --------------------------------------------------------------------- | ---------------------------------------------------------------------- | -| 400 | Malformed input; `WEBSITE_ACCESS_ERROR` (live site blocked/unscrapable); `WEBSITE_NOT_FOUND` (domain doesn't resolve); `NOT_FOUND` (no brand matched — not billed); `PDF_SKIPPED` (target is a PDF and `pdf[shouldParse]=false`); or `PDF_IMAGES_ONLY` (scanned PDF — retry with OCR enabled) | Validate input; treat `WEBSITE_ACCESS_ERROR`/`WEBSITE_NOT_FOUND`/`NOT_FOUND` as "no brand / not found"; retry `PDF_IMAGES_ONLY` with `ocr=true` | -| 401 | Missing/invalid key | Check `CONTEXT_DEV_API_KEY` | -| 403 | `FORBIDDEN` (e.g. prefetch or people enrich without a paid plan) / `USAGE_EXCEEDED` | Check plan / quota | -| 408 | Cold-hit or `timeoutMS` exceeded | Prefetch, raise `timeoutMS`, or retry | -| 413 | `CONTENT_TOO_LARGE` — requested content exceeds the 20 MB download cap | Permanent for that URL; don't retry | -| 422 | Free/disposable email on the `*-by-email` endpoints and `/people/enrich`; or `COLD_DOMAIN_TIMEOUT_TOO_LOW` (`timeoutMS` under 10s on an uncached domain) | Skip enrichment for personal emails; raise `timeoutMS` ≥ 10s or prefetch first | -| 429 | Rate limit | Exponential backoff | +```json +{ + "request_id": "", + "category": "docs_mismatch", + "note": "The response is missing a documented field; expected it per the API reference." +} +``` -Full catalog: [Troubleshooting](https://docs.context.dev/optimization/troubleshooting). +Send `url` instead of, or with, `request_id` for a docs page or one page of a crawl. Categories are `bug`, `docs_mismatch`, `friction`, `feature_gap`, `quality_degradation`, and `other`. Reporting the same `request_id` again returns the original `feedback_id` with `already_submitted: true`. In an MCP session, use the `submit-feedback` tool instead. -## Gotchas +## Completion checklist -- **Fields are nullable.** Logos, colors, stock, phone, links may be absent, so always fall back. -- **Pick the right logo.** Filter by `mode` (`light`/`dark`/`has_opaque_background`) and `type` (`logo` horizontal / `icon` square); don't assume `logos[0]`. -- **Color `name` is generated**, not the brand's official name; key off `hex`. -- **Bare domains only** (`stripe.com`), and **XOR `domain`/`directUrl`** on styleguide, fonts, screenshot, and `/brand/ai/products`. -- **Brand data caches ~3 months** server-side; pass `maxAgeMs: 0` to force a refresh where supported. -- **Logo Link is separate.** For high-volume logo embedding in a UI, use `https://logos.context.dev/?publicClientId=...&domain=...`: a front-end-safe `publicClientId`, no API key, its own quota. See [Get logos from a domain](https://docs.context.dev/guides/get-logo-from-url). +Before declaring an integration complete: -## Reference +- Confirm the path, method, and parameter names against the current API reference. +- Request only the `formats` the task needs, and read each output's `data` rather than assuming it is present. +- Keep the API key on the server and prove the missing-key path is intentional. +- Bound crawl size, latency, retries, and credit exposure. +- Handle missing fields and expected no-result states. +- Validate parsed and structured output and preserve provenance when needed. +- Run a focused success test and at least one relevant failure test. -- Machine index for agents: [docs.context.dev/llms.txt](https://docs.context.dev/llms.txt) -- [Introduction](https://docs.context.dev/introduction) · [Brand data](https://docs.context.dev/guides/get-brand-data) · [Scraping](https://docs.context.dev/guides/scrape-websites-to-markdown) · [Design system](https://docs.context.dev/guides/extract-design-system-from-website) · [Products](https://docs.context.dev/guides/extract-product-from-websites) · [Structured extraction](https://docs.context.dev/guides/extract-structured-data-from-websites) · [Transactions](https://docs.context.dev/guides/enrich-transaction-codes) · [Classification](https://docs.context.dev/guides/classification/overview) -- Optimization: [prefetching](https://docs.context.dev/optimization/prefetching) · [rate limits](https://docs.context.dev/optimization/rate-limits) · [best practices](https://docs.context.dev/optimization/best-practices) · [troubleshooting](https://docs.context.dev/optimization/troubleshooting) · [fair use](https://docs.context.dev/optimization/fair-use) +For long-lived integrations, read [Security and API stability](https://docs.context.dev/optimization/trust) and the [changelog](https://docs.context.dev/changelog).