@exa-labs/exa-contents

Call Exa Contents directly with cURL or raw HTTP. Use when an agent already has URLs and needs POST /contents without an SDK for extracted text, highlights, summaries, links, image links, subpages, freshness-controlled crawling, or per-URL status handling.

View in AI SkillSafe app
Scanned · no findings
0 downloads
0 stars
0 demos
SKILL.md
nameexa-contents
descriptionCall Exa Contents directly with cURL or raw HTTP. Use when an agent already has URLs and needs POST /contents without an SDK for extracted text, highlights, summaries, links, image links, subpages, freshness-controlled crawling, or per-URL status handling.

Exa Contents

Requires API key: Get one at https://dashboard.exa.ai/api-keys

Header: x-api-key: $EXA_API_KEY

Use POST https://api.exa.ai/contents when the agent already knows the URLs and needs clean, LLM-ready extraction without running a new search. Start with one content mode: highlights for compact agent context, text for broad page context, or summary for Exa-side compression.

Quick Start (cURL)

Basic text extraction

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://example.com"],
    "text": true
  }'

Highlights with freshness control

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://arxiv.org/abs/2307.06435"],
    "highlights": {
      "query": "methodology and results"
    },
    "maxAgeHours": 24,
    "livecrawlTimeout": 12000
  }'

Endpoint

POST https://api.exa.ai/contents

Authentication: x-api-key: <API_KEY> header. Exa also accepts Authorization: Bearer <API_KEY>, but prefer x-api-key in cURL examples for consistency.

Use this endpoint for known-URL extraction. If the agent needs discovery or ranking first, use POST /search.

Parameters

Core request parameters

Parameter Type Required Default Description
urls string[] Yes - URLs to extract content from. Use this for known URLs.
text boolean or object No - Return full page text as markdown. Object form supports maxCharacters, includeHtmlTags, verbosity, includeSections, and excludeSections.
highlights boolean or object No - Return key excerpts. Prefer true for agent workflows unless a custom focus is needed.
summary boolean or object No - Return per-page LLM summaries. Use when the caller wants Exa-side compression or structured extraction.
maxAgeHours integer No - Freshness control. 0 always live crawls; -1 uses cache only; omit for default cache-first behavior with crawl fallback.
livecrawlTimeout integer No 10000 Timeout for live crawling in milliseconds. Use 10000 to 15000 for most freshness-sensitive calls.
subpages integer No 0 Number of linked subpages to crawl from each URL.
subpageTarget string or string[] No - Terms used to prioritize which subpages matter, such as ["api", "reference", "pricing"].
extras.links integer No 0 Number of links to extract from each page.
extras.imageLinks integer No 0 Number of image URLs to extract from each page.
compliance string No - Enterprise-only compliance mode, such as hipaa, when enabled for the account.

Text object options

Parameter Type Default Description
maxCharacters integer - Character limit for returned text. Use this instead of tokensNum.
includeHtmlTags boolean false Preserve HTML tags in output.
verbosity string compact compact, standard, or full. Pair fresh section-aware extraction with maxAgeHours: 0.
includeSections string[] - Only include selected sections: header, navigation, banner, body, sidebar, footer, metadata.
excludeSections string[] - Exclude selected sections from the same section list.

Highlights object options

Prefer highlights: true for the highest-quality default. Only use object form when the agent needs a custom focus or budget.

Parameter Type Default Description
query string - Custom query guiding which excerpts are returned.
maxCharacters integer - Cap highlight characters per URL. Omit unless the caller has a strict budget.

Summary object options

Parameter Type Default Description
query string - Custom query for the summary.
schema object - JSON Schema for structured per-page summaries.

Content Modes

On /contents, text, highlights, and summary are top-level request fields.

Mode Best for Notes
text Deep analysis and broad page context Use maxCharacters to keep payloads bounded.
highlights Agent workflows and factual lookups Most token-efficient default. Excerpts are grounded in the source page.
summary Compression or structured per-page extraction Adds Exa-side synthesis per page.

Avoid requesting multiple modes unless the caller truly needs multiple views of the same page.

Freshness and Crawling

Use maxAgeHours as the normative freshness control.

Value Behavior
omitted Use default cache-first behavior with crawl fallback when needed.
positive integer Use cache if it is less than N hours old, otherwise live crawl.
0 Always live crawl. Highest freshness, higher latency.
-1 Cache only. Fastest, but fails if no cached content exists.

Set livecrawlTimeout when live crawling should not block past a fixed budget.

Subpages and extras

Use subpages and subpageTarget when linked pages matter.

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://docs.example.com"],
    "text": {
      "maxCharacters": 5000
    },
    "subpages": 10,
    "subpageTarget": ["api", "reference", "guide"],
    "extras": {
      "links": 10,
      "imageLinks": 5
    }
  }'

Start with subpages around 5 to 10, then increase only when the caller needs broader site coverage.

Response Fields and Statuses

Inspect statuses even when the HTTP status is 200. The endpoint can succeed for one URL and fail for another in the same request.

curl -sS -X POST "https://api.exa.ai/contents" \
  -H "Content-Type: application/json" \
  -H "x-api-key: $EXA_API_KEY" \
  -d '{
    "urls": ["https://example.com", "https://failed-url.example.com"],
    "highlights": true
  }' | jq '{results, statuses}'
Field Type Description
requestId string Unique request identifier.
results array Extracted content result objects.
results[].title string Page title.
results[].url string Page URL.
results[].publishedDate string or null Estimated publication date when available.
results[].author string or null Author when available.
results[].text string Returned when text is requested.
results[].highlights string[] Returned when highlights is requested.
results[].highlightScores number[] Similarity scores for highlights.
results[].summary string Returned when summary is requested.
results[].subpages array Nested result objects from subpage crawling.
results[].extras.links string[] Extracted links when requested.
statuses array Per-URL success or error states. Always inspect this field.
statuses[].id string Requested URL.
statuses[].status string success or error.
statuses[].error.tag string Error type for failed URLs.
statuses[].error.httpStatusCode integer or null HTTP code associated with a per-URL failure.
costDollars.total number Total request cost when returned.

Common per-URL error tags include CRAWL_NOT_FOUND, CRAWL_TIMEOUT, CRAWL_LIVECRAWL_TIMEOUT, SOURCE_NOT_AVAILABLE, UNSUPPORTED_URL, and CRAWL_UNKNOWN_ERROR.

Critical Pitfalls

  • Keep text, highlights, and summary at the top level on /contents.
  • Do not wrap extraction options in a contents object; that nesting belongs to /search.
  • Do not assume HTTP 200 means every URL succeeded; inspect statuses.
  • Do not send stream: true; /contents is not a streaming endpoint.
  • Do not send tokensNum; use text.maxCharacters to cap extracted text.
  • Do not use useAutoprompt, numSentences, highlightsPerUrl, or older livecrawl string values in new requests.
  • Prefer maxAgeHours for freshness and pair it with livecrawlTimeout when crawl latency matters.
  • Use subpageTarget with subpages; otherwise subpage selection is best effort.
  • Pick one of highlights, text, or summary by default. Stack modes only when the caller truly needs multiple views of each page.

Embed badges

Add these to your README to show the skill's verification status.

SkillSafe verified badge
Verified badge
[![SkillSafe verified badge](https://api.skillsafe.ai/v1/badge/@exa-labs/exa-contents/verified)](https://skillsafe.ai/skill/@exa-labs/exa-contents/)
Installs badge
Installs badge
[![Installs badge](https://api.skillsafe.ai/v1/badge/@exa-labs/exa-contents/installs)](https://skillsafe.ai/skill/@exa-labs/exa-contents/)
Scan badge
Scan badge
[![Scan badge](https://api.skillsafe.ai/v1/badge/@exa-labs/exa-contents/scan)](https://skillsafe.ai/skill/@exa-labs/exa-contents/)
Eval pass rate badge
Eval pass rate
[![Eval pass rate badge](https://api.skillsafe.ai/v1/badge/@exa-labs/exa-contents/eval)](https://skillsafe.ai/skill/@exa-labs/exa-contents/)