Skip to content
Documentation/Documents

Extract anchored chunks

Split extracted content into chunks that retain page, sheet, or section locations.

Markdown
POST/api/v1/extract/pageshttps://api.webcite.co

When to use it

Use these chunks for retrieval and citation display. Preserve the source anchors when you index the text. token_estimate is an estimate for context budgeting, not a provider billing count.

Request

1 credit for a billable extraction outcome.

Examples run on your server. Set WEBCITE_API_KEY first. Python examples use the requests package.

curl
curl --fail-with-body -X POST 'https://api.webcite.co/api/v1/extract/pages' \
  -H "x-api-key: $WEBCITE_API_KEY" \
  -H 'Content-Type: application/json' \
  -d '{
  "asset_id": "YOUR_ASSET_ID"
}'

Request body

asset_idstring

An uploaded asset id (from POST /upload). One of asset_id / asset_url is required.

asset_urlstring

A direct URL to the file. One of asset_id / asset_url is required.

Response

Returns chunks with contiguous zero-based ordinal, text, token_estimate, and available source anchors, plus extraction state.

200 response

application/json

chunksobject[]required

Anchored chunks, in document order. Empty both when the document held no text and when the read was refused , code is what tells those apart.

Show chunks fields
ordinalnumberrequired

Position of this chunk in the document, from 0.

pagenumber

1-based page the chunk came from.

sectionstring

Nearest preceding markdown heading.

sheetstring

Workbook tab the chunk came from.

textstringrequired

The chunk text.

token_estimatenumberrequired

Rough token count (~4 chars/token) for packing a retrieval budget.

codestring | nullrequired

The stable classification of reason, safe to switch on. The same vocabulary and the same values as POST /extract. Null when nothing refused.

CodeWhat it meansWhat to do
source_too_largea size or expansion ceiling refused the readreason carries the observed value and the limit; send a smaller file
source_encryptedthe container is password-protectedsend an unlocked copy
source_corruptthe container could not be opened as the format it declaresre-export the file
unsupported_formatno reader is registered for this formatconvert it
ocr_unavailablethe page had no text layer and no vision provider is configuredconfigure a vision key, or send a text-layer file
partial_extractionsome of the source was recovered and some was notuse what came back
empty_sourcethe source was read and held nothinga fact about the file, not a failure , nothing to retry
extraction_errorour failureretry

Values: "source_too_large""source_encrypted""source_corrupt""unsupported_format""ocr_unavailable""partial_extraction""empty_source""extraction_error"

reasonstring | nullrequired

Free text naming the cause, straight from whatever refused or failed , spreadsheet_magic_mismatch:xlsx, spreadsheet_too_large:bytes:33554433>33554432, extraction_failed:File is password-protected. Carries detail no enum can, and is NOT stable: it includes dependency error text. Display it; do not branch on it. Null when nothing refused.

statestringrequired

How much of the source the read recovered, in one word. complete on a healthy read, whatever the payload beside it turned out to contain.

Values: "complete""partial""unsupported""error"

Errors

For 400, check the request fields and source identifiers. For 401, check your API key. For 429, wait for the retry interval. See the errors guide before retrying a billable request.

400
Asset not found
401
Unauthorized - API key required
429
Rate limit exceeded
Error handling and retry guidance