Skip to content
ComPDF

PDF Generation Template Editor

Open-source visual PDF generation engine with customizable templates and developer-friendly APIs.

View on GitHub

Intelligent Document Extraction API

BASE URLhttps://api-server.compdf.com/server/

❖ Feature Description

Intelligently extract key fields and structured information from documents.

❖ Request Mode

Synchronous Request (Sync)
The API returns the result file directly after processing. Recommended for small files and real-time interactive scenarios that need immediate feedback.
Asynchronous Request (Async)
The API first returns task acceptance information, then you query progress and results with taskId. Suitable for large files and batch workloads.
Secure Request Mode
Upload and process files through secure mechanisms such as pre-signed URLs. Suitable for high-security and privacy compliance scenarios.

▎Call Flow

1Upload file
2Call API (sync)
3Get result URL
4Download file

▎Usage Limits

Download validity24 hours

Execute synchronously

POSThttps://api-server.compdf.com/server/v2/process/idp/documentExtract

❖ Request Parameters

Authentication credential sent in the header: x-api-key

Body Parameters multipart/form-data

No file selected
Upload file
File password (if the PDF is password-protected)
API error message language (1 = English, 2 = Chinese)
Page range. Page numbers start from 1, for example 1-3,6. Empty means all pages.
Enable OCR (0 = off, 1 = on)
OCR recognition language code. View supported languages.
OCR strategy: ALL, SCAN_PAGE, INVALID_CHARACTER, or INVALID_CHARACTER_AND_SCAN_PAGE
Whether to output one file per page (0 = no, 1 = yes)
Extraction mode: vision (default; performs page-by-page extraction using a vision model and offers better handwritten-text recognition) or layout (performs integrated layout-based extraction and supports large files, cross-page processing, and result traceability). Both modes use extract_fields to provide a fixed extraction schema.
JSON string that defines the extraction schema (snake_case alias). Both vision and layout modes use this field; layout mode no longer uses the document_types array.
Optional. Use snake_case.
Optional. options_json is a JSON string that can include model configuration parameters and ignore_labels. ignore_labels supports only number, footnote, header, header_image, footer, footer_image, and aside_text; any supplied label is ignored. Complete example: {"use_doc_unwarping":false,"use_chart_recognition":false,"use_seal_recognition":false,"use_ocr_for_image_block":false,"use_layout_detection":true,"layout_shape_mode":"auto","merge_tables":true,"relevel_titles":true,"concatenate_pages":false,"ignore_labels":[]}

❖ Response Properties

FieldTypeDescription
codeStringBusiness status code
msgStringMessage
dataObjectResponse data
data.fileKeyStringUnique key of the file in the storage system.
data.taskIdStringTask ID
data.fileNameStringSource file name. Required in presigned mode to generate the object storage upload URL.
data.downFileNameStringOutput file name after conversion.
data.fileUrlStringSource file storage URL or object storage key.
data.downloadUrlStringFile download URL
data.sourceTypeStringSource file type
data.targetTypeStringTarget file type
data.fileSizeIntegerSource file size in bytes.
data.convertSizeIntegerConverted file size in bytes.
data.convertTimeIntegerConversion time for a single file, typically in milliseconds.
data.statusStringFile processing status. Common values: success, failed, processing, etc.
data.failureCodeStringError code when file conversion fails.
data.failureReasonStringError reason when file conversion fails.
data.fileParameterStringConversion parameter JSON string submitted when creating the task.
🔗Request Example
curl --request POST \
  --url https://api-server.compdf.com/server/v2/process/idp/documentExtract \
  --header 'x-api-key: YOUR API-KEY' \
  --form [email protected] \
  --form mode=vision
Response Example
200 OK
{
  "code": "200",
  "msg": "success",
  "data": {
    "fileKey": "<string>",
    "taskId": "<string>",
    "fileName": "<string>",
    "downFileName": "<string>",
    "fileUrl": "<string>",
    "downloadUrl": "<string>",
    "sourceType": "<string>",
    "targetType": "<string>",
    "fileSize": 0,
    "convertSize": 0,
    "convertTime": 0,
    "status": "<string>",
    "failureCode": "<string>",
    "failureReason": "<string>",
    "fileParameter": "<string>"
  }
}