.png)
Azure AI Document Intelligence Integration Guide
Azure AI Document Intelligence integrates through REST APIs that asynchronously analyze document content or URLs with prebuilt and custom models.
Azure AI Document Intelligence integration options at a glance
Azure AI Document Intelligence provides HTTPS REST APIs for submitting documents or document URLs, selecting prebuilt or custom models, retrieving model information, and polling long-running analysis operations. Analyze requests return an operation location that the client checks until processing completes; the service does not provide a general-purpose webhook mechanism for completed analysis. Authentication uses either an Azure resource key or Microsoft Entra ID bearer tokens. Martini can consume these REST endpoints, store credentials in secrets, orchestrate polling with retries and timeouts, and map extracted fields, tables, confidence values, and page data into downstream applications or databases.
Common Azure AI Document Intelligence integration patterns
Common Azure AI Document Intelligence data objects used in integrations
Authentication and security considerations
Authentication options
Azure AI Document Intelligence supports an Azure resource key sent in the Ocp-Apim-Subscription-Key header or a Microsoft Entra ID OAuth 2.0 bearer token. Entra ID access requires an appropriate role assignment on the Document Intelligence resource.
Credential protection
Martini workflows should load keys, client credentials, tokens, and endpoints from environment configuration or secrets rather than embedding them in workflow definitions. Protect document URLs, raw analysis results, and extracted personal or financial information.
Network and data protection
- Verify the Azure resource region, endpoint, firewall rules, private networking, and managed identity configuration before deployment.
- Use appropriately scoped and time-limited document URLs when URL-based input is selected.
- Restrict logs and retained payloads so sensitive document content is not exposed unnecessarily.
Operational considerations for Azure AI Document Intelligence integrations
Asynchronous processing
Analyze requests return an operation location and require polling. Persist the operation reference, use a controlled interval with backoff, and enforce a timeout.
Throttling and retries
Azure quotas and throttling can affect submission and polling. Retry transient responses such as HTTP 429 with backoff, while separating authentication, invalid-input, and permanent model errors from retryable failures.
Idempotency and result size
Use a source identifier, content hash, or external correlation ID to prevent duplicate analysis. Large documents can produce substantial page, table, polygon, span, and cell data, so mappings should support multi-page results and configurable retention.
Model and schema changes
Configure the API version and model ID explicitly. Validate expected fields because model capabilities, regional availability, and response schemas can vary. Test prebuilt and custom model changes before production deployment.
Extraction quality
A technically successful analysis is not necessarily business-valid. Apply required-field, total, date, identifier, and confidence rules, and route incomplete or low-confidence results for review.
Why use Martini instead of scripts or point-to-point integrations?
Orchestration beyond a single API call
Scripts can call the Azure endpoint, but enterprise processing also requires intake, asynchronous polling, timeouts, retries, validation, target-system writes, duplicate prevention, and audit handling. Martini organizes these concerns in maintainable workflows.
Reusable integration logic
Martini can expose a controlled REST API for document intake, consume Azure REST endpoints, and reuse mappings and business rules across invoice, form, contract, and identity-document processes.
Reliable data movement
- Centralize secrets and environment-specific configuration.
- Map fields, tables, pages, and confidence values into different target models.
- Separate transport failures from extraction-quality exceptions.
- Persist operation state and raw results when resumability, auditability, or reprocessing is required.