Ellipse Gradient for Header

Amazon Textract Integration Guide

Integrate Amazon Textract with enterprise workflows through AWS-authenticated HTTPS APIs, Amazon S3 document processing, asynchronous jobs, and structured result mapping.

Amazon Textract integration options at a glance

Amazon Textract exposes regional HTTPS APIs with JSON payloads for text detection, document analysis, expense analysis, and identity-document analysis. Synchronous operations can process supported document bytes directly, while asynchronous operations use Amazon S3 for input and return a JobId for status and result retrieval. Textract can publish selected asynchronous completion notifications through Amazon SNS, although this is not a general webhook facility. Requests use AWS Signature Version 4 and IAM permissions. Martini can orchestrate S3, Textract, SNS or scheduled polling, transform Blocks and specialized result structures, validate confidence and business rules, and deliver results to downstream applications or databases.

Integration pointSupported by Amazon Textract?Common use casesHow Martini supports it
REST-style HTTPS APIsYesUse regional JSON APIs such as DetectDocumentText, AnalyzeDocument, AnalyzeExpense, AnalyzeID, and the corresponding asynchronous operations.Martini can consume REST APIs, map JSON requests and responses, and orchestrate calls in workflows. AWS Signature Version 4 signing may require suitable HTTP authentication configuration or custom JVM-compatible logic.
Bulk / async / batch APIsYesStartDocumentTextDetection, StartDocumentAnalysis, and StartExpenseAnalysis support longer-running processing using S3 input and JobId-based retrieval.Martini can start jobs, persist JobId and source identifiers, poll matching Get operations, retrieve paginated results, and route failed or partial jobs.
File / attachment APIsLimitedSynchronous operations accept supported document bytes, while asynchronous operations use Amazon S3 objects as document input. Textract is not a general-purpose file repository.Martini can coordinate document references and payloads, upload or identify S3 objects through the configured integration architecture, and pass the appropriate input to Textract.
Webhooks / outbound callbacksLimitedTextract can publish selected asynchronous job-completion notifications through Amazon SNS; this is not a general inbound webhook facility.Martini can consume the notification when an AWS-supported delivery path or intermediary exposes it compatibly, or can use scheduled polling instead.
AuthenticationYesTextract uses AWS Signature Version 4 with IAM access keys or temporary STS credentials, region-specific service details, and IAM policies.Martini can configure secure authentication and secrets, while custom JVM-compatible logic can support signing or SDK behavior not covered by standard REST configuration.
Amazon S3 integrationYesS3 supplies asynchronous input documents and can store large or paginated output when the relevant configuration is used.Martini can orchestrate S3 references, Textract jobs, region and permission checks, and downstream persistence.
Database accessNoTextract does not expose a database or SQL interface for extracted data.Martini can retrieve Textract results through the APIs and write normalized data to a supported database or downstream application.

How Amazon Textract exposes data and business events

Amazon Textract REST APIs

Amazon Textract exposes regional HTTPS APIs with JSON payloads for synchronous text detection and structured analysis. Operations include DetectDocumentText, AnalyzeDocument, AnalyzeExpense, AnalyzeID, and matching asynchronous Start and Get operations.

Martini implementation pattern

Martini implementation pattern: Martini builds the operation-specific request, supplies the regional endpoint and AWS-authenticated request, maps the response into an internal model, and routes the extracted result to the target workflow.

Implementation sequence

Receive a document payload or S3 object reference
Select the Textract operation and feature types
Sign and submit the regional HTTPS request
Parse the JSON response and operation-specific structures
Map the result to the target application model
Apply validation and route failures

Amazon Textract asynchronous jobs

Asynchronous Textract operations use an S3 document, return a JobId, and require a matching Get operation to retrieve status and results. Results can be paginated with NextToken.

Martini implementation pattern

Martini implementation pattern: A workflow starts the job, stores the JobId and source identity, then uses a notification or scheduler-driven polling path to retrieve every result page and complete downstream processing.

Implementation sequence

Confirm the source object and permissions
Start the asynchronous Textract operation
Persist the JobId and source document identifier
Receive a completion notification or poll job status
Retrieve result pages until NextToken is absent
Transform and deliver the completed result

Amazon SNS completion notifications

Textract supports completion notifications through Amazon SNS for selected asynchronous workflows. These notifications are service notifications rather than general Textract webhooks and require a compatible AWS delivery topology.

Martini implementation pattern

Martini implementation pattern: Martini receives a notification through an AWS-supported subscriber or intermediary when available, validates the referenced job, and calls the corresponding Get operation rather than treating the notification as the complete analysis result.

Implementation sequence

Publish or configure the SNS completion path
Receive the notification through a compatible endpoint
Validate the job and source correlation
Retrieve the corresponding Textract result
Deduplicate repeated notifications
Acknowledge or route processing failures

Amazon S3 document input

Amazon S3 is the standard input location for asynchronous Textract processing. Bucket, object, region, encryption, ownership, and IAM permissions affect whether the service can read the document.

Martini implementation pattern

Martini implementation pattern: Martini receives a document or object reference, verifies the expected S3 context, starts Textract with the reference, and records the object version or hash for traceability and idempotency.

Implementation sequence

Receive or create the S3 object reference
Validate bucket, region, encryption, and access context
Start the appropriate Textract operation
Persist the object key, version, or content hash
Correlate the result with the source document
Apply retention and sensitive-data handling rules

Common Amazon Textract integration patterns

Pattern 1: Extract invoices and receipts into finance systems

When to use this pattern

Use this pattern when invoices or receipts arrive as PDFs or images and finance teams need structured supplier, total, and line-item data. Asynchronous processing is appropriate when documents are stored in S3 or processing may exceed a normal request window.

Integration direction
Amazon S3
Martini
Amazon Textract
SAP S/4HANA
Example Mapping
Amazon Textract FieldCanonical FieldTarget Field
ExpenseDocuments.SummaryFields.VENDOR_NAMEsupplierNameSupplier
ExpenseDocuments.SummaryFields.INVOICE_RECEIPT_DATEdocumentDateDocumentDate
ExpenseDocuments.SummaryFields.TOTALtotalAmountGrossAmount
ExpenseDocuments.LineItemGroupslineItemsInvoiceItems
Martini implementation pattern

Martini starts StartExpenseAnalysis, tracks the JobId, polls or consumes a supported completion path, and retrieves all pages with GetExpenseAnalysis. It maps summary fields and line items, validates totals and required values, checks for duplicate source documents, and separates extraction errors from downstream posting failures for retry or review.

Martini capabilities used
  • workflows
  • API consumption
  • data mapping
  • business rules
  • scheduled execution
  • error handling
  • secure environment configuration

Pattern 2: Convert forms and tables into case data

When to use this pattern

Use this pattern when semi-structured forms or tabular documents must become searchable, normalized data in a case system or database. AnalyzeDocument with relevant feature types can provide forms, tables, queries, signatures, and relationships.

Integration direction
Document source
Martini
Amazon Textract
Database
Example Mapping
Amazon Textract FieldCanonical FieldTarget Field
Block.BlockTypeelementTypeElementType
Block.TextextractedTextValue
Block.ConfidenceconfidenceConfidence
Relationship.IdsrelatedElementIdsRelatedIds
Martini implementation pattern

Martini submits AnalyzeDocument or its asynchronous equivalent, resolves Block relationships, and transforms the operation-specific response into normalized JSON or relational rows. Validation rules can require key fields and confidence thresholds, while incomplete or inconsistent structures are routed to an exception workflow.

Martini capabilities used
  • workflows
  • API consumption
  • JSON handling
  • data mapping
  • validation
  • database integration
  • error handling

Pattern 3: Process asynchronous jobs with notification or polling

When to use this pattern

Use this pattern for longer-running or higher-volume document processing where synchronous calls could time out. Textract completion notifications through SNS can be used where the AWS delivery topology is compatible; scheduled polling remains an alternative.

Integration direction
Amazon S3
Amazon Textract
Amazon SNS
Martini
Example Mapping
Amazon Textract FieldCanonical FieldTarget Field
JobIdprocessingJobIdWorkflowJobId
JobStatusprocessingStatusDocumentStatus
NextTokenresultContinuationTokenPageCursor
DocumentLocation.S3Object.NamesourceObjectKeySourceDocument
Martini implementation pattern

Martini persists the JobId and source identity immediately after starting the job. It consumes a supported notification or schedules status checks, handles IN_PROGRESS, SUCCEEDED, FAILED, and PARTIAL_SUCCESS states, retrieves every page, and makes completion handling idempotent against duplicate notifications and retries.

Martini capabilities used
  • workflows
  • scheduler triggers
  • API consumption
  • orchestration
  • state correlation
  • retry handling
  • monitoring

Pattern 4: Validate identity-document extraction for onboarding

When to use this pattern

Use this pattern when identity-document images must be converted into structured onboarding or KYC data while preserving confidence checks and manual-review routing.

Integration direction
Onboarding application
Martini
Amazon Textract
Customer-management system
Example Mapping
Amazon Textract FieldCanonical FieldTarget Field
IdentityDocument.IdentityDocumentFieldsidentityFieldsApplicantIdentity
IdentityDocument.ConfidencefieldConfidenceVerificationConfidence
Document.PagespageCountSourcePageCount
DocumentLocationsourceDocumentReferenceDocumentReference
Martini implementation pattern

Martini calls AnalyzeID with the appropriate synchronous or S3-backed workflow, maps identity fields, checks required values and confidence thresholds, and applies application-specific rules. Low-confidence, missing, or contradictory results are routed for manual review rather than automatically submitted.

Martini capabilities used
  • workflows
  • API consumption
  • data mapping
  • validation
  • business rules
  • exception routing
  • secure data handling

Applications commonly integrated with Amazon Textract

Amazon Textract is commonly used alongside AWS storage, notification, orchestration, and compute services, as well as downstream enterprise applications that consume extracted document data. The following are practical integration targets; some represent broader AWS architecture patterns rather than Textract-specific native relationships.

Application Scenario Direction Martini Pattern
Amazon S3 Store source documents for asynchronous analysis and, where configured, store large or paginated Textract results. Amazon S3 → Martini → Amazon Textract Martini receives or identifies the S3 object, starts the appropriate Textract job, persists the JobId and object reference, and retrieves or routes the completed result.
Amazon SNS Deliver selected asynchronous Textract job-completion notifications for workflow correlation. Amazon Textract → Amazon SNS → Martini Martini consumes the notification through a compatible AWS-supported delivery path or intermediary, validates the job reference, and retrieves the corresponding Textract result.
Amazon SQS Buffer notification or document-processing work for reliable asynchronous handling. Amazon SNS → Amazon SQS → Martini A Martini workflow reads queued messages, deduplicates job references, retrieves Textract results, and applies retry and dead-letter handling around downstream processing.
AWS Lambda Run lightweight notification handling or document post-processing around Textract jobs. Amazon S3 → AWS Lambda → Martini → Amazon Textract Martini coordinates the API workflow while Lambda can provide an AWS-side handoff or preprocessing step; the workflow retains correlation and error context across the boundary.
AWS Step Functions Coordinate multi-step document processing, retries, polling, and manual-review branches. Amazon S3 → AWS Step Functions → Amazon Textract → Martini Martini can participate as an API or workflow endpoint, mapping Step Functions inputs and outputs while centralizing target-system validation and transformation.
Salesforce Send extracted customer, application, or document information into Salesforce objects and processes. Amazon Textract → Martini → Salesforce Martini maps Textract output into Salesforce payloads, validates required and confidence-sensitive fields, and routes rejected or incomplete documents for review.
SAP S/4HANA Convert invoice and receipt extraction into finance and procurement processing. Amazon Textract → Martini → SAP S/4HANA A Martini workflow maps ExpenseDocument summary fields and line items into SAP-facing structures, applies totals and duplicate checks, and handles posting failures separately from extraction failures.
Amazon EventBridge Route relevant AWS events to downstream targets where the required event source and rule are configured. AWS event source → Amazon EventBridge → Martini Martini can expose or consume an intermediary API endpoint for routed events, while the specific Textract event coverage is verified for the chosen AWS configuration.

How to build a Amazon Textract integration in Martini

Objective

Establish the regional Textract and supporting AWS integration path with the required IAM permissions and protected credentials.

Instructions in Martini

  • Configure the AWS region, Textract endpoint, and required IAM permissions
  • Store access keys or temporary credentials in Martini secrets or secure environment configuration
  • Confirm S3, SNS, and KMS permissions where the selected workflow requires them
  • Implement or configure AWS Signature Version 4 signing as required

Objective

Select the initiation and completion model based on document size, latency, and operational volume.

Instructions in Martini

  • Use an API or workflow trigger for synchronous processing
  • Use S3 plus a Start... operation for asynchronous processing
  • Choose compatible SNS delivery or scheduled polling for completion
  • Persist the source document identifier and operation type

Objective

Invoke the appropriate Textract operation and retain the identifiers needed for reliable result retrieval.

Instructions in Martini

  • Select DetectDocumentText, AnalyzeDocument, AnalyzeExpense, AnalyzeID, or the appropriate asynchronous operation
  • Pass document bytes or the S3 reference supported by the operation
  • Persist JobId, source object key or version, and an application-level idempotency key
  • Classify authentication, validation, throttling, and service errors

Objective

Transform operation-specific Textract structures into a canonical model suitable for downstream systems.

Instructions in Martini

  • Parse Blocks, Relationships, ExpenseDocuments, or IdentityDocuments according to the operation
  • Retrieve all pages when NextToken is returned
  • Normalize dates, amounts, line items, identity fields, confidence, and source references
  • Preserve relevant geometry and relationship information where downstream users need it

Objective

Validate extracted information before it is posted or used in business decisions.

Instructions in Martini

  • Check required fields and confidence thresholds
  • Reconcile invoice totals and line items where applicable
  • Detect duplicate source documents or repeated job notifications
  • Route low-confidence, incomplete, or inconsistent results for manual review

Objective

Write validated results to downstream applications or databases and provide operational recovery paths.

Instructions in Martini

  • Send mapped data to the target API, database, queue, or file process
  • Separate retryable throttling and transient failures from authorization or input errors
  • Configure backoff, replay, and dead-letter or exception handling as appropriate
  • Monitor workflow logs and avoid writing sensitive document content or PII unnecessarily

Common Amazon Textract data objects used in integrations

ObjectTypical UseCommon target systemsMartini handling
DocumentInput image or PDF supplied as bytes for synchronous processing or as an Amazon S3 object for asynchronous processing.Amazon S3, Amazon Textract, document repositories, downstream workflow applicationsMartini carries the document reference or payload, correlates it with the operation, and applies secure handling and retention rules.
BlockRepresents extracted pages, lines, words, tables, cells, key-value sets, selection elements, and geometry.Databases, case-management systems, search indexes, business APIsMartini maps Blocks and their optional fields and relationships into normalized JSON or relational structures, preserving confidence and geometry where required.
Document analysis jobTracks an asynchronous operation started by a Start... API and retrieved using a JobId.Workflow state stores, queues, monitoring systems, downstream applicationsMartini persists JobId, source identifiers, operation type, and status, then polls or correlates notifications and retrieves all result pages.
RelationshipLinks extracted Blocks, such as key-value sets to values or tables to cells.Normalized document models, databases, case systemsMartini resolves relationship references during transformation and validates that required form, table, or hierarchy connections are present.
ExpenseDocumentContains expense-analysis summary fields, line-item groups, labels, and values for invoices and receipts.SAP S/4HANA, finance applications, accounts-payable workflowsMartini maps and validates totals, dates, suppliers, and line items, then applies duplicate and exception rules before posting.
IdentityDocumentContains normalized identity fields and detected fields from identity-document analysis.Onboarding applications, KYC workflows, customer-management systemsMartini maps identity fields, evaluates confidence and required-field rules, and routes incomplete or low-confidence results for review.

Authentication and security considerations

AWS authentication and IAM

Amazon Textract uses AWS Signature Version 4 rather than OAuth or API keys. Requests require the correct access key or temporary STS credentials, region, service name, timestamp, signed headers, and payload details.

  • Use least-privilege IAM permissions for Textract operations and supporting S3, SNS, and KMS actions.
  • Store credentials in Martini secrets or secure environment configuration rather than workflow definitions or logs.
  • Account for temporary credential session tokens, clock skew, region alignment, and signed-header consistency.
  • Protect invoices, identity documents, and extracted PII with encryption, access controls, retention policies, and log redaction.

Document access

Asynchronous workflows require Textract and the calling integration to have the appropriate access to the S3 bucket and object. Validate cross-account access, object ownership, versioning, encryption, and regional requirements.

Operational considerations for Amazon Textract integrations

Reliability and throughput

  • Respect regional and operation-specific quotas, concurrent-job limits, document-size limits, and page limits.
  • Use exponential backoff and queue-based throttling for high-volume workloads.
  • Persist JobId, source identifiers, operation type, and idempotency keys before polling or processing notifications.
  • Retrieve every result page using NextToken and make page processing safe to repeat.

Validation and schema handling

  • Handle IN_PROGRESS, SUCCEEDED, FAILED, and PARTIAL_SUCCESS states explicitly.
  • Use operation-specific mappings because text detection, document analysis, expense analysis, and identity analysis return different structures.
  • Validate confidence, required fields, dates, amounts, totals, and duplicate documents before downstream posting.
  • Test representative document variations and tolerate optional fields, missing relationships, and new response properties.

Monitoring and recovery

Distinguish retryable throttling and transient service errors from invalid input, authorization, missing jobs, and encryption failures. Retain enough source and job context to replay a failed document safely without duplicating downstream results.

Why use Martini instead of scripts or point-to-point integrations?

Orchestrate the complete document lifecycle

Scripts often handle a single Textract call but leave S3 coordination, asynchronous job state, pagination, notification delivery, validation, and downstream delivery scattered across separate components. Martini provides a workflow-based approach for coordinating these stages.

  • Consume Textract HTTPS APIs and coordinate S3, notification, and polling paths.
  • Map Blocks, Relationships, ExpenseDocuments, and IdentityDocuments into reusable target models.
  • Apply confidence checks, duplicate detection, totals validation, and manual-review routing as explicit business rules.
  • Centralize error handling, retries, monitoring, secure configuration, and operational context.
  • Expose controlled APIs when other applications need a consistent document-processing façade instead of calling Textract directly.

This approach separates vendor-specific AWS behavior from enterprise workflow logic, making document-processing integrations easier to maintain as target systems and business rules change.

Frequently asked questions

How can Amazon Textract be integrated with enterprise systems?

Amazon Textract can be integrated through its regional HTTPS APIs with JSON payloads. Synchronous operations accept supported document bytes, while asynchronous operations use Amazon S3, return a JobId, and expose results through matching Get operations. Selected asynchronous workflows can publish completion notifications through Amazon SNS. Enterprise workflows can then map Blocks, ExpenseDocuments, IdentityDocuments, and related structures into applications, databases, queues, or APIs.

Can Martini integrate with Amazon Textract?

Yes. Martini can integrate with Amazon Textract through its HTTPS APIs, Amazon S3 document workflows, and notification or scheduled-polling patterns. Martini can orchestrate jobs, map operation-specific responses, apply validation and business rules, and deliver results to downstream systems. AWS Signature Version 4 signing may require appropriate REST authentication configuration or custom JVM-compatible implementation.

Do I need a connector to integrate Amazon Textract with Martini?

No. A dedicated Amazon Textract connector is not required. Martini can use Amazon Textract's native HTTPS APIs, S3-based asynchronous processing, IAM authentication, and selected SNS completion-notification patterns. Custom JVM-compatible logic is an option when AWS signing or SDK behavior is not covered by the configured API consumption approach.

Is there any extra Lonti cost to integrate Amazon Textract with Martini?

Lonti does not charge an additional per-connector or per-vendor fee to integrate Amazon Textract. Integrations are subject to the provisioned capacity of the Martini environment. Separate costs may apply from AWS, including Textract, S3, SNS, compute, storage, or infrastructure usage, and from other third-party systems according to their pricing and deployment models.

Which Amazon Textract integration methods should an enterprise use?

Use the regional HTTPS APIs as the primary integration method. Choose synchronous operations for suitable smaller or immediate requests, and asynchronous Start and Get operations with Amazon S3 for longer-running or larger document workflows. AWS Signature Version 4 and IAM are required for authentication. Textract does not provide a confirmed GraphQL, SOAP, or database interface.

Does Amazon Textract provide events or webhooks?

Textract supports selected asynchronous job-completion notifications through Amazon SNS. This is an AWS service notification pattern, not a general webhook facility for arbitrary Textract events. Martini can process a notification when an AWS-supported delivery path or intermediary exposes it compatibly, or can poll the matching Get operation with a scheduled workflow.

How should Amazon Textract synchronization and pagination work?

For asynchronous processing, persist the source document identifier, operation type, and JobId. Retrieve results through the corresponding Get operation and continue while NextToken is present. Use idempotency keys based on the source object, document hash, business transaction, or JobId so duplicate notifications and workflow retries do not create duplicate downstream records.

How does Martini handle Amazon Textract mapping, errors, and retries?

Martini maps operation-specific structures such as Blocks, Relationships, ExpenseDocuments, and IdentityDocuments into canonical and target models. Workflows can validate confidence, required fields, totals, and relationships before delivery. Retryable throttling and transient service errors should be handled with backoff, while invalid input, authorization failures, and low-confidence results should be routed to correction or manual-review paths.