.png)
Apify Integration Guide
Integrate Apify Actors, asynchronous runs, datasets, files, and selected webhook events with enterprise applications through Martini workflows and APIs.
Apify integration options at a glance
Apify’s primary integration mechanism is its REST API, which can start Actors and Tasks, inspect asynchronous runs, retrieve Dataset items, access Key-value store records, manage Request queues, and configure Webhooks. Selected Apify events can notify a Martini API when a run reaches a configured status, reducing the need for continuous polling. Dataset results can be retrieved as JSON or exported in formats such as CSV, XML, or RSS, while Key-value stores can contain screenshots, HTML, binary files, and other artifacts. Martini stores Apify API tokens as secrets, orchestrates run-and-retrieve workflows, maps Actor-specific schemas, and delivers validated results to downstream systems.
Common Apify integration patterns
Common Apify data objects used in integrations
Authentication and security considerations
API-token protection
Apify’s public API primarily uses API tokens, commonly supplied as Bearer credentials or, where appropriate, a token query parameter. Martini should store tokens in secrets or protected environment configuration rather than workflow definitions, URLs, payloads, or logs.
Least-privilege access
Use credentials with access limited to the Actors, runs, Datasets, Key-value stores, Request queues, and Webhooks required by each workflow. Separate development, test, and production credentials where isolation is needed.
Webhook validation
Secure Martini’s receiving API and validate the shared secret, signature, token, or other configured authentication value according to the Apify webhook configuration. Do not process output solely because a callback was received.
Data governance
Extracted pages, screenshots, HTML, and datasets may contain sensitive or regulated information. Assess site permissions, terms of service, privacy obligations, robots directives, copyright, regional restrictions, and retention requirements before storing or distributing results.
Operational considerations for Apify integrations
Rate limits and capacity
Apify request limits and Actor capacity can vary by account plan, resource, and workload. Use bounded backoff for HTTP 429 and transient 5xx responses, avoid unnecessary polling, and account for Actor compute, concurrency, proxy, storage, and platform usage costs.
Asynchronous execution
Do not assume an Actor completes during the submission request. Persist run and output identifiers, handle transitional and terminal statuses explicitly, and use webhooks or scheduled polling according to event coverage.
Pagination and checkpoints
Process large Datasets in pages or batches rather than loading everything into memory. Track the relevant offset, cursor, or last successful item and use stable identifiers when output may change between retrieval operations.
Idempotency and retries
Webhook delivery or workflow replay can cause duplicate processing. Use run IDs, Dataset item identifiers, source URLs, or business keys, prefer downstream upserts, and ensure retries do not unintentionally submit duplicate Actor runs.
Schema and testing
Actor input and output schemas are Actor-specific and may change when an Actor is upgraded or replaced. Validate required fields, test representative success and failure states, and monitor API responses, run status, mapping failures, and downstream writes.
Why use Martini instead of scripts or point-to-point integrations?
Orchestrate complete integration lifecycles
Scripts often combine submission, polling, output retrieval, transformation, and target writes in one fragile process. Martini separates these concerns into reusable workflows and APIs that can support scheduled, event-driven, and on-demand execution.
Manage changing Actor schemas
Apify output varies by Actor. Martini provides explicit mappings, validation, business rules, and transformation steps so Actor-specific data can be converted into governed canonical and target models.
Improve operational recovery
Martini can correlate run identifiers, apply bounded retries, checkpoint batches, handle webhook reconciliation, and route failures for review without exposing credentials or sensitive extracted content in logs.
Expose controlled enterprise APIs
Martini can provide a stable API façade over Apify workloads, applying authentication, authorization, input validation, and business rules while keeping Apify implementation details behind a maintainable integration boundary.