Ellipse Gradient for Header

Apify Integration Guide

Integrate Apify Actors, asynchronous runs, datasets, files, and selected webhook events with enterprise applications through Martini workflows and APIs.

Apify integration options at a glance

Apify’s primary integration mechanism is its REST API, which can start Actors and Tasks, inspect asynchronous runs, retrieve Dataset items, access Key-value store records, manage Request queues, and configure Webhooks. Selected Apify events can notify a Martini API when a run reaches a configured status, reducing the need for continuous polling. Dataset results can be retrieved as JSON or exported in formats such as CSV, XML, or RSS, while Key-value stores can contain screenshots, HTML, binary files, and other artifacts. Martini stores Apify API tokens as secrets, orchestrates run-and-retrieve workflows, maps Actor-specific schemas, and delivers validated results to downstream systems.

Integration pointSupported by Apify?Common use casesHow Martini supports it
REST APIsYesManage Actors, Actor runs, Tasks, Datasets, Key-value stores, Request queues, and Webhooks; submit JSON input and retrieve status or output references.Martini can consume the Apify REST API, map request and response payloads, store run identifiers, and orchestrate downstream processing.
Webhooks / outbound callbacksLimitedNotify an external endpoint when a configured Actor run reaches a selected status or other supported Apify event.Martini can expose a secured REST API, validate the callback, correlate it with a run, and retrieve the referenced output.
Bulk / async / batch APIsLimitedAsynchronously execute Actors and retrieve collection-style Dataset results; this is not a general-purpose bulk transaction API for arbitrary business records.Martini can submit runs, poll or await notifications, process Dataset pages in batches, and checkpoint successful work.
File / attachment APIsYesRetrieve JSON, text, binary files, screenshots, HTML, and other Actor artifacts stored in Key-value stores.Martini can retrieve objects, validate metadata and content, transform files, and forward them to APIs, repositories, databases, or object storage.
Dataset APIs and exportsYesRead structured Actor output and request formats such as JSON, CSV, XML, or RSS where supported by the endpoint.Martini can page through results, map Actor-specific schemas, validate fields, deduplicate items, and deliver them to target systems.
Request queuesYesInspect, add, or manage URLs and other crawl requests used by crawling Actors.Martini can use the Apify API to seed or inspect queues and apply business rules around crawl inputs and processing state.
AuthenticationYesAuthenticate public API requests with Apify API tokens supplied as Bearer credentials or, where appropriate, a token query parameter.Martini stores tokens as environment secrets and injects them into requests without hard-coding credentials in workflows or logs.
GraphQL APIsNot confirmedNo public GraphQL mechanism was confirmed for the standard Apify platform API.Martini should use the confirmed Apify REST API instead of assuming GraphQL support.
SOAP APIsNot confirmedNo SOAP API was confirmed for the standard Apify platform API.Martini should use the confirmed Apify REST API and selected Webhooks instead.

How Apify exposes data and business events

Apify REST APIs

Apify exposes a REST API for managing Actors, Actor runs, Tasks, Datasets, Key-value stores, Request queues, and Webhooks. It is the primary integration mechanism for starting workloads and retrieving their status and output.

Martini implementation pattern

Martini implementation pattern: a workflow authenticates with an Apify API token, submits an Actor or Task request, persists returned identifiers, retrieves status and output references, and sends mapped results to downstream systems.

Implementation sequence

Authenticate with an Apify API token stored as a Martini secret
Submit the Actor or Task input through the Apify REST API
Persist the returned run identifier and correlation data
Retrieve run status and output references
Read Dataset or Key-value store output
Map validated results to the target system

Apify Webhooks

Apify supports webhook-style notifications for selected platform events and configured Actor run conditions. Coverage is event- and configuration-dependent rather than universal across all Apify operations.

Martini implementation pattern

Martini implementation pattern: Martini exposes a secured REST endpoint, validates the Apify notification, correlates the run, retrieves current status and output, and starts the downstream processing workflow. A reconciliation workflow can cover missed callbacks.

Implementation sequence

Expose a secured Martini REST endpoint for the callback
Validate the configured webhook authentication value and payload
Correlate the notification with the Apify run
Retrieve current run status and output references
Process Dataset or Key-value store output
Record completion and support reconciliation for missed notifications

Asynchronous Actor runs

Apify Actors and Tasks can run asynchronously, which is useful for long-running crawls and browser automation. The initial request returns identifiers that can be used for later status and output retrieval.

Martini implementation pattern

Martini implementation pattern: a workflow submits the run and branches to webhook-driven or scheduled polling logic. It handles transitional and terminal statuses separately, prevents duplicate submissions, and processes output only after successful completion.

Implementation sequence

Submit an asynchronous Actor or Task run
Store the Actor, Task, and run identifiers
Poll status or wait for a supported webhook
Handle SUCCEEDED, FAILED, ABORTED, and transitional states
Retrieve the associated output
Retry or route unrecoverable failures for operational review

Datasets and exports

Datasets hold structured results produced by Actors. Depending on the endpoint and requested format, results can be returned as JSON, CSV, XML, or RSS, with the exact schema determined by the Actor.

Martini implementation pattern

Martini implementation pattern: Martini retrieves Dataset items in pages or batches, validates the Actor-specific contract, transforms fields into a canonical model, and writes idempotently to a target application or database.

Implementation sequence

Identify the Dataset associated with the completed run
Retrieve items or an appropriate export format
Process pages or batches with checkpoints
Validate required Actor-specific fields
Map and transform values for the target model
Upsert results and record the processed position

Key-value stores and artifacts

Key-value stores can expose JSON, text, binary files, screenshots, HTML, and other artifacts created by an Actor. They provide a practical output path for browser and extraction workloads.

Martini implementation pattern

Martini implementation pattern: a workflow retrieves the named object, validates content type and size, applies retention or sensitivity rules, and forwards the artifact to a repository, object store, case system, or internal API.

Implementation sequence

Identify the Key-value store and record key
Retrieve the object through the Apify API
Validate content type, size, and metadata
Apply business and data-governance rules
Forward or transform the artifact
Record the destination and processing result

Common Apify integration patterns

Pattern 1: Run an Actor and load extracted data

When to use this pattern

Use this pattern when a scheduled or business-triggered process must collect permitted web data and load it into a CRM, database, warehouse, or other enterprise application. The Actor-specific output should be treated as a versioned contract.

Integration direction
Scheduler
Martini
Apify
Salesforce or Snowflake
Example Mapping
Apify FieldCanonical FieldTarget Field
Actor output companyNamecompany.nameAccount.Name
Actor output sourceUrlsource.urlExternal_Source_URL__c
Actor output lastSeenPriceproduct.priceCurrent_Price__c
Martini implementation pattern

A scheduled Martini workflow submits JSON input to an Actor, stores the run identifier, polls or receives a supported completion notification, and retrieves Dataset pages. It validates required fields, normalizes values, deduplicates by source URL or business key, and upserts the target records. Bounded retries handle transient API failures, while Actor execution failures are routed separately from mapping failures.

Martini capabilities used
  • workflows
  • scheduler triggers
  • API consumption
  • data mapping
  • business rules
  • validation
  • error handling

Pattern 2: Process completed-run notifications

When to use this pattern

Use this pattern when crawl duration is variable and polling would create unnecessary API traffic. It is appropriate only when the required Apify run event and status are supported by the configured webhook.

Integration direction
Apify
Martini
Slack or Jira
Example Mapping
Apify FieldCanonical FieldTarget Field
run.idjob.idExternal_Run_ID
run.statusjob.statusIssue_Status
run.defaultDatasetIdoutput.datasetIdDataset_Reference
Martini implementation pattern

Martini exposes a secured API to receive the callback, validates the configured authentication value, and retrieves current Apify state rather than trusting the notification alone. A workflow applies status rules, retrieves output when successful, and creates a Slack notification or Jira issue for failures. Run IDs and correlation keys provide replay protection and reconciliation.

Martini capabilities used
  • API exposure
  • webhook consumption
  • workflows
  • business rules
  • data mapping
  • error handling

Pattern 3: Publish competitive-intelligence changes

When to use this pattern

Use this pattern for permitted monitoring of product pages, listings, prices, rankings, or other public web data. It compares current output with prior processed data and publishes only material changes.

Integration direction
Scheduler
Martini
Apify
Slack or database
Example Mapping
Apify FieldCanonical FieldTarget Field
product.urlitem.sourceUrlSource_URL
product.priceitem.currentPriceCurrent_Price
product.availabilityitem.availabilityAvailability_Status
Martini implementation pattern

A scheduled workflow starts the Actor, retrieves a completed Dataset, and compares stable business keys and relevant values with a prior snapshot. Martini applies thresholds, suppresses duplicate alerts, writes the new checkpoint, and sends approved changes to Slack or a database. The design should separately account for site permissions, privacy, terms of service, and retention requirements.

Martini capabilities used
  • scheduler triggers
  • API consumption
  • data mapping
  • business rules
  • deduplication
  • validation
  • monitoring

Pattern 4: Retrieve browser artifacts and archive them

When to use this pattern

Use this pattern when an Actor produces screenshots, HTML, JSON documents, or other files that must be retained or passed to another process.

Integration direction
Apify
Martini
Amazon S3
Example Mapping
Apify FieldCanonical FieldTarget Field
key-value record keyartifact.nameObject_Key
content typeartifact.contentTypeContent_Type
run.idartifact.sourceRunIdApify_Run_ID
Martini implementation pattern

Martini retrieves the Key-value store object after run completion, validates its type and size, adds run and retention metadata, and forwards it to approved object storage or an internal API. Failed transfers are retried safely using the run ID and artifact key as an idempotency key, while sensitive extracted content is excluded from operational logs.

Martini capabilities used
  • API consumption
  • workflows
  • file handling
  • data mapping
  • validation
  • error handling
  • secrets management

Applications commonly integrated with Apify

Apify can provide extracted data, monitoring results, and browser artifacts to business applications through its REST API and storage resources. The following are common or practical enterprise targets; the exact flow depends on the Actor, output schema, permissions, and target API.

Application Scenario Direction Martini Pattern
Salesforce Load extracted company, contact, lead, or account intelligence into Salesforce for sales research and enrichment. Apify → Martini → Salesforce A scheduled Martini workflow starts an Actor, monitors the run through polling or a selected webhook, retrieves Dataset items, validates the Actor-specific fields, and upserts Salesforce objects using stable business keys.
HubSpot Enrich Companies and Contacts or deliver public company research to marketing and sales processes. Apify → Martini → HubSpot Martini submits JSON Actor input, retrieves the completed Dataset, normalizes company and contact fields, applies deduplication rules, and sends approved updates to HubSpot APIs.
Shopify Support approved monitoring of public product listings, pricing, availability, or catalog information for competitive-analysis workflows. Shopify or public web sources → Apify → Martini → Internal systems An Apify Actor collects permitted source data, while Martini retrieves and compares Dataset results, applies change thresholds, and forwards relevant changes to internal APIs or data stores.
Google Sheets Deliver smaller Dataset outputs to analysts for review, reconciliation, or operational handoff. Apify → Martini → Google Sheets Martini retrieves paginated Dataset items, converts the Actor-specific schema into tabular rows, validates required columns, and writes batches through the Google Sheets API.
Slack Notify teams about completed crawls, failed Actor runs, detected changes, or threshold breaches. Apify → Martini → Slack A Martini webhook or polling workflow evaluates run status and business rules, then sends concise notifications to Slack through its API or an incoming webhook without exposing Apify credentials.
Jira Create issues for detected website, catalog, or content changes that require investigation. Apify → Martini → Jira Martini compares current and prior Dataset values, deduplicates alerts by source URL or business key, and creates or updates Jira issues through its REST API.
Snowflake Load larger Dataset outputs into an analytical environment for trend analysis and reporting. Apify → Martini → Snowflake Martini processes Dataset pages or exports, maps Actor-specific fields to a governed analytical model, checkpoints successful batches, and delivers data using approved Snowflake ingestion or database interfaces.
Amazon S3 Archive Dataset exports, screenshots, HTML, and other Actor artifacts for retention or downstream processing. Apify → Martini → Amazon S3 Martini retrieves Key-value store objects or Dataset exports, validates content type and metadata, applies retention rules, and writes artifacts to S3 with correlation identifiers.

How to build a Apify integration in Martini

Objective

Configure secure access to Apify and any downstream applications without embedding credentials in workflow definitions.

Instructions in Martini

  • Store the Apify API token in Martini secrets or environment configuration
  • Use the token as a protected Bearer credential for Apify API calls
  • Limit token permissions to required Actors, runs, storage, queues, or Webhooks
  • Configure separate credentials for development, testing, and production where needed

Objective

Select a trigger that matches the workload duration and notification coverage.

Instructions in Martini

  • Use a scheduler for recurring Actor or Task execution
  • Use an API trigger for on-demand runs
  • Use a Martini API to receive supported Apify webhook notifications
  • Retain scheduled polling or reconciliation for missed or unsupported events

Objective

Submit the workload and reliably obtain its status and output references.

Instructions in Martini

  • Send Actor or Task input as JSON through the Apify REST API
  • Persist Actor, Task, run, Dataset, and Key-value store identifiers as applicable
  • Poll status or process a validated webhook notification
  • Handle successful, failed, aborted, and transitional run states separately

Objective

Convert Actor-specific output into a governed internal or target model.

Instructions in Martini

  • Retrieve Dataset pages, exports, or Key-value store objects
  • Validate required fields, content types, and expected schema versions
  • Map Actor-specific fields to canonical and target fields
  • Apply deduplication, enrichment, normalization, and business rules

Objective

Deliver validated data or artifacts to enterprise applications and storage systems.

Instructions in Martini

  • Upsert records using stable identifiers or business keys
  • Process large outputs in batches with checkpoints
  • Forward screenshots, HTML, or files with appropriate metadata
  • Avoid writing incomplete or failed Actor output to target systems

Objective

Make execution observable and safe to retry without creating duplicate runs or outputs.

Instructions in Martini

  • Capture API status codes, run status, correlation IDs, and target identifiers
  • Back off on rate-limit and transient server responses
  • Use run IDs, source URLs, or business keys for idempotency
  • Route unrecoverable errors to operational notifications or queues
  • Reconcile missed webhook events and review schema changes before deployment

Common Apify data objects used in integrations

ObjectTypical UseCommon target systemsMartini handling
ActorsReusable scraping, browser automation, crawling, and data-extraction programs.Salesforce, HubSpot, Snowflake, internal APIsMartini submits Actor input, records the Actor identifier, starts runs, and routes output according to the workflow’s business rules.
Actor runsIndividual executions containing status, timing, resource usage, input, and output references.Operational databases, Slack, Jira, monitoring systemsMartini persists run identifiers and correlation data, polls or receives selected notifications, distinguishes terminal states, and applies retry policies.
Actor TasksSaved Actor configurations with predefined input and settings.Scheduled workflows, CRM enrichment processes, reporting pipelinesMartini invokes Tasks through the REST API, supplies runtime parameters where applicable, and processes the resulting run like an Actor execution.
DatasetsStructured, usually tabular output collections produced by Actors.Salesforce, HubSpot, Google Sheets, Snowflake, databasesMartini retrieves pages or exports, validates Actor-specific schemas, transforms fields, removes duplicates, and upserts target data.
Key-value storesJSON documents, text, binary files, screenshots, HTML, and other named Actor outputs.Amazon S3, document repositories, case systems, internal APIsMartini retrieves objects through the API, checks content and metadata, and forwards or transforms the artifact.
Request queuesURLs and crawl requests managed by crawling Actors.Crawling workflows, internal URL inventories, databasesMartini adds, reads, or manages queue items through the API and tracks processing or reconciliation state.

Authentication and security considerations

API-token protection

Apify’s public API primarily uses API tokens, commonly supplied as Bearer credentials or, where appropriate, a token query parameter. Martini should store tokens in secrets or protected environment configuration rather than workflow definitions, URLs, payloads, or logs.

Least-privilege access

Use credentials with access limited to the Actors, runs, Datasets, Key-value stores, Request queues, and Webhooks required by each workflow. Separate development, test, and production credentials where isolation is needed.

Webhook validation

Secure Martini’s receiving API and validate the shared secret, signature, token, or other configured authentication value according to the Apify webhook configuration. Do not process output solely because a callback was received.

Data governance

Extracted pages, screenshots, HTML, and datasets may contain sensitive or regulated information. Assess site permissions, terms of service, privacy obligations, robots directives, copyright, regional restrictions, and retention requirements before storing or distributing results.

Operational considerations for Apify integrations

Rate limits and capacity

Apify request limits and Actor capacity can vary by account plan, resource, and workload. Use bounded backoff for HTTP 429 and transient 5xx responses, avoid unnecessary polling, and account for Actor compute, concurrency, proxy, storage, and platform usage costs.

Asynchronous execution

Do not assume an Actor completes during the submission request. Persist run and output identifiers, handle transitional and terminal statuses explicitly, and use webhooks or scheduled polling according to event coverage.

Pagination and checkpoints

Process large Datasets in pages or batches rather than loading everything into memory. Track the relevant offset, cursor, or last successful item and use stable identifiers when output may change between retrieval operations.

Idempotency and retries

Webhook delivery or workflow replay can cause duplicate processing. Use run IDs, Dataset item identifiers, source URLs, or business keys, prefer downstream upserts, and ensure retries do not unintentionally submit duplicate Actor runs.

Schema and testing

Actor input and output schemas are Actor-specific and may change when an Actor is upgraded or replaced. Validate required fields, test representative success and failure states, and monitor API responses, run status, mapping failures, and downstream writes.

Why use Martini instead of scripts or point-to-point integrations?

Orchestrate complete integration lifecycles

Scripts often combine submission, polling, output retrieval, transformation, and target writes in one fragile process. Martini separates these concerns into reusable workflows and APIs that can support scheduled, event-driven, and on-demand execution.

Manage changing Actor schemas

Apify output varies by Actor. Martini provides explicit mappings, validation, business rules, and transformation steps so Actor-specific data can be converted into governed canonical and target models.

Improve operational recovery

Martini can correlate run identifiers, apply bounded retries, checkpoint batches, handle webhook reconciliation, and route failures for review without exposing credentials or sensitive extracted content in logs.

Expose controlled enterprise APIs

Martini can provide a stable API façade over Apify workloads, applying authentication, authorization, input validation, and business rules while keeping Apify implementation details behind a maintainable integration boundary.

Frequently asked questions

How can Apify be integrated with enterprise systems?

Apify can be integrated through its REST API, which starts Actors and Tasks, monitors asynchronous runs, retrieves Dataset items, accesses Key-value stores, manages Request queues, and configures Webhooks. Martini can orchestrate these calls, process selected webhook notifications, transform Actor-specific output, and deliver results to applications, databases, files, or messaging destinations.

Can Martini integrate with Apify?

Yes. Martini can integrate with Apify by consuming the Apify REST API, storing API tokens as secrets, orchestrating asynchronous Actor runs, retrieving Datasets and Key-value store objects, and receiving supported Apify webhook notifications through a Martini API. No native Martini Apify connector is documented in the supplied research.

Do I need a connector to integrate Apify with Martini?

No dedicated Apify connector is required. Martini can use Apify’s confirmed REST APIs, selected webhook callbacks, Dataset and Key-value store endpoints, and API-token authentication to implement the integration.

Is there any extra Lonti cost to integrate Apify with Martini?

Lonti does not charge an additional per-connector or per-vendor fee to integrate Apify. The integration is subject to the provisioned capacity of the Martini environment. Separate costs may apply from Apify, cloud infrastructure, destination systems, or other third parties based on subscription, usage, and deployment model.

Which Apify integration methods should be used?

The Apify REST API is the primary method for starting Actors and Tasks, monitoring runs, and retrieving output. Use selected Webhooks when the required event is supported and the callback can be secured; use scheduled polling or reconciliation when webhook coverage is unavailable or operational control is required. GraphQL and SOAP were not confirmed for the standard Apify platform API.

Are Apify events or webhooks available?

Apify supports webhook-style notifications for selected platform events and configured Actor run conditions. Coverage is not universal, so the workflow should confirm the required event and retain polling or reconciliation for missed notifications and unsupported scenarios.

How does synchronization between Apify and another system work?

A Martini workflow starts or monitors an Actor, retrieves Dataset pages or Key-value store objects, validates the Actor-specific schema, and maps output to the target model. Incremental synchronization can use prior run identifiers, stable source URLs, business keys, timestamps where reliable, and checkpoints. Upserts and duplicate detection help make reprocessing safe.

How are Apify errors, retries, and duplicate runs handled?

Martini can distinguish API submission errors, transient HTTP failures, Actor execution failures, output validation failures, and downstream write errors. Bounded backoff handles rate limits and transient failures, while run IDs, Dataset references, source URLs, or business keys support idempotency. Retries should avoid unintentionally starting duplicate Actor runs.

Can Martini expose an API façade for Apify?

Yes. Martini can expose a controlled REST API that accepts business-level requests, validates input, starts an Apify Actor or Task, and returns or tracks the run reference. The façade can hide Apify credentials, apply authorization and business rules, and provide a stable contract to consuming applications.