.png)
Replicate Integration Guide
Connect enterprise applications to Replicate’s REST API to submit model predictions, process files, and handle asynchronous results through webhooks or polling.
Replicate integration options at a glance
Replicate provides a REST API over HTTPS for models, model versions, predictions, deployments, collections, and related files. Martini can securely consume the API with a Replicate bearer token stored in protected secrets or environment configuration. Prediction requests can run asynchronously, with Martini either receiving selected prediction lifecycle notifications through a Replicate webhook or polling outstanding predictions on a schedule. File-oriented model inputs and outputs can be routed through Martini and copied to durable storage such as Amazon S3 or Google Cloud Storage. Martini workflows can validate model-specific payloads, correlate prediction IDs with business transactions, transform outputs, and update downstream applications.
| Integration point | Supported by Replicate? | Common use cases | How Martini supports it |
|---|---|---|---|
| REST APIs | Yes | Create and retrieve predictions, list or retrieve models and versions, manage deployments, cancel predictions, and work with related resources. | Martini can consume Replicate’s REST API, map model-specific payloads, persist response identifiers, and expose a controlled API façade for calling applications. |
| Webhooks / outbound callbacks | Limited | Receive selected prediction lifecycle notifications, including progress, completion, failure, or cancellation depending on configuration and API behavior. | Martini can expose an authenticated REST endpoint, validate callback data, correlate the prediction ID, enforce idempotency, and continue downstream processing. |
| Asynchronous predictions and polling | Yes | Submit predictions that complete later, then poll prediction resources or reconcile outstanding work when callbacks are delayed or missed. | Martini workflows can persist prediction state, use scheduled polling, apply controlled intervals, and route terminal failures or timeouts for review. |
| File and media inputs | Limited | Submit images, audio, video, documents, or other model-specific file inputs using supported URLs or documented file-input formats. | Martini can retrieve or receive files, construct the required model input, transform metadata, and route generated outputs to durable storage. |
| Authentication | Yes | Authenticate API requests with a Replicate bearer API token over HTTPS. | Martini can keep the token in secrets or protected environment configuration and attach it to outbound requests without exposing it to clients. |
| SDKs | Yes | Replicate provides client libraries for common environments including Python and JavaScript/TypeScript. | Martini can use the REST API directly; custom JVM-compatible logic can be used when specialized request or file handling is required. |
| GraphQL APIs | Not confirmed | No official Replicate GraphQL API was identified in the supplied research. | Martini should use the documented REST API rather than assume GraphQL availability. |
| SOAP APIs | Not confirmed | No official Replicate SOAP API was identified in the supplied research. | Martini should use REST over HTTPS for Replicate integration. |
How Replicate exposes data and business events
Replicate REST APIs
Replicate’s primary integration mechanism is a REST API over HTTPS. It supports models, model versions, predictions, deployments, files, and related resources, including creation, retrieval, cancellation, and status operations.
Martini implementation pattern
Martini implementation pattern: a Martini API or workflow receives a business request, validates the selected model and version, builds the documented JSON payload, adds the bearer token from protected configuration, calls Replicate, and persists the prediction ID and correlation data for later processing.
Implementation sequence
Replicate prediction webhooks
Replicate supports webhook-style notifications for selected prediction lifecycle events. These callbacks are focused on prediction progress and terminal states rather than being a universal event stream for every Replicate resource.
Martini implementation pattern
Martini implementation pattern: expose an authenticated REST API endpoint, validate the callback and any supported signing information, correlate the prediction ID with stored state, reject duplicate or malformed deliveries, and continue output processing only when the status and payload are acceptable.
Implementation sequence
Asynchronous predictions and polling
Replicate predictions commonly complete asynchronously. Clients can poll the prediction resource, and scheduled reconciliation is useful when callbacks are delayed, duplicated, or missed.
Martini implementation pattern
Martini implementation pattern: persist incomplete prediction state, use a scheduled workflow to retrieve outstanding predictions at a controlled interval, update terminal results, and apply retry or exception rules for transient and permanent failures.
Implementation sequence
Replicate file inputs and outputs
Replicate models can accept file-oriented inputs such as images, audio, video, or documents, and may return URLs or lists of generated files. The exact representation depends on the selected model schema.
Martini implementation pattern
Martini implementation pattern: obtain the source artifact, construct the model-specific input representation, submit the prediction, retrieve output files after completion, and copy business-critical artifacts to durable storage instead of relying indefinitely on hosted output URLs.
Implementation sequence
Common Replicate integration patterns
Pattern 1: Process documents or images asynchronously
When to use this pattern
Use this pattern when an application needs extraction, classification, generation, or transformation of documents and images without blocking the initiating business transaction. It accommodates variable model execution times and durable artifact retention.
Integration direction
Example Mapping
| Replicate Field | Canonical Field | Target Field |
|---|---|---|
| input | sourceArtifactReference | model input file |
| model | modelReference | owner/model-name |
| version | modelVersion | version ID or digest |
| output | processedArtifactReference | durable object URL |
Martini implementation pattern
Martini receives the source request, validates the model-specific file input, submits a prediction, and stores the prediction ID with a business correlation key. A webhook or scheduled polling workflow processes the terminal result, copies important output to durable storage, maps extracted values, and routes failed or low-confidence results for review. Duplicate callbacks are ignored using prediction state.
Martini capabilities used
- API creation
- API consumption
- workflows
- webhook consumption
- scheduled workflows
- data mapping
- validation
- secrets management
- error handling
Pattern 2: Enrich customer-support cases with model results
When to use this pattern
Use this pattern when support content must be classified, summarized, or enriched before routing or reporting. It separates model execution from the originating case update and permits manual review for failures or low-confidence outcomes.
Integration direction
Example Mapping
| Replicate Field | Canonical Field | Target Field |
|---|---|---|
| input | caseText | prediction input |
| prediction.status | processingStatus | case processing status |
| prediction.output | classificationResult | category, priority, or summary |
| prediction.id | externalExecutionId | integration audit field |
Martini implementation pattern
Martini receives eligible ServiceNow content, selects and records a tested model version, submits the prediction, and returns or stores an integration status. A callback or reconciliation workflow validates the result, applies confidence and routing rules, updates ServiceNow, and sends failed or incomplete work to an exception path with controlled retries.
Martini capabilities used
- REST API consumption
- API façade
- workflow orchestration
- business rules
- data mapping
- correlation
- idempotency
- error routing
Pattern 3: Generate and publish media assets
When to use this pattern
Use this pattern when prompts or product content are submitted to an image, video, audio, or other media model and the generated artifact must be published to a business application.
Integration direction
Example Mapping
| Replicate Field | Canonical Field | Target Field |
|---|---|---|
| prompt | generationPrompt | prediction input prompt |
| parameters | generationOptions | model-specific input fields |
| output | generatedAsset | durable object URL |
| prediction.id | generationJobId | asset processing reference |
Martini implementation pattern
Martini receives the prompt and generation parameters, validates them against the selected model version, submits the asynchronous prediction, and handles completion through a webhook or polling. It copies the output to Amazon S3, applies file and content rules, and updates the originating product or collaboration application only after durable storage succeeds.
Martini capabilities used
- REST API consumption
- file handling
- asynchronous workflows
- webhook consumption
- mapping and transformation
- business rules
- retry handling
Pattern 4: Reconcile scheduled prediction workloads
When to use this pattern
Use this pattern for larger or recurring workloads where the organization must control submission concurrency, persist every prediction ID, and recover from missed callbacks or transient API failures.
Integration direction
Example Mapping
| Replicate Field | Canonical Field | Target Field |
|---|---|---|
| sourceRowId | businessCorrelationId | prediction tracking key |
| model/version | executionConfiguration | prediction request |
| status | predictionStatus | processing table status |
| error | processingError | exception details |
Martini implementation pattern
A scheduler selects eligible Snowflake rows, submits controlled batches of predictions, and persists identifiers and retry counts. Later workflow runs poll incomplete predictions and reconcile webhook results, normalize output variants, write terminal results back to Snowflake, and isolate rate-limit, invalid-input, and prediction failures for retry or manual action.
Martini capabilities used
- scheduler triggers
- workflow orchestration
- REST API consumption
- SQL/database integration
- controlled concurrency
- mapping
- reconciliation
- error handling
Applications commonly integrated with Replicate
Replicate is commonly used as a model-execution layer within broader enterprise workflows. Martini can orchestrate requests, callbacks, file handling, correlation, and downstream updates between Replicate and named business applications or storage platforms.
| Application | Scenario | Direction | Martini Pattern |
|---|---|---|---|
| Amazon S3 | Store source files and copy generated images, video, audio, or documents from temporary model output locations into durable object storage. | Amazon S3 → Martini → Replicate | A Martini workflow retrieves or receives an input artifact, submits the model-specific file reference to Replicate, receives or polls the prediction result, and copies important output files to Amazon S3 with prediction metadata and a business correlation ID. |
| Google Cloud Storage | Retain generated media and input artifacts while Replicate performs model execution. | Google Cloud Storage → Martini → Replicate | Martini obtains the input object, invokes the Replicate REST API, handles asynchronous completion, and writes output files and durable metadata to Google Cloud Storage. |
| Salesforce | Classify case text, summarize customer interactions, generate content, or enrich Salesforce records with model-derived information. | Salesforce → Martini → Replicate → Salesforce | Martini receives a Salesforce business event or API request, builds the selected model and version payload, submits a prediction, then maps the completed output back to Salesforce with validation and exception routing. |
| ServiceNow | Classify incidents, summarize tickets, extract structured fields, or recommend routing and priority. | ServiceNow → Martini → Replicate → ServiceNow | A Martini workflow receives eligible ServiceNow content, submits it to Replicate, correlates the prediction webhook or polling result, applies confidence and business rules, and updates the originating ServiceNow record or an exception queue. |
| Slack | Submit prompts or content for processing and post generated summaries, classifications, or media back to channels. | Slack → Martini → Replicate → Slack | Martini accepts a Slack-originated request, invokes Replicate asynchronously, stores the prediction correlation, and posts a completion or failure message after validating and transforming the output. |
| Shopify | Generate or enrich product imagery and descriptions, classify catalog content, or process product media. | Shopify → Martini → Replicate → Shopify | Martini retrieves selected Shopify product data or media, submits model-specific inputs to Replicate, stores generated files durably, and updates Shopify only after output validation and duplicate checks. |
| Snowflake | Send selected data or unstructured content for model processing and store prediction results for analysis. | Snowflake → Martini → Replicate → Snowflake | A scheduled Martini workflow reads eligible Snowflake data, submits controlled prediction workloads, persists prediction IDs and statuses, and writes normalized outputs and processing metadata back to Snowflake. |
How to build a Replicate integration in Martini
Objective
Establish the Replicate API connection without exposing the bearer token to calling applications or logs.
Instructions in Martini
- Store the Replicate API token in Martini secrets or protected environment configuration.
- Configure HTTPS requests to Replicate’s REST base URL.
- Keep client-facing API authentication separate from the Replicate credential.
Objective
Select the business or operational event that starts prediction processing.
Instructions in Martini
- Use a Martini API for synchronous intake from an application.
- Use a webhook endpoint for inbound Replicate prediction callbacks.
- Use a scheduler for polling, reconciliation, or recurring workloads.
Objective
Construct and send a valid model-specific prediction request.
Instructions in Martini
- Validate the owner/model reference and tested version where applicable.
- Map prompts, parameters, and file references to the selected model schema.
- Persist the prediction ID, model, version, timestamp, and business correlation ID.
Objective
Handle asynchronous model execution without assuming that creation means completion.
Instructions in Martini
- Receive selected prediction lifecycle callbacks or poll incomplete predictions.
- Apply controlled polling intervals and concurrency limits.
- Treat duplicate callbacks as safe, repeatable processing attempts.
Objective
Convert model-specific outputs into a stable internal or target representation.
Instructions in Martini
- Branch mappings for text, JSON-like output, URLs, and lists of file URLs.
- Validate required output fields and business confidence rules.
- Copy important output files to durable storage before publishing references.
Objective
Decide whether a result can update downstream systems or needs review.
Instructions in Martini
- Route invalid, canceled, failed, or low-confidence predictions to an exception path.
- Prevent duplicate downstream updates using prediction and business correlation identifiers.
- Apply target-specific enrichment and authorization rules.
Common Replicate data objects used in integrations
| Object | Typical Use | Common target systems | Martini handling |
|---|---|---|---|
| Models | Identify the machine-learning model available for a prediction and select its owner and name. | Salesforce, ServiceNow, Shopify, Snowflake, Amazon S3 | Martini validates the model reference, applies routing rules, and stores the selected model alongside the business correlation ID. |
| Model versions | Pin a prediction to a tested version or digest with a known input schema. | Snowflake, ServiceNow, Salesforce, operational databases | Martini records the version used, validates required inputs, and supports version-specific mappings and regression testing. |
| Predictions | Represent asynchronous or synchronous model executions, including inputs, status, output, errors, timestamps, and identifiers. | Salesforce, ServiceNow, Snowflake, Slack, Shopify | Martini submits and correlates predictions, processes webhooks or polling results, applies idempotency, and routes failures or retries. |
| Deployments | Represent dedicated model deployments with configurable infrastructure and scaling behavior. | Operational databases, Snowflake, monitoring systems | Martini can call deployment-related REST resources and map deployment identifiers or status into operational workflows. |
| Collections | Represent curated groups of models available through Replicate. | Internal model catalogs, databases, administration applications | Martini can retrieve collection information through supported API operations where required and normalize it for internal catalogs. |
| Files | Carry image, audio, video, document, or other generated artifacts used as prediction inputs or outputs. | Amazon S3, Google Cloud Storage, Salesforce, Shopify, Slack | Martini passes model-specific file references, retrieves important outputs, copies them to durable storage, and stores durable URLs with prediction metadata. |
Authentication and security considerations
Bearer token authentication
Replicate API requests use a bearer API token over HTTPS. Store the token in Martini secrets or protected environment configuration rather than exposing it to client applications.
Protected callback processing
Protect Martini endpoints that receive Replicate callbacks with authentication and request validation. Where configured and supported, verify Replicate’s documented webhook signing information.
Data protection
- Do not log Replicate tokens or sensitive model inputs.
- Limit access to prediction data, output files, and durable storage locations.
- Keep client authentication separate from the credential used for outbound Replicate requests.
Operational considerations for Replicate integrations
Asynchronous execution
Persist prediction IDs, model versions, business correlation IDs, statuses, timestamps, and retry state. Use webhooks where practical and scheduled reconciliation for missed or delayed callbacks.
Rate limits and retries
Control submission concurrency, avoid overly frequent polling, and use exponential backoff for retryable throttling and transient failures. Separate permanent validation errors from recoverable transport or service errors.
Idempotency and duplicates
Webhook delivery can occur more than once. Use the prediction ID with an application-level correlation key to prevent duplicate downstream updates.
Schema and retention
Model input and output schemas can differ by model and version. Validate each payload and pin tested versions where appropriate. Copy business-critical output files to durable storage because hosted output URLs may have limited retention.
Pagination and observability
Follow Replicate pagination fields or continuation URLs for list operations. Record model and version, processing duration, status, and error details while excluding tokens and sensitive content from logs.
Why use Martini instead of scripts or point-to-point integrations?
Orchestrate the complete lifecycle
Scripts often handle submission but leave callbacks, polling, reconciliation, file retention, and downstream updates fragmented. Martini coordinates these stages in maintainable workflows.
Separate vendor and business models
Martini maps model-specific inputs and outputs into stable enterprise representations, applies validation and business rules, and supports different models without spreading Replicate-specific logic across applications.
Improve operational reliability
Workflows can persist correlation state, handle duplicate callbacks, apply retry and exception paths, and provide monitoring context for asynchronous predictions.
Expose controlled interfaces
Martini can provide an authenticated API façade so applications do not need direct access to Replicate tokens or model-specific endpoint details.
Frequently asked questions
Replicate integrates through a REST API over HTTPS. Enterprise systems can submit model predictions, retrieve models and versions, check asynchronous prediction status, manage supported resources, and process file-oriented inputs and outputs. Selected prediction lifecycle notifications can be delivered through Replicate webhooks, with polling used for reconciliation.
Yes. Martini can consume Replicate’s REST API, submit and track predictions, receive selected Replicate webhook callbacks through a Martini API, poll incomplete predictions, handle file outputs, and map results into applications or databases.
No. A dedicated Replicate connector is not required. Martini can integrate using Replicate’s confirmed native mechanisms: REST over HTTPS, bearer-token authentication, prediction webhooks, asynchronous polling, and model-specific file inputs and outputs.
Lonti does not charge an additional per-connector or per-vendor fee to integrate Replicate. The integration is subject to the provisioned capacity of the Martini environment. Separate costs may apply from Replicate, storage providers, cloud infrastructure, or other third-party services based on their pricing, usage, and deployment models.
Use Replicate’s REST API as the primary mechanism for models, versions, predictions, deployments, and related resources. Use webhooks for selected prediction lifecycle notifications and retain scheduled polling or reconciliation for missed callbacks and operational recovery. No official Replicate GraphQL or SOAP API was identified.
Replicate supports webhook-style notifications for selected prediction lifecycle events such as progress, completion, failure, or cancellation depending on configuration and API behavior. They are not a general event stream for every model, deployment, collection, or account event, so handlers should be idempotent and supported by reconciliation.
A Martini workflow submits a prediction, stores its prediction ID and business correlation ID, and then receives a callback or polls the prediction resource until a terminal state. It maps the output to the target system, copies important files to durable storage, and records status, timestamps, errors, and retry state.
Martini can transform model-specific JSON, text, URLs, file lists, and metadata into target schemas, while validation and business rules determine whether results are accepted. Workflows can use controlled retries and error routing for rate limits, transient failures, output-download errors, invalid inputs, and failed predictions. Prediction IDs and correlation keys support duplicate prevention.
Yes. Martini can expose an authenticated API that accepts a stable enterprise request, validates inputs, invokes Replicate’s REST API, and returns or stores an internal processing status. This keeps Replicate tokens and model-specific implementation details behind a controlled enterprise interface.
Related Martini documentation
Replicate APIs
Workflows
Files and security
Build reliable Replicate integrations with Martini
Use Martini to connect enterprise applications with Replicate’s model APIs, coordinate asynchronous predictions, secure credentials, retain important outputs, and manage downstream automation in maintainable workflows.