.png)
Dataiku Integration Guide
Integrate Dataiku DSS with enterprise systems through REST APIs, selected scenario webhook triggers, asynchronous jobs, and governed Martini workflows.
Dataiku integration options at a glance
Dataiku DSS provides documented REST APIs for working with projects, datasets, recipes, scenarios, jobs, models, deployments, and other platform resources. Selected scenario automations can be initiated through webhook-style HTTP triggers, while many operations run asynchronously and return execution identifiers for later status checks. Dataiku also supports connection-based access to databases and analytical platforms, plus resource-specific file and dataset operations. Martini can consume these APIs, securely manage Dataiku API keys, invoke scenario triggers, poll job status, transform payloads, and expose controlled APIs around Dataiku operations. Exact resource coverage depends on the DSS version, enabled components, and deployment topology.
Common Dataiku integration patterns
Common Dataiku data objects used in integrations
Authentication and security considerations
API keys and service identities
Dataiku DSS commonly uses API keys for programmatic access. Keys may be associated with a user, project, service account, API service, API node, or governed deployment, so the required identity depends on the endpoint and topology.
Least-privilege access
Use a dedicated service identity with only the permissions required for the relevant Projects, Datasets, Recipes, Scenarios, jobs, models, deployments, or API services. Network reachability alone does not establish authorization.
Secret management
- Store Dataiku API keys and endpoint credentials in Martini secure environment configuration.
- Separate development, test, and production credentials.
- Plan for key rotation, expiration, revocation, and auditability.
- Do not place secrets in mappings, URLs, source code, or logged payloads.
Deployment boundaries
Confirm whether the target is a design node, automation node, API node, or another deployment component. Resource availability and permissions can differ across Dataiku deployment topologies.
Operational considerations for Dataiku integrations
Asynchronous execution
Many Dataiku operations start jobs or Scenarios and return before processing completes. Persist execution identifiers, poll deliberately, set timeouts, and distinguish accepted, running, successful, failed, canceled, and abandoned states.
Pagination and payload size
List endpoints may paginate Projects, Datasets, jobs, deployments, or other resources. Large datasets and artifacts should use supported staging, files, pagination, or asynchronous mechanisms rather than oversized synchronous requests.
Retries and idempotency
Coordinate retries with Dataiku job state so a transient HTTP failure does not launch duplicate processing. Use source event IDs, business keys, or execution IDs to prevent duplicate Scenario runs and downstream writes.
Rate and workload control
An API request may be lightweight while the resulting Dataiku job consumes substantial compute and storage. Limit concurrency and avoid repeatedly starting equivalent Scenarios.
Schema and freshness
Projects, Datasets, Recipes, and model outputs can evolve. Use explicit mappings, validate required fields and output freshness, and avoid relying on undocumented fields or column order.
Testing and observability
Test against the intended DSS version and deployment component. Capture correlation IDs, Dataiku execution IDs, response status, and sanitized error details in Martini logs without exposing credentials.
Why use Martini instead of scripts or point-to-point integrations?
Reusable orchestration
Martini provides a workflow layer for receiving events, calling Dataiku APIs, coordinating asynchronous execution, applying business rules, and publishing results to multiple enterprise systems.
Controlled API exposure
Martini can expose a governed business-level API so applications do not need direct access to Dataiku credentials, project identifiers, or internal deployment details.
Explicit transformation
Mappings and validation keep Dataiku Dataset structures, model outputs, and Scenario results separate from target application schemas. This reduces coupling as projects and downstream systems change.
Operational reliability
Centralized retries, timeout handling, duplicate prevention, error routing, logging, and environment-specific configuration are easier to maintain than scattered scripts or point-to-point calls.
Flexible integration boundaries
Martini can consume REST APIs, invoke selected webhook triggers, connect to supported databases or files when appropriate, and add custom JVM-compatible logic when a specialized transformation is required.