.png)
Databricks Integration Guide
Integrate Databricks with enterprise systems through REST APIs, SQL connectivity, asynchronous Jobs, file exchange, and selected webhook notifications.
Databricks integration options at a glance
Databricks provides REST APIs for workspace administration, Jobs, SQL warehouses, Unity Catalog, files, identity, and serving endpoints. SQL workloads can use JDBC or ODBC, while the SQL Statement Execution API supports synchronous and asynchronous queries. Jobs and long-running statements can be started and monitored through identifiers and polling. Databricks also supports webhook destinations for selected notifications, rather than universal events for every object. File APIs and cloud storage support batch exchange and staging. Martini can consume these APIs, connect through JDBC, receive supported webhook notifications, schedule incremental workflows, transform data, and apply retries and reconciliation rules.
| Integration point | Supported by Databricks? | Common use cases | How Martini supports it |
|---|---|---|---|
| REST APIs | Yes | Manage Jobs, runs, SQL warehouses, Unity Catalog objects, workspace resources, files, identities, permissions, and serving endpoints. | Martini can consume Databricks REST APIs over HTTPS, map JSON responses, paginate results, and orchestrate dependent calls in workflows. |
| Database / analytics access | Yes | Use SQL warehouses, JDBC, or ODBC for set-oriented reads, writes, parameterized SQL, incremental extraction, and staging. | Martini can use configured JDBC connectivity and SQL workflow nodes where the driver, network path, warehouse, and authentication are available. |
| Bulk / async / batch APIs | Limited | Run large-scale processing through Jobs, multi-task workflows, Spark, Delta processing, file ingestion, and asynchronous SQL statement execution. | Martini can submit work, retain run or statement identifiers, poll status, retrieve results, and apply timeout, retry, and compensation logic. |
| Webhooks / outbound callbacks | Limited | Receive selected Jobs-related and product-specific notifications through configured webhook destinations. | Martini can expose a REST API to receive supported notifications, validate them, retrieve current Databricks details, and trigger downstream workflows. |
| File APIs | Limited | Transfer or stage workspace files, DBFS-related files, Unity Catalog volumes, and batch exchange artifacts. | Martini can orchestrate file validation, transfer, metadata capture, and subsequent Databricks processing through APIs or storage workflows. |
| Authentication | Yes | Authenticate with personal access tokens, OAuth 2.0, service principals, and applicable cloud identity controls. | Martini can keep credentials in secure configuration, use supported authentication settings, and separate environment-specific secrets from workflow logic. |
| Scheduled synchronization | Yes | Run incremental SQL extraction, file processing, metadata synchronization, and downstream application updates on a schedule. | Martini scheduler-triggered workflows can maintain watermarks, process pages or partitions, and commit state only after successful target operations. |
| GraphQL APIs | No | Databricks documents REST-based platform APIs rather than an official GraphQL API. | Martini can use REST, JDBC, SQL, files, or supported webhooks instead; a GraphQL mechanism should not be assumed for Databricks. |
How Databricks exposes data and business events
Databricks REST APIs
Databricks REST APIs cover workspace, account, Jobs, SQL, Unity Catalog, files, identity, serving, and other platform operations. They commonly use JSON over HTTPS and require the appropriate workspace or account host.
Martini implementation pattern
Martini consumes the relevant REST definition through an API-oriented workflow, authenticates with a protected token or OAuth configuration, handles pagination and response validation, and maps the result into a canonical model or downstream API.
Implementation sequence
Databricks SQL access
Databricks supports SQL warehouses through JDBC and ODBC, and provides a SQL Statement Execution API for submitting statements. SQL is appropriate for set-oriented reads, writes, staging, and analytical queries.
Martini implementation pattern
Martini uses JDBC where driver and network configuration permit, or consumes the Statement Execution API for API-led SQL operations. It applies parameterization, controls result sizes, and uses incremental predicates or staged files for large transfers.
Implementation sequence
Databricks Jobs and asynchronous execution
Jobs and long-running SQL statements can execute asynchronously. A request may return a run or statement identifier before processing completes, requiring later status checks and result retrieval.
Martini implementation pattern
Martini starts the operation, persists the returned identifier, and uses bounded polling intervals until success, failure, cancellation, or timeout. The workflow can notify an operations system or start a compensating process based on the terminal state.
Implementation sequence
Databricks webhook notifications
Databricks supports webhook destinations for selected Jobs-related and product-specific notifications. Coverage is not universal for every table mutation or platform object.
Martini implementation pattern
A Martini API receives the supported notification, authenticates or validates the request according to the configured contract, and retrieves authoritative Job, run, or alert details from Databricks before triggering downstream processing.
Implementation sequence
Databricks file exchange
Databricks supports file operations for workspace files, DBFS-related locations, and Unity Catalog volumes. External cloud storage can also support staging and batch exchange, depending on architecture.
Martini implementation pattern
Martini validates file metadata and content, transfers or stages the artifact through the agreed endpoint, and starts a Databricks Job or SQL process. The workflow records paths, checksums or identifiers where available, and isolates failed files for replay.
Implementation sequence
Common Databricks integration patterns
Pattern 1: Extract Databricks data into an operational application
When to use this pattern
Use this pattern when Databricks is the analytical source and an application needs selected customer, operational, or finance data without repeatedly transferring full tables. A timestamp, Delta version, Change Data Feed sequence, or other watermark identifies incremental changes.
Integration direction
Example Mapping
| Databricks Field | Canonical Field | Target Field |
|---|---|---|
| customer_id | customerId | account_reference |
| updated_at | lastUpdatedAt | source_updated_time |
| status | lifecycleStatus | record_status |
| segment | customerSegment | u_segment |
Martini implementation pattern
A scheduled Martini workflow queries a SQL warehouse through JDBC or the Statement Execution API, retrieves only rows after the stored watermark, maps and validates the result, and upserts the target objects. It persists the watermark only after successful writes and routes rejected rows to reconciliation.
Martini capabilities used
- scheduled workflows
- JDBC connectivity
- API consumption
- data mapping
- business rules
- error handling
Pattern 2: Trigger and monitor a Databricks Job
When to use this pattern
Use this pattern when an upstream application, file arrival, or completed integration should launch a Databricks data refresh, feature computation, or reprocessing task and receive a reliable completion status.
Integration direction
Example Mapping
| Databricks Field | Canonical Field | Target Field |
|---|---|---|
| job_id | processingJobId | job_id |
| request_id | correlationId | run_name or correlation metadata |
| run_id | executionId | run_id |
| life_cycle_state | executionStatus | workflow_status |
Martini implementation pattern
Martini accepts the request, validates the requested Job and input partition, starts the Databricks Job, and stores the returned run identifier. A polling branch handles success, failure, cancellation, timeout, transient retries, and status publication to the caller or operations platform.
Martini capabilities used
- API workflows
- asynchronous orchestration
- polling
- correlation tracking
- business rules
- retry handling
Pattern 3: Ingest application data into Databricks
When to use this pattern
Use this pattern when application events or extracts need to be validated, transformed, and loaded into Databricks for analytics or batch processing. JDBC, APIs, or staged files can be selected according to volume and network design.
Integration direction
Example Mapping
| Databricks Field | Canonical Field | Target Field |
|---|---|---|
| Id | sourceId | source_id |
| LastModifiedDate | sourceUpdatedAt | source_updated_at |
| AnnualRevenue | annualRevenue | annual_revenue |
| Industry | industry | industry |
Martini implementation pattern
Martini retrieves source data, normalizes dates and types, validates required fields, and writes a staging table or file. It then starts the Databricks processing workflow and records accepted, rejected, and replayable partitions with a batch correlation identifier.
Martini capabilities used
- REST API consumption
- file processing
- JDBC connectivity
- mapping and transformation
- validation
- reconciliation
Pattern 4: Process a selected Databricks notification
When to use this pattern
Use this pattern for supported Job or alert-related webhook notifications where an operational platform must be updated or a downstream workflow must begin. It should not be treated as a universal table-change event design.
Integration direction
Example Mapping
| Databricks Field | Canonical Field | Target Field |
|---|---|---|
| event_type | eventType | notification_type |
| run_id | executionId | u_databricks_run_id |
| result_state | executionStatus | state |
| error_message | failureReason | close_notes or work_notes |
Martini implementation pattern
A Martini API receives the notification, validates its origin and correlation data, retrieves authoritative run details, and applies event-specific routing. It updates ServiceNow or another operations system idempotently and records duplicate notifications separately from processing failures.
Martini capabilities used
- exposed REST APIs
- webhook consumption
- API orchestration
- idempotency rules
- data mapping
- monitoring
Applications commonly integrated with Databricks
Databricks commonly participates in data, analytics, and operational architectures. Martini can orchestrate exchanges with named applications using their APIs, JDBC, files, or other supported interfaces, while keeping authentication, mappings, retries, and workflow state in a maintainable integration layer.
| Application | Scenario | Direction | Martini Pattern |
|---|---|---|---|
| Salesforce | Consolidate customer, account, opportunity, and service data for analytics or publish selected scores and segments back to Salesforce. | Salesforce → Martini → Databricks | Martini consumes Salesforce data, validates and maps it to staging tables or files, and invokes Databricks SQL or Jobs workflows. A reverse flow can retrieve calculated attributes and call Salesforce APIs with upsert and retry handling. |
| ServiceNow | Analyze incidents, requests, configuration data, and operational metrics in Databricks, or publish selected analytics and remediation outputs to ServiceNow. | ServiceNow → Martini → Databricks | A Martini workflow retrieves ServiceNow data through its API, applies an incremental watermark, and loads Databricks through JDBC, SQL, or staged files. Results can be validated and written back to ServiceNow with correlation and error handling. |
| Snowflake | Exchange governed datasets between lakehouse and cloud data warehouse environments or support phased data-platform migration. | Snowflake → Martini → Databricks | Martini coordinates extracts, staged files, or SQL operations between the platforms, applies canonical mappings, and records batch identifiers and reconciliation results. The exact exchange method depends on the deployed network and storage architecture. |
| Amazon S3 | Stage raw files, exports, checkpoints, and datasets used by Databricks workloads. | Amazon S3 → Martini → Databricks | Martini schedules or receives file-processing requests, validates object metadata, transfers or stages files, and starts a Databricks Job or SQL process. Failed partitions are routed for replay without repeating successful work. |
| Azure Data Lake Storage Gen2 | Provide cloud object storage for Databricks data, Delta tables, and file-based exchanges. | Azure Data Lake Storage Gen2 → Martini → Databricks | Martini orchestrates file arrival, validation, metadata capture, and Databricks processing through APIs or supported storage patterns. Workflow state records the source path, batch, schema version, and processing outcome. |
| Tableau | Provide dashboards and governed analytics using Databricks SQL warehouse data. | Databricks → Martini → Tableau | Martini can coordinate dataset refresh requests, publish curated extracts through the agreed data interface, and expose status APIs for downstream operations. Databricks SQL warehouses remain the primary analytical access layer. |
| Microsoft Power BI | Provide governed Databricks datasets and SQL queries for reporting and business intelligence. | Databricks → Martini → Microsoft Power BI | Martini orchestrates refresh or publication workflows around Databricks SQL data, applies business rules to selected outputs, and records completion or failure state for operational monitoring. |
| Jira | Correlate engineering and delivery data with operational and business datasets, or synchronize selected project metrics. | Jira → Martini → Databricks | Martini retrieves Jira data through its APIs, normalizes project and issue attributes, and loads incremental data into Databricks. Calculated metrics can be sent back selectively with duplicate detection and audit tracking. |
How to build a Databricks integration in Martini
Objective
Establish the Databricks workspace or account endpoint and select REST, JDBC, SQL Statement Execution, file, or webhook interaction according to the integration requirement.
Instructions in Martini
- Identify the correct workspace or account host
- Configure OAuth, a service principal, or another confirmed authentication method
- Store tokens and client credentials in Martini secrets or secure configuration
- Confirm network access, warehouse, Unity Catalog, and Jobs permissions
Objective
Select the execution model that matches the business process, such as an API request, supported Databricks notification, scheduled extraction, or file arrival.
Instructions in Martini
- Use an API trigger for request-driven processing
- Expose a Martini API for supported Databricks webhook notifications
- Use a scheduler for watermark-based synchronization
- Use a file or workflow trigger for batch exchange
Objective
Read Databricks data or platform state using the appropriate API or SQL interface while accounting for pagination, asynchronous execution, and result size.
Instructions in Martini
- Call the relevant REST endpoint or submit SQL
- Continue through page tokens when present
- Store Job or statement identifiers for asynchronous operations
- Use incremental predicates, Delta versions, or other approved watermarks
Objective
Coordinate dependent calls, polling, file handling, target writes, and status publication in a reusable Martini workflow.
Instructions in Martini
- Separate submission, monitoring, and completion stages
- Define terminal success, failure, cancellation, and timeout states
- Correlate every request, batch, file, run, or statement
- Route exceptional records to reconciliation or replay handling
Objective
Transform Databricks JSON, SQL rows, metadata, or files into the target model and enforce schema and data-quality rules before writing.
Instructions in Martini
- Map explicit source fields to canonical fields
- Normalize dates, numeric types, nulls, and nested values
- Validate required fields and permitted values
- Handle schema additions, removals, and type changes deliberately
Objective
Persist results in the destination application, database, file location, or operational platform using idempotent or upsert-oriented behavior where possible.
Instructions in Martini
- Use upsert or merge semantics when supported
- Commit synchronization watermarks only after successful target writes
- Capture accepted and rejected counts
- Preserve source identifiers and correlation data for reconciliation
Common Databricks data objects used in integrations
| Object | Typical Use | Common target systems | Martini handling |
|---|---|---|---|
| Jobs | Define and launch notebooks, Python scripts, SQL tasks, pipelines, and other automated processing. | Martini workflows, operational applications, incident platforms, data stores | Martini calls the Jobs API, stores identifiers and correlation data, monitors runs, and routes terminal outcomes. |
| Runs | Represent individual Job executions, including state, timing, task results, and output metadata. | Monitoring platforms, ServiceNow, audit stores, notification systems | Martini polls or retrieves run details, maps status and failure information, and applies bounded retry and alerting rules. |
| SQL warehouses | Provide compute for SQL execution, BI access, and application queries. | Martini JDBC workflows, Tableau, Microsoft Power BI, operational applications | Martini selects the appropriate warehouse, submits SQL through JDBC or APIs, and controls query and concurrency behavior. |
| Notebooks | Contain executable SQL, Python, Scala, or R used in collaborative processing and Jobs tasks. | Jobs, source-control or deployment processes, operational orchestration | Martini can invoke related Jobs or workspace APIs and treat notebook execution as an asynchronous workflow activity. |
| Unity Catalog objects | Organize and govern catalogs, schemas, tables, views, volumes, functions, and external locations. | Data catalogs, governance processes, BI applications, audit stores | Martini consumes metadata APIs, validates permissions and schema mappings, and synchronizes selected metadata or data outputs. |
| Users, service principals, and groups | Represent identities and access-control subjects at account or workspace level. | Identity governance, audit platforms, access reviews | Martini can retrieve or manage supported identity resources through APIs while protecting credentials and applying least-privilege rules. |
Authentication and security considerations
Use controlled service identities
Databricks supports personal access tokens, OAuth 2.0, service principals, and applicable cloud identity controls. For unattended production integrations, prefer a service principal and OAuth where supported rather than embedding long-lived personal credentials.
Protect credentials and scope access
- Store tokens and OAuth client values in Martini secure configuration or secrets management.
- Grant only the required workspace, account, Jobs, SQL warehouse, Unity Catalog, and object permissions.
- Plan for token rotation, expiration, and revocation.
- Use the correct workspace or account host and confirm private-network reachability.
Operational considerations for Databricks integrations
Plan for asynchronous work
Job runs and long SQL statements may complete after the initial request. Persist run or statement identifiers, poll at bounded intervals, and define success, failure, cancellation, and timeout states.
Control throughput and state
- Respect API limits, warehouse capacity, concurrency, and retry-after guidance where provided.
- Implement pagination explicitly and avoid transferring large result sets through one response.
- Use JDBC, staged files, or Databricks processing for large datasets.
- Apply idempotent writes, correlation identifiers, and watermarks to prevent duplicate execution.
- Validate schema changes, nested structures, nullability, and data types before target writes.
- Test private endpoints, firewall rules, SQL permissions, and Unity Catalog privileges independently.
Why use Martini instead of scripts or point-to-point integrations?
Coordinate more than one API call
Databricks integrations often combine authentication, SQL or REST calls, asynchronous polling, file exchange, target updates, and operational notifications. Martini represents this behavior as maintainable workflows rather than isolated scripts.
Make synchronization reliable
- Centralize mappings, transformations, validation, and business rules.
- Persist watermarks, run identifiers, and correlation data for replay and reconciliation.
- Apply consistent retries, timeout handling, duplicate protection, and error routing.
- Expose controlled APIs for requests and supported Databricks notifications.
- Keep environment-specific endpoints and secrets separate from reusable integration logic.
Frequently asked questions
Databricks can integrate through REST APIs, JDBC or ODBC SQL access, the SQL Statement Execution API, Jobs and asynchronous processing, file APIs, cloud-storage staging, and webhook notifications for selected events. The right method depends on whether the requirement is platform administration, set-oriented data exchange, batch processing, or event notification.
Yes. Martini can consume Databricks REST APIs, use supported JDBC connectivity for SQL workloads, orchestrate Jobs and SQL statement polling, process files, and receive supported Databricks webhook notifications through a Martini API.
No dedicated Databricks connector is required. Martini can integrate using Databricks REST APIs, JDBC or SQL access, file mechanisms, supported webhook notifications, and Databricks authentication methods confirmed for the deployment.
Lonti does not charge an additional per-connector or per-vendor fee to integrate Databricks. The integration is subject to the provisioned capacity of the Martini environment. Separate costs may apply from Databricks, cloud infrastructure, networking, drivers, or other third-party systems.
Use REST APIs for Jobs, runs, metadata, workspace, identity, and platform operations. Use JDBC or ODBC for set-oriented SQL access, and the SQL Statement Execution API for API-led SQL execution. Use file exchange for batch staging and webhooks only for supported notification types.
Martini can receive Databricks webhook notifications for selected Jobs-related and product-specific events through an exposed API. Databricks does not provide a universal event stream for every table mutation or platform object, so coverage must be confirmed for the specific feature.
Incremental synchronization can use an update timestamp, Delta table version, Change Data Feed where enabled, a Job timestamp, or another reliable watermark. Martini stores the watermark and advances it only after the corresponding target writes succeed, with explicit handling for updates, late-arriving data, and deletions.
Martini can distinguish authentication, permission, transport, SQL, data-quality, and target-system failures, then apply bounded retries and alerting for transient conditions. Correlation identifiers, deterministic partitions, upserts, and persisted workflow state help prevent duplicate processing when a response is lost or an asynchronous operation is retried.
Related Martini documentation
Databricks APIs
Workflows
Data access
Build a reliable Databricks integration with Martini
Use Martini to connect Databricks APIs, SQL workloads, files, and supported notifications with enterprise applications through secure, observable, and maintainable workflows.