.png)
Deepgram Integration Guide
Connect Deepgram speech-to-text and text-to-speech capabilities to enterprise workflows through REST APIs, WebSocket streaming, asynchronous callbacks, and secure API-key authentication.
Deepgram integration options at a glance
Deepgram provides HTTPS REST APIs for prerecorded transcription, text-to-speech, project administration, API-key management, and usage information. It also supports live transcription through WebSocket sessions and asynchronous prerecorded processing with callback URLs for selected operations. Audio can be submitted as request content or through an accessible media URL, while results are returned as structured JSON or generated audio. Martini can consume these APIs, expose an endpoint for supported callbacks, store credentials as secrets, orchestrate processing workflows, and map transcript or usage data into enterprise applications and databases. WebSocket designs require separate runtime and session-management assessment.
| Integration point | Supported by Deepgram? | Common use cases | How Martini supports it |
|---|---|---|---|
| REST APIs | Yes | Prerecorded speech-to-text, text-to-speech, project administration, API-key management, and usage operations use Deepgram HTTPS APIs. | Martini can consume REST APIs from workflows, add authentication headers, submit media or URLs, and transform JSON responses. |
| WebSocket streaming | Yes | Live transcription sends audio frames during a session and returns interim and final results over a bidirectional WebSocket connection. | Martini-based designs require runtime and workflow assessment for persistent WebSocket handling; REST and callbacks are generally simpler for batch processing. |
| Webhooks / outbound callbacks | Limited | Selected asynchronous prerecorded-audio requests can provide a callback URL for completed processing results. | Martini can expose a REST endpoint, validate and correlate the callback, and route the result through a workflow. |
| Bulk / async / batch APIs | Limited | Prerecorded and asynchronous processing supports batch transcription patterns, subject to endpoint, duration, size, concurrency, and account limits. | Martini can schedule or queue submissions, persist job identifiers, process callbacks, and retry transient failures. |
| File / attachment APIs | Limited | Supported prerecorded operations accept audio as request content or through an accessible media URL; Deepgram is not a general-purpose file repository. | Martini can validate media metadata, send content or URLs, and archive results in an external storage system. |
| Authentication | Yes | Project API keys are the principal authentication method, with permissions or scopes; temporary credentials are available for selected short-lived scenarios. | Martini can store keys and token configuration in secure environment settings and apply them to outbound requests. |
| Database / analytics access | Limited | Project, usage, and operational information can be accessed through documented APIs; direct customer SQL access is not provided. | Martini can retrieve usage data and write normalized results to supported SQL databases or analytics platforms. |
| SDKs | Yes | Deepgram publishes language SDKs, particularly useful for live streaming, audio-device integration, and client-specific session behavior. | Martini can use documented REST and callback interfaces directly, with custom JVM-compatible logic considered for protocol-specific requirements. |
How Deepgram exposes data and business events
Deepgram REST APIs
Deepgram's primary integration surface is a set of HTTPS APIs for prerecorded speech-to-text, text-to-speech, project administration, API-key management, and usage operations. Prerecorded audio can be supplied as request content or through a supported accessible URL.
Martini implementation pattern
Martini implementation pattern: a workflow authenticates with a project API key stored as a secret, validates the input, calls the selected Deepgram endpoint, maps the JSON response, and routes the result to an enterprise application, database, or storage platform.
Implementation sequence
Deepgram asynchronous callbacks
Deepgram supports callback URLs for selected asynchronous prerecorded-audio requests. This is a targeted completion mechanism rather than a universal event stream for projects, usage, credentials, or all platform events.
Martini implementation pattern
Martini implementation pattern: expose a controlled REST API, validate the callback, correlate it with the original Deepgram request or internal job, and process the result idempotently. Callback coverage must be confirmed for the selected endpoint and request mode.
Implementation sequence
Deepgram WebSocket streaming
Deepgram live transcription uses bidirectional WebSocket sessions in which a client sends audio frames and receives interim and final transcription messages. It is distinct from an HTTP webhook or asynchronous REST callback.
Martini implementation pattern
Martini implementation pattern: assess runtime suitability for persistent WebSocket connections, define session and reconnect behavior, and use Martini orchestration around the streaming component for authorization, persistence, enrichment, and downstream delivery.
Implementation sequence
Deepgram media processing
Deepgram accepts audio content or supported media URLs for prerecorded operations and returns structured JSON results or generated audio for text-to-speech. It does not serve as the long-term media repository.
Martini implementation pattern
Martini implementation pattern: coordinate external storage, media validation, Deepgram submission, result processing, and archival in a workflow that keeps source references and processing identifiers together.
Implementation sequence
Common Deepgram integration patterns
Pattern 1: Transcribe customer-service calls to a CRM
When to use this pattern
Use this pattern when recorded calls or media URLs must become searchable CRM activities, case notes, or conversation metadata. It supports asynchronous processing and keeps recording correlation separate from transcript content.
Integration direction
Example Mapping
| Deepgram Field | Canonical Field | Target Field |
|---|---|---|
| request_id | speechJobId | External_Job_Id__c |
| results.channels.alternatives.transcript | transcriptText | Call_Transcript__c |
| results.channels.alternatives.words | wordSegments | Transcript_Metadata__c |
| metadata.duration | audioDuration | Call_Duration__c |
Martini implementation pattern
Martini receives a recording reference, validates access and metadata, submits the audio to Deepgram, and stores the job identifier. A callback workflow validates and correlates the result, maps transcript fields to the CRM activity or case, applies consent and retention rules, and retries transient downstream failures without creating duplicate updates.
Martini capabilities used
- workflows
- API consumption
- API exposure
- data mapping
- business rules
- error handling
- secure environment configuration
Pattern 2: Build an asynchronous media transcription pipeline
When to use this pattern
Use this pattern for recordings arriving in object storage that need batch transcription, durable transcript storage, and operational status tracking. It is suitable when low latency is less important than controlled throughput and replayability.
Integration direction
Example Mapping
| Deepgram Field | Canonical Field | Target Field |
|---|---|---|
| media_url | sourceMediaUrl | SOURCE_URL |
| request_id | deepgramJobId | DEEPGRAM_JOB_ID |
| results.channels.alternatives.transcript | transcriptText | TRANSCRIPT_TEXT |
| results.channels.alternatives.utterances | utteranceRows | UTTERANCES |
Martini implementation pattern
A scheduled or event-driven workflow discovers eligible objects, bounds concurrency, submits supported URLs, and records job state. The callback path validates results, flattens optional transcript structures, writes normalized rows, preserves raw output where required, and uses the job identifier as an idempotency key.
Martini capabilities used
- scheduled workflows
- workflow orchestration
- API consumption
- JSON transformation
- SQL/database integration
- retry handling
- monitoring
Pattern 3: Deliver live contact-center transcription
When to use this pattern
Use this pattern when interim and final speech results are needed for captions, agent guidance, or quality monitoring during a live session. The architecture must support persistent bidirectional WebSocket behavior.
Integration direction
Example Mapping
| Deepgram Field | Canonical Field | Target Field |
|---|---|---|
| channel_index | audioChannel | channel |
| is_final | segmentFinal | final |
| alternatives.transcript | segmentText | transcript |
| words[].speaker | speakerNumber | speaker |
Martini implementation pattern
Assess the Martini runtime and place the WebSocket session component where persistent streaming is supported. Martini can orchestrate authorization, session metadata, final-segment processing, enrichment, downstream API delivery, reconnect handling, and durable storage of completed results while avoiding duplicate interim updates.
Martini capabilities used
- workflow orchestration
- API exposure
- data mapping
- business rules
- custom JVM-compatible logic where required
- error handling
Pattern 4: Generate voice notifications from business workflows
When to use this pattern
Use this pattern when an enterprise application needs generated audio for notifications, accessibility, customer responses, or outbound communications.
Integration direction
Example Mapping
| Deepgram Field | Canonical Field | Target Field |
|---|---|---|
| message_text | speechText | text |
| voice | voiceSelection | model_or_voice |
| audio_format | outputFormat | encoding |
| audio_result | generatedAudio | object_content |
Martini implementation pattern
Martini receives validated text, applies voice and format rules, calls the Deepgram text-to-speech API, and stores or forwards the generated audio. The workflow records source identifiers and delivery status, limits payload sizes, and retries only transient failures.
Martini capabilities used
- API consumption
- data mapping
- business rules
- file handling
- workflow orchestration
- error handling
Applications commonly integrated with Deepgram
Deepgram can be placed between audio sources, speech-processing workflows, and enterprise applications that need transcripts, generated audio, conversation metadata, or usage information. The following are practical integration targets based on Deepgram's role as a speech-processing platform and common enterprise architecture patterns.
| Application | Scenario | Direction | Martini Pattern |
|---|---|---|---|
| Salesforce | Attach transcripts, summaries, speaker information, or extracted conversation data to Accounts, Contacts, Leads, Cases, or Opportunities. | Deepgram → Martini → Salesforce | Martini submits prerecorded audio or receives a supported Deepgram callback, normalizes the transcript, correlates it with the Salesforce object, and updates the relevant record with retry and duplicate protection. |
| ServiceNow | Transcribe service-desk calls or voice interactions and associate the result with Incidents, Cases, or customer-service records. | Deepgram → Martini → ServiceNow | A Martini workflow submits an audio URL, validates the asynchronous result, maps transcript and speaker data to ServiceNow fields, and applies idempotent updates to the target record. |
| Zendesk | Add transcripts and speech-derived metadata to support Tickets and customer interactions. | Deepgram → Martini → Zendesk | Martini retrieves or receives the transcript, applies field and privacy rules, and writes a normalized comment, attachment reference, or custom field to Zendesk. |
| Microsoft Teams | Process meeting or call recordings for searchable transcripts, summaries, and compliance workflows where recording access is available. | Microsoft Teams → Martini → Deepgram | Martini obtains an authorized recording reference, submits it to Deepgram, stores the correlation identifier, and routes the completed transcript to the required repository or collaboration process. |
| Amazon S3 | Read recordings for batch transcription and store transcript artifacts or generated audio in durable object storage. | Amazon S3 → Martini → Deepgram | A scheduled or event-driven Martini workflow reads eligible media metadata, submits an accessible URL to Deepgram, processes the callback, and writes normalized output and processing status back to S3. |
| Google Cloud Storage | Process stored media and return transcripts or generated audio to cloud storage for downstream applications. | Google Cloud Storage → Martini → Deepgram | Martini validates the media object, calls the Deepgram API, transforms the response, and stores the resulting transcript or audio artifact with source and job metadata. |
| Snowflake | Load transcripts, utterances, speaker segments, and usage information for analytics and reporting. | Deepgram → Martini → Snowflake | Martini flattens optional Deepgram JSON structures into relational rows, applies schema and privacy rules, and performs idempotent inserts or merges into Snowflake. |
| Jira | Create or update issues from transcribed incident calls, product feedback, or engineering recordings. | Deepgram → Martini → Jira | A Martini workflow converts transcript content into a validated Jira issue payload, applies routing rules, and records the Deepgram job identifier to prevent duplicate issue creation. |
How to build a Deepgram integration in Martini
Objective
Configure the Deepgram project API key or approved temporary-token flow without exposing credentials in workflow logic or logs.
Instructions in Martini
- Create secure environment configuration for the Deepgram credential
- Use HTTPS for all vendor requests
- Limit project-key permissions where supported
- Keep media and transcript access within approved security boundaries
Objective
Select an event, schedule, API request, callback, or streaming session model that matches the processing requirement.
Instructions in Martini
- Use a workflow trigger for new media or business requests
- Use a scheduler for controlled batch submission
- Expose a Martini API for supported Deepgram callbacks
- Assess WebSocket runtime requirements separately for live transcription
Objective
Acquire audio, text, callback data, or usage information and validate it before processing.
Instructions in Martini
- Validate media URL accessibility or binary content
- Check content type, size, duration, and required options
- Persist Deepgram request identifiers and internal correlation keys
- Treat callbacks as independently retriable inputs
Objective
Coordinate Deepgram calls, asynchronous state, downstream writes, and operational status in a maintainable workflow.
Instructions in Martini
- Call the appropriate Deepgram REST endpoint
- Persist submitted, processing, completed, and failed states
- Bound concurrency for batch jobs
- Separate streaming session management from ordinary callback processing
Objective
Convert Deepgram JSON, transcript structures, usage data, or generated audio metadata into downstream models.
Instructions in Martini
- Map only the optional transcript fields required by the target
- Flatten words, utterances, speakers, and confidence values when relational storage is needed
- Preserve raw responses when auditability or replay is important
- Normalize timestamps, identifiers, and status values
Objective
Apply privacy, consent, retention, routing, validation, and idempotency rules before external writes.
Instructions in Martini
- Use stable job identifiers to prevent duplicate processing
- Quarantine malformed or unexpected callbacks
- Restrict logging of raw audio, API keys, and sensitive transcripts
- Route failures according to recoverable and non-recoverable error types
Common Deepgram data objects used in integrations
| Object | Typical Use | Common target systems | Martini handling |
|---|---|---|---|
| Audio files and audio URLs | Input sources for prerecorded transcription and related speech-processing requests. | Amazon S3, Google Cloud Storage, Microsoft Teams, Salesforce, ServiceNow | Martini validates accessibility, content type, size, and correlation metadata before submitting content or a URL to Deepgram. |
| Transcripts | Structured speech-to-text output containing text and optional words, utterances, confidence, speakers, paragraphs, summaries, topics, and entities. | Salesforce, ServiceNow, Zendesk, Snowflake, Jira | Martini maps required fields, preserves raw JSON when useful, applies privacy rules, and writes idempotently to downstream systems. |
| Live streaming sessions | WebSocket sessions for incremental audio transmission and interim or final transcription results. | Contact-center applications, captioning services, quality-monitoring platforms | Martini designs must account for session lifecycle, ordering, reconnects, partial results, and runtime suitability for persistent bidirectional streaming. |
| Projects | Organizational containers for Deepgram resources, credentials, usage, and configuration. | Administrative databases, reporting platforms, governance workflows | Martini can retrieve project information through management APIs and synchronize selected metadata to governance or reporting systems. |
| API keys | Project-scoped credentials used to authenticate Deepgram API requests. | Martini secrets, security administration systems | Keys remain in secure environment configuration; workflows avoid exposing them in mappings, URLs, logs, or payloads. |
| Usage records | Project-level consumption and request information for monitoring, reporting, and cost allocation. | Snowflake, SQL databases, monitoring and finance reporting systems | Martini retrieves usage data, normalizes endpoint-specific responses, and loads it using scheduled or on-demand workflows. |
Authentication and security considerations
Authentication model
Deepgram primarily uses project API keys supplied in the Authorization header over HTTPS. Temporary access tokens are available for selected short-lived or client-side scenarios and should generally be issued by a protected backend workflow.
Credential protection
- Store API keys and token configuration in Martini secrets or secure environment configuration.
- Do not place credentials in mappings, URLs, source code, payloads, or logs.
- Use project-scoped permissions where supported and rotate keys through controlled administrative processes.
- Protect recordings and transcripts because they may contain personal, financial, health, or confidential business information.
Operational considerations for Deepgram integrations
Reliability and scale
- Confirm Deepgram request, duration, payload, concurrency, and account-plan limits before designing high-volume processing.
- Bound parallel submissions and use backoff for transient HTTP failures.
- Implement endpoint-specific pagination for project, member, key, or usage APIs rather than assuming one pagination model.
- Persist request identifiers and make callback processing idempotent.
Media and schema handling
- Validate MIME type, encoding, sample rate, channels, duration, size, and media URL accessibility before submission.
- Expect optional transcript fields to vary with model and requested features such as diarization, utterances, summaries, topics, and entities.
- Define retention, regional processing, deletion, and replay policies for audio and transcript data.
Streaming and testing
Live transcription requires session lifecycle, audio framing, ordering, interim-versus-final handling, reconnect, and termination logic. Test asynchronous callbacks, duplicate delivery, expired media URLs, rate limits, unsupported formats, partial results, and downstream failures before production deployment.
Why use Martini instead of scripts or point-to-point integrations?
Beyond point-to-point scripts
Martini separates Deepgram access from downstream application logic through reusable workflows, APIs, mappings, and environment configuration. This avoids duplicating authentication, correlation, validation, retries, and transformation logic across individual scripts or applications.
Operational control
- Coordinate scheduled, event-driven, API-led, asynchronous, and callback-based processing.
- Apply consistent idempotency, business rules, privacy controls, and error handling.
- Map nested transcript structures into CRMs, service platforms, databases, storage, and analytics models.
- Expose a controlled API façade when applications should not call Deepgram directly.
- Retain reusable integration assets and monitor workflow execution as requirements evolve.
Frequently asked questions
Deepgram can be integrated through HTTPS REST APIs for prerecorded transcription, text-to-speech, project administration, and usage; WebSocket sessions for live transcription; and callback URLs for selected asynchronous prerecorded-audio operations. Audio may be supplied as content or an accessible URL, with results returned as structured JSON or generated audio.
Yes. Martini can consume Deepgram REST APIs, expose an API endpoint for supported asynchronous callbacks, orchestrate transcription and text-to-speech workflows, map Deepgram JSON, and write results to applications, databases, and storage systems. Live WebSocket processing requires a separate runtime and architecture assessment.
No. A dedicated Deepgram connector is not required. Martini can use Deepgram's confirmed native REST APIs, supported callback URLs, authentication methods, media inputs, and other documented endpoints through workflows and APIs.
Lonti does not charge an additional per-connector or per-vendor fee to integrate Deepgram. The integration uses the provisioned capacity of the Martini environment. Separate costs may apply from Deepgram, cloud infrastructure, storage, databases, or other third-party services based on subscription, usage, and deployment model.
REST APIs are the primary choice for prerecorded transcription, text-to-speech, administration, and usage. Use selected asynchronous callback patterns for long-running prerecorded jobs, and use WebSocket streaming when interim live results are required and the runtime can support persistent bidirectional sessions.
Deepgram supports callback URLs for selected asynchronous prerecorded-audio requests. This is not a universal webhook system for project, credential, usage, or streaming events. Live transcription delivers results over WebSocket sessions instead.
Store the Deepgram request or job identifier together with an internal correlation key, then process the result or callback through a Martini workflow. Map required transcript fields such as text, words, utterances, speakers, confidence, and duration into the target model while preserving raw JSON when auditability or replay is important.
Martini can classify authentication, validation, unsupported-format, inaccessible-URL, rate-limit, payload, callback, and streaming errors. Use bounded backoff for transient failures, avoid retrying permanent validation errors, and use the Deepgram request identifier as an idempotency key to prevent duplicate transcript or downstream updates.
Yes. Martini can expose a controlled REST API that accepts an enterprise request, validates and enriches it, calls Deepgram, and returns or persists the result. It can also expose an endpoint for supported Deepgram asynchronous callbacks, with authentication, correlation, validation, and error handling around the vendor interaction.
Related Martini documentation
Workflows
Data
Connect Deepgram with Martini
Use Martini to orchestrate Deepgram speech-processing workflows, receive supported callbacks, secure API access, transform transcript data, and deliver results to enterprise applications and data platforms.