Ellipse Gradient for Header

Azure AI Speech Integration Guide

Integrate Azure AI Speech with enterprise systems through REST APIs, Speech SDKs, asynchronous transcription, audio files, and secure authentication.

Azure AI Speech integration options at a glance

Azure AI Speech provides REST APIs and Speech SDKs for speech-to-text, batch transcription, text-to-speech, speech translation, custom Speech models, and related capabilities. Martini can consume the REST APIs, securely manage subscription keys or bearer tokens, submit audio for asynchronous batch processing, poll job status, retrieve result files, and map outputs into business applications or databases. Audio inputs and generated files can be handled through documented storage and file patterns. Selected asynchronous operations may support callback-style notifications, but polling is the broadly applicable approach. Persistent real-time streaming may require a custom JVM-compatible component using an Azure Speech SDK.

Integration pointSupported by Azure AI Speech?Common use casesHow Martini supports it
REST APIsYesSpeech-to-text, text-to-speech, token acquisition, batch transcription, and custom Speech operations use documented REST endpoints.Martini can consume the Azure AI Speech REST APIs, manage request headers and payloads, map responses, and orchestrate multi-step workflows.
Speech SDKsYesSDKs support language-specific recognition, synthesis, and streaming scenarios across supported programming languages.Martini can use REST directly or invoke custom JVM-compatible logic when an SDK-specific capability is required.
Bulk / async / batch APIsYesBatch transcription processes one or more audio files asynchronously and exposes job status and result artifacts.Martini can submit jobs, persist job identifiers, poll status, retrieve result files, and route completed transcripts.
File / attachment APIsLimitedAudio files can be supplied for transcription and batch jobs can produce result files; this is not a general-purpose attachment API.Martini can handle file references and result artifacts, validate media metadata, and forward outputs to storage or downstream applications.
Webhooks / outbound callbacksLimitedCallback or notification behavior is available for selected asynchronous operations and API surfaces, rather than every Speech event.Martini can receive a documented callback, but should use scheduled polling when the selected operation does not provide a reliable notification mechanism.
AuthenticationYesSpeech APIs support subscription keys, short-lived bearer tokens, and supported Microsoft Entra ID or managed identity configurations depending on the operation.Martini can store credentials in secrets or environment configuration and apply the required authentication to outbound API calls.
Streaming protocolsLimitedReal-time recognition and synthesis use streaming connections and Speech SDK transport mechanisms rather than ordinary request-response REST calls.Martini can orchestrate surrounding APIs, while persistent bidirectional streaming may require a custom JVM-compatible component or application service.
Database / analytics accessNot confirmedAzure AI Speech does not expose a documented relational database interface; processed results can be sent to a database by Martini.Martini can map Speech results into SQL operations or downstream analytics ingestion paths.

How Azure AI Speech exposes data and business events

Azure AI Speech REST APIs

Azure AI Speech exposes REST APIs for speech recognition, text-to-speech, token acquisition, batch transcription, and related Speech operations. REST is the most direct integration surface for server-to-server workflows.

Martini implementation pattern

Martini implementation pattern: Martini uses a configured REST-consuming workflow to authenticate, construct the Speech request, submit audio or text, validate the response, and map the result into a canonical model or downstream API.

Implementation sequence

Receive an audio, text, or processing request
Load the regional endpoint and credentials from secure configuration
Call the documented Azure AI Speech REST endpoint
Validate the response and capture Azure correlation information
Map the Speech result into the target application model
Write the result or return the API response

Batch transcription

Batch transcription processes one or more audio files asynchronously. A client submits a job, checks its status, and retrieves result files after completion.

Martini implementation pattern

Martini implementation pattern: Martini submits the batch request, persists the Azure job identifier and source correlation key, then uses a scheduled or queue-oriented workflow to poll status and retrieve results without creating duplicate jobs.

Implementation sequence

Validate the audio reference and processing parameters
Submit the batch transcription job
Persist the Azure job identifier and source recording key
Poll the documented job status
Retrieve result files after successful completion
Map transcript content and update the downstream system

Audio and result files

Azure AI Speech accepts audio input and can return audio output or transcription result files, depending on the selected operation. Durable storage such as Azure Blob Storage may be used where documented by the chosen API.

Martini implementation pattern

Martini implementation pattern: Martini manages file references rather than embedding large binaries in ordinary messages when appropriate, validates access and media properties, and routes result artifacts to approved storage or enterprise applications.

Implementation sequence

Receive or locate the approved media reference
Validate storage access, format, size, and duration
Submit the documented file or reference to Azure AI Speech
Retrieve the generated audio or transcription artifact
Store or forward the artifact with controlled retention
Record the source and result identifiers for reconciliation

Selected callbacks and notifications

Azure Speech supports callback or webhook-style notifications for selected asynchronous operations and API surfaces. This behavior does not apply to every Speech event or every real-time recognition result.

Martini implementation pattern

Martini implementation pattern: Martini can expose an API for a documented callback, validate the notification, and reconcile it with persisted job state. Scheduled polling remains the fallback for batch operations without a suitable callback.

Implementation sequence

Expose a controlled Martini callback API when the Speech operation supports it
Authenticate and validate the notification
Match the notification to the stored Azure job identifier
Retrieve the current job or result state
Process the completed artifact
Mark the job complete or route it for retry

Speech SDK and streaming scenarios

Real-time recognition and synthesis use streaming connections and Speech SDK transport mechanisms. These are materially different from ordinary REST request-response calls.

Martini implementation pattern

Martini implementation pattern: Martini can orchestrate the surrounding request, authentication, storage, and downstream processing, while a custom JVM-compatible component can encapsulate the Azure Speech SDK when persistent bidirectional streaming is required.

Implementation sequence

Determine whether the selected operation requires persistent streaming
Expose a stable workflow-facing interface
Invoke the approved custom JVM-compatible Speech component
Normalize interim and final recognition results
Apply completion and failure rules
Forward final results to downstream systems

Common Azure AI Speech integration patterns

Pattern 1: Transcribe contact-center audio into case records

When to use this pattern

Use this pattern when recordings from a contact center or collaboration platform must be transcribed and attached to customer-service records. The workflow should support asynchronous processing, speaker metadata where available, and reconciliation of retries against the original interaction.

Integration direction
Audio source
Martini
Azure AI Speech
ServiceNow
Example Mapping
Azure AI Speech FieldCanonical FieldTarget Field
sourceInteractionIdinteractionIdcorrelation_id
transcription.phrasestranscriptTextwork_notes
transcription.timestampstranscriptSegmentstranscript_metadata
job.statusprocessingStatusu_speech_status
Martini implementation pattern

A Martini API or scheduled workflow receives an audio reference, validates the media and source identifier, submits a batch transcription job, and stores the job state. A follow-up workflow polls for completion, maps phrases and timestamps, applies rules for missing or low-quality output, and updates ServiceNow. Stable interaction identifiers and Azure job IDs prevent duplicate case updates; transient failures use backoff and retry.

Martini capabilities used
  • APIs
  • workflows
  • API consumption
  • data mapping
  • business rules
  • scheduled execution
  • error handling

Pattern 2: Convert voice intake into structured business requests

When to use this pattern

Use this pattern when an application accepts an audio recording for a service, maintenance, or operational request and downstream systems require structured fields rather than unprocessed speech output.

Integration direction
Voice intake application
Martini
Azure AI Speech
SAP S/4HANA
Example Mapping
Azure AI Speech FieldCanonical FieldTarget Field
audioReferencemediaLocationdocumentReference
transcription.textrequestDescriptionlongText
transcription.languagelocalelanguage
sourceRequestIdrequestIdexternalReference
Martini implementation pattern

Martini receives the request through an exposed API, validates the audio location and required metadata, invokes Azure AI Speech, and applies extraction and validation rules to the transcript. Only requests meeting required-field and confidence policies are sent to SAP S/4HANA; incomplete results are routed for review. The workflow records correlation identifiers and retries transient API failures without resubmitting an already accepted job.

Martini capabilities used
  • API exposure
  • API consumption
  • workflows
  • validation
  • mapping and transformation
  • business rules
  • retry handling

Pattern 3: Generate audio notifications from enterprise text

When to use this pattern

Use this pattern when a business application needs Azure AI Speech to convert approved text into audio for a notification, call-center, or accessibility workflow.

Integration direction
Business application
Martini
Azure AI Speech
Storage or notification application
Example Mapping
Azure AI Speech FieldCanonical FieldTarget Field
textmessageTextinput_text
voicevoiceNamevoice
languagelocalelanguage
outputFormataudioFormatcontent_type
Martini implementation pattern

A Martini API validates text length, locale, voice, and output-format parameters before calling the Speech synthesis endpoint. The workflow returns the audio for smaller payloads or stores it in an approved location for larger outputs, then forwards a controlled reference to the target application. Invalid parameters are rejected before the Azure call and transient failures are retried according to the operation's safety rules.

Martini capabilities used
  • API exposure
  • API consumption
  • validation
  • data mapping
  • workflow orchestration
  • secure configuration
  • error handling

Pattern 4: Manage custom Speech model operations

When to use this pattern

Use this pattern when domain terminology, pronunciation data, or approved training datasets must be registered and incorporated into later transcription workflows. Exact operations and permissions depend on the selected API version and region.

Integration direction
Dataset repository
Martini
Azure AI Speech
Transcription workflows
Example Mapping
Azure AI Speech FieldCanonical FieldTarget Field
datasetReferencetrainingDataLocationdataset_uri
modelNamemodelIdentifiercustom_model_id
operation.statusmodelStatusdeployment_status
approvedLocalelocalespeech_locale
Martini implementation pattern

A scheduled Martini workflow detects approved dataset changes, calls the documented custom Speech endpoints, tracks operation and model identifiers, and records the resulting state. Business rules prevent unapproved models from being selected by production transcription workflows. Authentication failures, invalid datasets, region mismatches, and transient service errors are separated for remediation or retry.

Martini capabilities used
  • scheduled workflows
  • API consumption
  • data mapping
  • business rules
  • state tracking
  • secrets management
  • error handling

Applications commonly integrated with Azure AI Speech

Azure AI Speech can be used alongside recording platforms, customer-service applications, collaboration tools, document repositories, and enterprise business systems. The exact access path depends on the source application, its APIs, tenant configuration, and the selected Speech operation. Martini can coordinate those endpoints, normalize speech results, and route the output to downstream systems.

Application Scenario Direction Martini Pattern
Microsoft Teams Transcribe meeting, call, or customer-interaction audio and route transcripts for retention, search, or follow-up processing. Microsoft Teams → Martini → Azure AI Speech → Microsoft Teams Martini receives an approved recording reference, submits or retrieves audio through the configured Microsoft APIs, starts Azure AI Speech processing, polls asynchronous jobs when required, and writes normalized transcript metadata back to Microsoft 365 or another target.
Salesforce Attach call transcripts to Leads, Contacts, Accounts, Cases, or Opportunities to support sales and service analysis. Salesforce → Martini → Azure AI Speech → Salesforce A Martini workflow accepts an audio reference, validates language and media metadata, calls Azure AI Speech, maps phrases, timestamps, and speaker data, and writes the completed transcript to the appropriate Salesforce object with duplicate protection.
ServiceNow Convert service-call recordings into incident notes, case comments, or searchable knowledge content. ServiceNow → Martini → Azure AI Speech → ServiceNow Martini receives a ServiceNow or contact-center request, submits audio to Azure AI Speech, monitors the batch job, applies validation and routing rules, and updates the related ServiceNow record with transcript content and processing status.
Microsoft Dynamics 365 Enrich customer-service or sales records with transcriptions and extracted action items. Microsoft Dynamics 365 → Martini → Azure AI Speech → Microsoft Dynamics 365 Martini orchestrates the audio reference, Speech request, asynchronous status checks, transcript transformation, and Dynamics 365 update while preserving the source interaction identifier for reconciliation.
SharePoint Store audio and generated transcript files in document libraries with metadata and retention controls. SharePoint → Martini → Azure AI Speech → SharePoint A Martini workflow retrieves or receives a controlled SharePoint file reference, submits the audio to Azure AI Speech, retrieves the result artifact, and writes the transcript and processing metadata to the appropriate library location.
Power BI Provide normalized transcription metrics, speaker statistics, processing KPIs, or downstream classifications for reporting. Azure AI Speech → Martini → Power BI Martini extracts relevant Speech result fields, normalizes timestamps and processing states, stores the data in an approved analytical destination, and exposes or forwards the resulting dataset through the selected Power BI ingestion path.
SAP S/4HANA Use voice-derived data to initiate or enrich service, maintenance, or operational processes. Audio source → Martini → Azure AI Speech → SAP S/4HANA Martini submits approved audio for transcription, applies business rules to extract structured values, validates required fields, and calls the relevant SAP API while recording the source interaction and Azure job identifiers.
Zendesk Add call transcripts to tickets and use extracted text to support triage, categorization, or quality review. Zendesk → Martini → Azure AI Speech → Zendesk Martini receives a ticket-linked audio reference, submits it for Speech processing, transforms the completed transcript, and updates the Zendesk ticket or attachment workflow with retry and duplicate controls.

How to build a Azure AI Speech integration in Martini

Objective

Configure the Azure region, Speech endpoint, API version, and authentication method for each environment without embedding credentials in workflow payloads.

Instructions in Martini

  • Store subscription keys, bearer-token settings, or supported Microsoft Entra ID configuration in Martini secrets or environment configuration.
  • Keep development, testing, and production Speech resources separated where appropriate.
  • Confirm that the selected authentication method is supported by the Speech operation and region.

Objective

Select the trigger that matches the processing model: an exposed API for incoming requests, a schedule for batch status checks, or a documented callback for selected asynchronous operations.

Instructions in Martini

  • Use an API trigger for audio locations, synthesis requests, or external processing requests.
  • Use a scheduler or queue-oriented workflow for batch transcription polling.
  • Use a callback endpoint only when the selected Azure operation explicitly documents notification support.

Objective

Obtain the audio, text, job state, or result artifact required by the Speech operation while controlling payload size and storage access.

Instructions in Martini

  • Validate audio references, storage permissions, format, duration, language, and channel information.
  • Retrieve current job status and result files through documented Azure endpoints.
  • Avoid placing large binary content or sensitive media URLs in ordinary logs and messages.

Objective

Coordinate the Azure request, asynchronous state transitions, result retrieval, and downstream writes as a durable Martini workflow.

Instructions in Martini

  • Persist the source identifier, Azure job identifier, processing state, and correlation information.
  • Separate submission, status polling, result retrieval, and downstream delivery where asynchronous processing requires it.
  • Use conditional routing for succeeded, failed, canceled, partial, and retryable states.

Objective

Convert Azure Speech request and response structures into canonical enterprise models while preserving useful metadata.

Instructions in Martini

  • Map phrases, timestamps, confidence-related information, speakers, channels, locale, and result-file references where available.
  • Tolerate optional speaker or channel metadata because not every operation returns the same fields.
  • Normalize audio output references and transcript segments before sending them to target systems.

Objective

Validate business requirements and control duplicate processing before data is written to downstream applications.

Instructions in Martini

  • Use a stable interaction ID, source recording ID, checksum, or upstream event ID for idempotency.
  • Reject unsupported audio or incomplete structured requests before invoking downstream business APIs.
  • Route low-quality, incomplete, or policy-sensitive results for review rather than silently accepting them.

Common Azure AI Speech data objects used in integrations

ObjectTypical UseCommon target systemsMartini handling
Speech resourceDefines the Azure endpoint, region, credentials, quotas, and billing scope used for Speech operations.Martini configuration, Azure administration, monitoring storesMartini keeps the endpoint, region, API version, and credentials in environment configuration or secrets rather than workflow payloads.
Audio inputProvides the stream, request body, file, or externally accessible media location submitted for recognition or transcription.Azure AI Speech, Azure Blob Storage, SharePoint, Microsoft Teams, contact-center platformsMartini validates format, size, duration, language, channel information, and access permissions before submitting the audio or its documented reference.
TranscriptionContains recognized phrases, timestamps, confidence-related information, and speaker or channel metadata where supported.Salesforce, ServiceNow, Microsoft Dynamics 365, Zendesk, databases, analytics storesMartini maps required fields into a canonical transcript model, tolerates optional speaker metadata, and applies business rules before writing downstream.
Batch transcription jobRepresents an asynchronous request that processes one or more audio files and exposes status and result artifacts.Martini workflow state, SQL databases, queues, operational applicationsMartini stores the job identifier and source correlation key, polls documented status fields, prevents duplicate submissions, and retrieves results after completion.
Speech synthesis request and audio outputCombines text, language, voice, speaking style, and output format to produce synthesized audio.Notification services, call-center platforms, storage systems, enterprise applicationsMartini validates request parameters, calls the synthesis endpoint, and returns or stores the generated audio according to payload-size and retention requirements.
Custom Speech modelRepresents a domain-specific recognition model trained or adapted with supported datasets such as audio, transcripts, or pronunciation data.Azure AI Speech, model registries, transcription workflowsMartini can orchestrate documented dataset and model operations, record model identifiers, and select approved models in later transcription workflows subject to API and region availability.

Authentication and security considerations

Credential protection

Azure AI Speech supports subscription keys, short-lived bearer tokens, and supported Microsoft Entra ID or managed identity configurations depending on the selected operation and deployment. Store these values in Martini secrets or environment configuration rather than workflow payloads, URLs, client applications, or ordinary logs.

Endpoint and resource controls

Keep the Azure region, Speech resource, API version, and authentication configuration aligned. Use separate resources or credentials for development, testing, and production where appropriate, and grant only the permissions required by the selected Speech operation.

Speech data handling

  • Protect audio, transcripts, generated audio, and temporary media URLs as potentially sensitive data.
  • Apply retention, deletion, residency, encryption, and access-control policies appropriate to the content.
  • Use private storage and short-lived access URLs when external file references are required.
  • Avoid exposing credentials or sensitive speech payloads through client-facing responses and logs.

Operational considerations for Azure AI Speech integrations

Quotas and retries

Speech resources can impose quotas and throttling. Limit concurrent submissions and use exponential backoff for transient 429 and 5xx responses. Separate interactive and batch workloads where the Azure architecture and governance model permit it.

Asynchronous state

Batch transcription requires durable job state. Store the source identifier, Azure job ID, status, and result reference, and distinguish running, succeeded, failed, and canceled states. Reconcile existing jobs before retrying submission.

Media and schema validation

  • Validate codec, container, sampling rate, channel layout, size, duration, language, and storage accessibility.
  • Do not assume every transcript includes speaker, channel, confidence, or timestamp metadata.
  • Follow documented pagination fields, continuation links, result manifests, and downloadable artifacts.
  • Keep API versions and regional settings configurable and test changes to response structures.

Monitoring and testing

Record correlation identifiers and Azure request identifiers where available, while excluding sensitive content from logs. Test authentication failures, invalid media, quota responses, inaccessible storage, partial result files, duplicate submissions, and downstream write failures.

Why use Martini instead of scripts or point-to-point integrations?

Centralized orchestration

Martini coordinates Azure AI Speech calls, asynchronous job monitoring, file handling, downstream APIs, databases, and business rules in maintainable workflows rather than scattering behavior across scripts.

Reusable integration assets

Teams can expose a consistent API façade, reuse authentication and transformation logic, and standardize transcript, audio, and job-state models across applications such as customer service, collaboration, and enterprise operations.

Reliability and control

  • Apply validation, idempotency, retries, conditional routing, and durable status tracking.
  • Keep credentials and environment-specific endpoints outside application code.
  • Separate provider-specific Speech behavior from downstream system mappings.
  • Monitor and troubleshoot integration workflows through centralized operational controls.

Frequently asked questions

How can Azure AI Speech be integrated with enterprise systems?

Azure AI Speech can be integrated through its REST APIs, Speech SDKs, asynchronous batch transcription endpoints, audio and result-file workflows, and supported authentication methods. Enterprise workflows can submit audio or text, monitor processing, retrieve transcripts or synthesized audio, and route results to applications, databases, storage, or messaging systems.

Can Martini integrate with Azure AI Speech?

Yes. Martini can consume Azure AI Speech REST APIs, manage subscription-key or token-based authentication, orchestrate batch transcription, poll job status, retrieve results, transform Speech data, and expose APIs for upstream applications. Specialized persistent streaming scenarios may require custom JVM-compatible code using a Speech SDK.

Do I need a connector to integrate Azure AI Speech with Martini?

No. A dedicated Azure AI Speech connector is not required. Martini can integrate using Azure AI Speech's native REST APIs, authentication methods, batch endpoints, documented file or result interfaces, and selected callback mechanisms.

Is there any extra Lonti cost to integrate Azure AI Speech with Martini?

Lonti does not charge an additional per-connector or per-vendor fee to integrate Azure AI Speech with Martini. Integrations are subject to the provisioned capacity of the Martini environment. Separate costs may apply from Microsoft Azure, infrastructure providers, storage services, or other third-party systems based on subscription, usage, and deployment model.

Which Azure AI Speech integration methods should be used?

REST APIs are the primary choice for speech recognition, synthesis, token acquisition, batch transcription, and related operations. Speech SDKs are relevant for language-specific or streaming capabilities, while batch APIs are appropriate for asynchronous file processing. GraphQL and SOAP APIs are not confirmed for Azure AI Speech.

Are webhooks or callbacks available for Azure AI Speech?

Callback or webhook-style notifications are available for selected asynchronous operations and API surfaces, but they should not be assumed for every Speech event or real-time recognition result. Martini can receive a documented callback, while scheduled polling remains the broadly applicable pattern for batch transcription.

How does synchronization and duplicate prevention work?

Batch transcription is synchronized by persisting the source recording identifier, Azure job identifier, processing state, and result reference. Martini can poll documented status fields, follow result manifests or pagination where required, and reconcile retries against stable identifiers so that the same audio is not submitted or written downstream more than once.

Can Martini expose an API façade for Azure AI Speech?

Yes. Martini can expose a controlled REST API that accepts an audio location, text-to-speech request, metadata, or batch request and then invokes Azure AI Speech behind the API. It can validate inputs, hide Speech credentials, apply business rules, standardize responses, and route failures without exposing provider-specific details to every client.