TTS and Voice Transcription
Speech Services connect the platform to Text-to-Speech (TTS) and Voice Transcription engines so it can turn text into spoken audio and spoken audio into text. TTS lets you generate professional voice prompts for IVRs and announcements without hiring a voice artist and re-record them instantly when wording changes; Voice Transcription converts voicemail and call recordings into searchable, readable text so users can triage messages by reading and supervisors can review calls without listening end to end. This screen is where an administrator provisions those provider connections (Google, AWS, Custom, or Thirdlane); once configured, the capabilities are enabled per tenant and per user.
In Configuration Manager, open General Settings & Tools > Services > Speech Services to configure Text-to-Speech and Voice Transcription providers for voicemail and call recordings, and for recording Voice Prompts used in IVR and elsewhere. While we describe the configuration steps here, please refer to your provider documentation for more detail.
To enable TTS and Voice Transcription for voicemail and recorded calls, configure the related services in your cloud provider (Google or AWS) or use Custom or Thirdlane providers as applicable, then assign services in Configuration Manager. Once configuration is in place, you can enable Voice Transcription at Tenant and User Extension level.
Custom provider (Text-to-Speech and Transcription)
You can define services whose Provider is Custom for either Text-to-Speech or Transcription.
- Set Purpose to Text-to-Speech or Transcription, then set Provider to Custom.
- Choose a script from the dropdown. Only executable files under the following directories are allowed:
- Text-to-Speech:
/usr/local/share/thirdlane/service/tts - Transcription:
/usr/local/share/thirdlane/service/transcribe
The form stores the full path; the list shows the file name only.
- Text-to-Speech:
- Optionally add Environment variables as grid rows. Each row is stored as one
KEY=valueline (newline-separated in the database). Duplicate variable names are not allowed. Empty values are allowed.
At runtime, custom environment entries are applied when the script runs. Custom Text-to-Speech uses the same generic CLI contract as other CLI-based TTS integrations (Polly-style arguments and JSON on stdout).
Custom Transcription runs your script once per recording as your-script --action transcribe --file <audio path> --key <job key> and takes whatever the script prints on standard output as the transcript. The full command line, what the script must return, and how to test and troubleshoot it are documented in Custom Transcription Script.
For Amazon and Google cloud setup, see the sections below.
AWS TTS and Voice Transcription setup
Setup on the AWS web site
The platform signs in to AWS as an IAM user - an identity a program logs in as, rather than a person - and needs its access key. Create one dedicated to speech services rather than reusing an existing login.
- Sign up for an AWS account at portal.aws.amazon.com/billing/signup if you do not have one.
- Sign in and open the IAM console.
- Select Users in the left pane and click Add user.
- Select the credential type Access key - Programmatic access, then click Next: Permissions.
- Click Attach existing policies directly and select
AmazonTranscribeFullAccess,AmazonPollyFullAccessandAmazonS3FullAccess. Amazon Transcribe reads audio from an S3 bucket rather than from an upload, which is why S3 access is needed as well. - Click Next: Tags, then Next: Review, then Create user.
- Save the Access key ID and Secret access key. AWS shows the secret only once.
AWS setup in Configuration Manager
Add a service in Speech Services and fill in:
- Description - short description of this service.
- Region - the AWS region used for Polly and Transcribe, for example
us-east-1. - Access Key ID - the Access Key ID saved in the previous step.
- Secret Access Key - the Secret access key saved in the previous step.
- 1st Language - the first language considered when transcribing voice to text.
- 2nd Language, 3rd Language, 4th Language - additional languages to consider. Leave them empty if your callers speak only one language.
Google TTS and Voice Transcription setup
Setup on the Google web site
The platform authenticates to Google as a service account - an identity a program logs in as, rather than a person - using a JSON key file you download once and paste into Configuration Manager.
- Sign in to your Google account at accounts.google.com, or create one.
- Open the Google Cloud Console and create a new project. The examples here use the name
thirdlane-transcribe. Google derives a project ID from the name and requires that ID to be unique across all of Google Cloud, so you may be offered a version with digits appended. - Open APIs & Services, select Library, and enable both the Text-to-Speech API and the Speech-to-Text API.
- Select Credentials in the left pane, click Create credentials, and choose Service account from the drop-down.
- Click Create and Continue and grant the service account the roles Google documents for the Text-to-Speech and Speech-to-Text APIs, then click Done.
- Click the service account you just created and open its Keys tab.
- Click Add Key, select Create new key, choose JSON as the key type, and click Create. The key file downloads to your computer.
Keep the JSON file - you paste its contents into Key File in the next section.
Google setup in Configuration Manager
Add a service in Speech Services and fill in:
- Description - short description of this service.
- Key File - the full contents of the JSON key file downloaded from Google, pasted in as text.
- 1st Language - the first language considered when transcribing voice to text.
- 2nd Language, 3rd Language, 4th Language - additional languages to consider. Leave them empty if your callers speak only one language.
Enabling Recorded Calls Transcription
Enabling Recorded Calls Transcription on a Tenant level
Recorded Calls Transcription can be enabled or disabled on a Tenant level.
Allow Recorded Calls Transcription. Specify whether Recorded Calls Transcription will be available for this Tenant.
Enable Transcription by default? Specify whether Recorded Calls Transcription will be enabled by default when creating User Extensions for this Tenant.
Enabling Recorded Calls Transcription for User Extension
If Recorded Calls Transcription is enabled for Tenant, you can enable it for User Extensions for that Tenant.
Transcribe to Text. Specify whether Recorded Calls Transcription will be enabled.
Worked example: Google TTS for IVR greetings + transcription for one team
Goal: build IVR prompts by typing text instead of recording audio, and let the support team read voicemail/recording transcripts - piloted on one tenant first.
- Create the Google credentials. In the GCP Console, create a project, enable the Text-to-Speech and Speech-to-Text APIs, create a service account, and download its JSON key (see the Google steps above).
- Add the service in Configuration Manager. Open Speech Services, add a service, paste the JSON into Key File, set a Description, and set 1st Language (and more if your callers are multilingual). Save.
- Generate an IVR greeting with TTS. Go to a tenant’s Voice Prompts, create a prompt, and choose the TTS option - type “Thanks for calling Acme. Press 1 for Sales, 2 for Support.” and generate. Re-typing and regenerating is how you tweak wording later, with no studio session. Use that prompt as the announcement on your IVR.
- Turn on transcription for the pilot tenant. On the Tenant, set Allow Recorded Calls Transcription = yes. Leave Enable Transcription by default off so you can opt users in one by one.
- Enable it for the team. On each support User Extension, set Transcribe to Text = yes. Their voicemail and recordings now come with readable transcripts.
- Cap the spend. Transcription bills per audio minute, so set the daily/monthly minute limits and maximum audio length in Default Values > Call Recording before rolling out more widely.
Once the pilot looks good, flip Enable Transcription by default on at the tenant so newly created users get it automatically.
Best practices
- Grant least privilege on cloud accounts. Create a dedicated IAM user (AWS) or service account (Google) scoped to only the speech APIs you use, rather than reusing broad admin credentials. Store the keys only in Configuration Manager.
- Match transcription languages to your callers. Configure the language slots (1st through 4th) for the languages your users actually speak; adding languages you never receive can reduce accuracy and add cost.
- Set transcription limits to control spend. Transcription is billed by audio minute at the provider. Use the tenant-level daily/monthly limits and maximum audio length (see Default Values) to cap usage.
- Enable transcription per tenant first, then per user, so you can pilot with a small group before a broad rollout.
- Keep the “convert files to supported format” option in mind if you store compressed recordings - it lets transcription work at the cost of extra server load.
See also
- Speech Services overview — summary of the Speech Services section.
- Custom Transcription Script — the command-line contract for transcribing with your own speech-to-text engine.
- AI Services — Chat AI for Connect (composer rewrite, summarize, translate) and Recording AI for post-call analysis (summary, sentiment, categorization, etc.).
- Default Values — system-wide transcription defaults and limits.