API reference / Voice and AI
Voices, speech and transcription
Use these routes to pick a voice, turn text into audio, clone or design a voice, and turn a short recording into text.
A voice works anywhere a voice_id is accepted: spoken sends, voice
broadcasts, AI agents and text-to-speech. To send a spoken message, you don't
need to synthesize it first. Pass tts_body and voice_id on the send, as
described in Sending basics.
Routes
| Method | Route | Scope | What it does |
|---|---|---|---|
GET |
/voice/public/voices |
media:read |
List your voices and the catalog |
POST |
/voice/public/tts/synthesize |
voice:send |
Turn text into an audio file |
POST |
/voice/public/voices/clone |
media:write |
Clone a voice from a recording |
POST |
/voice/public/voices/design |
media:write |
Design a voice from a description |
DELETE |
/voice/public/voices/{voice_id} |
media:write |
Delete one of your voices |
POST |
/voice/public/asr/transcribe |
media:read |
Turn a recording into text |
Pick a voice
The list holds your cloned and designed voices, plus the platform catalog.
| Field | Type | Required | Description |
|---|---|---|---|
include_urls |
boolean | No | Include each voice's sample url. Default false. |
include_pro_voices |
boolean | No | Include the catalog. Default true. |
skip |
integer | No | Voices to skip. Default 0. |
limit |
integer | No | Page size. Default 100. |
curl "https://api-v2.dropcowboy.com/voice/public/voices?include_urls=true" \
-H "x-key: $DC_KEY" -H "x-secret: $DC_SECRET"
{
"data": {
"voices": [
{
"voice_id": "7d2a9e4b-1c6f-4b3a-8e5d-2f9c7a1b4e63",
"team_id": "3f6c2a1e-8b4d-4c7a-9e2f-5a1b3c4d6e7f",
"name": "Jordan (cloned)",
"type": null,
"gender": "female",
"status": "ready",
"pro_voice": false,
"url": "https://media.example.com/samples/jordan.wav",
"created_at": 1774041600000,
"failed_at": null
}
],
"total": 1
},
"meta": { "request_id": "c2a6e8f4-3b9d-4c1e-8a7f-5d3b1e9c6a24" }
}
Catalog voices have a null team_id. Designed voices have
type: "designed". data.total counts your account's voices, not the
catalog. Timestamps are epoch milliseconds.
Only voices with status: "ready" can speak, catalog voices included:
status |
Meaning | What to do |
|---|---|---|
ready |
The voice can speak. | Use its voice_id. |
processing |
A clone is still being built. | Poll every few seconds. |
pending_payment |
The voice slot hasn't been bought. | Complete checkout in the dashboard. |
failed |
The clone didn't finish. | Clone again from a cleaner recording. |
Text-to-speech
Turn text into an audio file with any voice the list shows as ready. Use it
to preview a script, to cache audio you play many times, or to produce files
for another system.
| Field | Type | Required | Description |
|---|---|---|---|
voice_id |
string | Yes | One of your voices or a catalog voice. |
text |
string | Yes | The words to speak. |
language |
string | No | A language hint, for example en. |
curl https://api-v2.dropcowboy.com/voice/public/tts/synthesize \
-H "x-key: $DC_KEY" -H "x-secret: $DC_SECRET" \
-H "content-type: application/json" \
-d '{
"voice_id": "7d2a9e4b-1c6f-4b3a-8e5d-2f9c7a1b4e63",
"text": "Thanks for calling Example Dental. How can I help?"
}'
{
"data": {
"audio_url": "https://media.example.com/tts/1b7e3c9a.mp3",
"expires_at": 1774045200000,
"tts_characters": 50,
"content_type": "audio/mpeg"
},
"meta": { "request_id": "c2a6e8f4-3b9d-4c1e-8a7f-5d3b1e9c6a24" }
}
The call is synchronous. audio_url works for about an hour, until
expires_at, so download the file if you need it later. You're billed per
character, and tts_characters reports the billed count.
Clone a voice
Get consent from the person whose voice you clone. Send one of sample_url,
media_id or voice_id.
| Field | Type | Required | Description |
|---|---|---|---|
sample_url |
string | One of three | Public URL of a clean recording. |
media_id |
string | One of three | A recording in your media library. |
ext |
string | No | The extension of the media_id file, for example .mp3. |
voice_id |
string | One of three | Re-clone one of your voices instead of creating a new one. |
name |
string | No | Defaults to "My Voice". |
gender |
string | No | For example female. |
instructions |
string | No | A description of how the voice should sound. |
speed |
number | No | 0.25 to 4. |
curl https://api-v2.dropcowboy.com/voice/public/voices/clone \
-H "x-key: $DC_KEY" -H "x-secret: $DC_SECRET" \
-H "content-type: application/json" \
-d '{
"name": "Jordan",
"sample_url": "https://media.example.com/samples/jordan.wav",
"gender": "female"
}'
The route answers 202 and builds the clone in the background:
{
"data": {
"voice_id": "7d2a9e4b-1c6f-4b3a-8e5d-2f9c7a1b4e63",
"long_job_id": "1d5b9f3e-7a2c-4e8d-b6f1-3c7a9e5d2b84"
},
"meta": { "request_id": "c2a6e8f4-3b9d-4c1e-8a7f-5d3b1e9c6a24" }
}
If the voice slot still needs buying, the response also carries
pending_payment: true. Complete checkout in the dashboard.
Wait for it to finish
Poll GET /voice/public/voices every few seconds and find
your voice_id. Stop when its status is ready or failed.
Design a voice
Describe a voice in words and hear a preview. Previews are free.
| Field | Type | Required | Description |
|---|---|---|---|
instructions |
string | Yes | Up to 500 characters describing the voice. |
text |
string | Yes | The line the preview reads. |
language |
string | No | A language hint, for example en. |
name |
string | No | A name for the voice. |
speed |
number | No | 0.25 to 4. |
curl https://api-v2.dropcowboy.com/voice/public/voices/design \
-H "x-key: $DC_KEY" -H "x-secret: $DC_SECRET" \
-H "content-type: application/json" \
-d '{
"instructions": "Warm, friendly receptionist in her thirties, relaxed pace.",
"text": "Thanks for calling Example Dental. How can I help?"
}'
The route answers 202 with a long_job_id:
{
"data": { "long_job_id": "8a4c2e6f-9b1d-4f3a-a7e5-2c8b6d4f1e90" },
"meta": { "request_id": "c2a6e8f4-3b9d-4c1e-8a7f-5d3b1e9c6a24" }
}
Listen to the preview and save it in the dashboard. A saved design appears in
GET /voice/public/voices with type: "designed".
Delete a voice
Wait until a clone's status is ready or failed before you delete it.
curl -X DELETE https://api-v2.dropcowboy.com/voice/public/voices/7d2a9e4b-1c6f-4b3a-8e5d-2f9c7a1b4e63 \
-H "x-key: $DC_KEY" -H "x-secret: $DC_SECRET"
{
"data": {
"voice_id": "7d2a9e4b-1c6f-4b3a-8e5d-2f9c7a1b4e63",
"deleted": true
},
"meta": { "request_id": "e4b8a2c6-9d3f-4a1e-b7c5-8f2d6a9e3b17" }
}
The voice leaves the list and can't be used in sends or synthesis. If it
counted toward your monthly saved-voice subscription, that count drops by
one. A voice in pending_payment is removed from your checkout cart instead.
You can delete only your own voices, not catalog voices.
Transcription
Turn a short recording into text, for example a voicemail reply or a recorded
message you want to check before a campaign. Send media_id or url.
| Field | Type | Required | Description |
|---|---|---|---|
media_id |
string | One of two | An MP3 or WAV in your media library. |
ext |
string | No | The extension the media_id file was uploaded with: .wav (the default) or .mp3. |
url |
string | One of two | Public URL of a WAV file with PCM 16-bit samples. |
curl https://api-v2.dropcowboy.com/voice/public/asr/transcribe \
-H "x-key: $DC_KEY" -H "x-secret: $DC_SECRET" \
-H "content-type: application/json" \
-d '{ "media_id": "1b7e3c9a-4d2f-4a8b-9e6c-7f2a1d5b3c80" }'
{
"data": {
"text": "Hi, this is Jordan returning your call about the appointment.",
"language": "en"
},
"meta": { "request_id": "c2a6e8f4-3b9d-4c1e-8a7f-5d3b1e9c6a24" }
}
The call is synchronous and times out after 29 seconds. Keep recordings to a few minutes and under 10 MB.
Audio formats
media_idtakes any MP3 or WAV you uploaded as media. An MP3 needs a few seconds after the upload completes before it can be transcribed. Until then it answers409: wait two or three seconds and retry.urlmust point at a WAV file with PCM 16-bit samples. Any other format answers400. For MP3 or other formats, upload the file as media and sendmedia_idinstead.
Errors
The errors specific to these routes are below. For the error format and everything else, see Responses, errors and limits.
Text-to-speech errors
| Status | Cause | What to do |
|---|---|---|
400 |
voice_id or text is missing. |
Send both. |
402 |
No funds, or the voice is pending_payment. |
Add funds, or complete checkout in the dashboard. See Payment required. |
404 |
The voice doesn't exist, was deleted, or belongs to another account. | Pick a voice_id from the list. |
409 |
voice-not-ready: the voice is processing or failed. |
Wait for ready, or pick another voice. |
Clone, design and delete errors
| Status | Cause | What to do |
|---|---|---|
400 |
No sample given, speed outside 0.25 to 4, or instructions over 500 characters (instructions_too_long). On delete, a voice_id that isn't a UUID. |
Fix the field. |
402 |
Billing blocks the clone, for example a voice still awaiting checkout. | Add funds or complete checkout. |
404 |
The media_id doesn't exist or was never uploaded, or the voice_id isn't one of your voices. |
Check the id. Finish the media upload first. |
Transcription errors
| Status | Cause | What to do |
|---|---|---|
400 |
Neither url nor media_id was sent, the url audio isn't WAV, the audio is over 10 MB, or it couldn't be fetched or transcribed. |
Send a supported file under 10 MB, or upload it as media. |
402 |
No funds for speech to text. | Add funds. |
404 |
The media_id doesn't exist or its file was never uploaded. |
Check the id. Finish the upload first. |
409 |
The media_id file is an MP3 still being converted. |
Retry in a few seconds. |
502 |
Speech to text is briefly unavailable. | Retry with backoff. |