Building Blocks / Voice
Speech-to-Text (ASR)
Turn a recording into a transcript.
This block is the Speech-to-text API. The full product reference is the Speech-to-text API reference. This page is the short version: when to use it, one working request, and its limits. There is no widget to embed.
When to use this vs REST
This block is the REST API. Call it from your server with your API key. There is no browser widget.
Do not rebuild
- Speech recognition or audio format conversion.
- Transcript storage. Save the text on your own record.
Drop-in
Pass a media_id from Recording Studio or an upload, or a url. Store the transcript on your own record. There is no browser widget for this block.
curl -X POST https://api-v2.dropcowboy.com/voice/public/asr/transcribe \
-H "x-key: $KEY" -H "x-secret: $SECRET" \
-H "Content-Type: application/json" \
-d '{"media_id":"5a4b3c2d-1e0f-4a9b-8c7d-6e5f4a3b2c1d"}'
From Node on your server:
async function transcribe(mediaId) {
const transcript = await fetch('https://api-v2.dropcowboy.com/voice/public/asr/transcribe', {
method: 'POST',
headers: { 'x-key': process.env.DC_KEY, 'x-secret': process.env.DC_SECRET, 'Content-Type': 'application/json' },
body: JSON.stringify({ media_id: mediaId })
}).then(function (r) { return r.json(); });
return transcript.data.text;
}
With Twilio
Twilio's RecordingUrl needs your Twilio Account SID and Auth Token to fetch. If our media upload cannot reach it, download the file yourself and pass a public URL instead.
app.post('/twilio/recording-complete', async function (req, res) {
const media = await fetch('https://api-v2.dropcowboy.com/media/public/media', {
method: 'POST',
headers: { 'x-key': process.env.DC_KEY, 'x-secret': process.env.DC_SECRET, 'Content-Type': 'application/json' },
body: JSON.stringify({ url: req.body.RecordingUrl + '.wav', name: 'Twilio call ' + req.body.CallSid })
}).then(function (r) { return r.json(); });
const transcript = await fetch('https://api-v2.dropcowboy.com/voice/public/asr/transcribe', {
method: 'POST',
headers: { 'x-key': process.env.DC_KEY, 'x-secret': process.env.DC_SECRET, 'Content-Type': 'application/json' },
body: JSON.stringify({ media_id: media.data.media_id })
}).then(function (r) { return r.json(); });
saveTranscriptToCrm(req.body.CallSid, transcript.data.text);
res.sendStatus(200);
});
From a browser
Mint a site token with the media scope on your server, then call the same route with Authorization: Bearer <token>.
# On your server: trade your API key (needs numbers:write) for a site token. 1 hour max.
curl -s -X POST https://api-v2.dropcowboy.com/phone/public/embed/token \
-H "x-key: $KEY" -H "x-secret: $SECRET" \
-H "Content-Type: application/json" \
-d '{"site_id":"YOUR_SITE_ID","scope":["media"],"ttl_seconds":900}'
JS API / HTML tag
No browser widget and no HTML tag. Routes:
| Route | Body | What it does |
|---|---|---|
POST /voice/public/asr/transcribe |
{ media_id | url } |
Returns the transcript for that recording. Scope media:read. |
Auth and scopes
API key and secret on your server (media:read). From a browser, send Authorization: Bearer with a site token minted with the media scope.
Site token scope: media. Mint it with POST https://api-v2.dropcowboy.com/phone/public/embed/token using an API key with numbers:write. Tokens last up to one hour.
Limits
- Each call checks prepaid balance and returns 402 when funds are short.
- A url must be a WAV file (PCM 16-bit). Upload MP3 and other formats as media and pass media_id.
- A media_id MP3 answers 409 for a few seconds after upload while it is converted. Retry after a short wait.
- Audio can be up to 10 MB. url downloads also time out after 15 seconds.
Related REST
Routes are on https://api-v2.dropcowboy.com.
- Speech-to-text API reference - full reference
- POST /voice/public/asr/transcribe
- POST /media/public/media
Full reference: Media API.