Deepgram (Audio Transcription)
Deepgram is a speech-to-text API. In Velaclaw it is used for inbound audio/voice note transcription viatools.media.audio.
When enabled, Velaclaw uploads the audio file to Deepgram and injects the transcript
into the reply pipeline ({{Transcript}} + [Audio] block). This is not streaming;
it uses the pre-recorded transcription endpoint.
Getting started
1
Set your API key
Add your Deepgram API key to the environment:
2
Enable the audio provider
3
Send a voice note
Send an audio message through any connected channel. Velaclaw transcribes it
via Deepgram and injects the transcript into the reply pipeline.
Configuration options
- With language hint
- With Deepgram options
Notes
Authentication
Authentication
Authentication follows the standard provider auth order.
DEEPGRAM_API_KEY is
the simplest path.Proxy and custom endpoints
Proxy and custom endpoints
Override endpoints or headers with
tools.media.audio.baseUrl and
tools.media.audio.headers when using a proxy.Output behavior
Output behavior
Output follows the same audio rules as other providers (size caps, timeouts,
transcript injection).
Deepgram transcription is pre-recorded only (not real-time streaming). Velaclaw
uploads the complete audio file and waits for the full transcript before injecting
it into the conversation.
Related
Media tools
Audio, image, and video processing pipeline overview.
Configuration
Full config reference including media tool settings.
Troubleshooting
Common issues and debugging steps.
FAQ
Frequently asked questions about Velaclaw setup.