Fast, encrypted Speech-to-Text API

Transcribe and diarize audio through a high-speed REST API.
Hosted in France (EU), GDPR-compliant, AES-256 encrypted, and never trained on your data.

POST /api/v1/transcription-jobs
GET /transcriptions/{id}/transcript · 200 OK
Read full API documentation

Strictly isolated from any AI model training

Hosted on European infrastructure with AES-256 encryption and full GDPR compliance.

EU infrastructure

All servers, processing pipelines, and storage operate within European data centers, fully compliant with GDPR.

AES-256 encryption

Data is encrypted at rest across all storage layers, from upload to retrieval.

Zero model training

Audio files and output transcripts are never retained or used to train internal or third-party AI models.

Configurable retention

Retain transcripts as long as needed, or set zero-retention for immediate deletion after processing.

Need legal documentation for your compliance team?

Access our DPA →

Powered by the Vook engine

Get direct access to the transcription engine that powers Vook.ai.

Direct file ingestion

Up to 99% accuracy on clear audio. Send raw audio directly without preprocessing or output cleanup steps.

100+ languages

Guaranteed accuracy across a wide range of languages.

Speaker identification

Partition audio streams by speaker with exact start and end timestamps on every transcript.

Automatic language detection

Identify the source language automatically from the incoming audio stream.

Flexible export formats

Receive structured JSON alongside pre-formatted SRT, PDF, Word, or Markdown files.

Batch & webhooks

Submit large audio files asynchronously and receive an immediate HTTP callback when processing finishes.

Integration in 3 steps

1

Get your API key

Create an account to generate your key. Every account includes one free transcription per day.

2

Upload the file

Call the init endpoint, PUT the file to the signed URL, and trigger the transcription job.

3

Retrieve the transcript

Poll the job status or handle the webhook event to receive your formatted output.

Ready to build?

Sign up, generate an API key, and use your daily free transcription.

Need high volume or custom DPA? Talk to sales

API Frequently Asked Questions

Can't find your question? Send us a message!

What audio and video formats does the API support?

The same formats as the app: common audio (MP3, WAV, M4A, FLAC, OGG, OPUS, AAC) and video (MP4, MOV, MKV, WEBM). Files up to 5 hours long are supported.

Which languages does the API support?

The API transcribes 100+ languages. Pass 'auto' as the language to detect the spoken language automatically, or set a specific language code (like 'en' or 'fr') to force it.

How accurate is the transcription?

Vook.ai reaches up to 99% accuracy on clean audio and guarantees at least 90% across all transcriptions, with automatic speaker identification and timestamps on every result.

What does the API return?

A structured JSON transcript with speaker labels and word-level timestamps. You can also export to SRT, Word, PDF, and Markdown from the same endpoint.

How fast is transcription?

A one-hour file is transcribed in about a minute. High volumes are processed asynchronously and each job carries its own status.

Where can I find the documentation?

Read the REST API reference for full endpoint schemas, payload structures, and runnable examples.

How is the API priced?

Standard pricing is €0.60 per hour of audio, transcription and diarization included. For à la carte features or high volume, custom plans start from €0.20 per hour. Talk to our team for a quote.