Every word out of every video.

Drop a file or paste a link. A timestamped transcript comes back in a fraction of the runtime. Three free, no card.

How carefully should we listen?
  • Timestamps included
  • Runs on our own hardware
  • Files up to six hours

Paste a link from

From a link to text, while you watch.

A real run on this site: a YouTube link goes in, the words appear while the engine listens. The waiting is shortened, nothing else is.

Built for speech.

Interviews, podcasts, lectures and meetings are what the engine is trained on and where it is strong.

Podcasts

A full episode back as searchable text, with the timestamps to cut from.

It is bad at music. The engine listens for speech, and that is where it is strong: interviews, podcasts, lectures, meetings. Sung vocals over a mix leave it almost nothing to hold on to, so a song can come back nearly empty. That is the model's limit, not something wrong with the file you sent.

Three engines

Choose the trade you actually want.

Speed and accuracy pull against each other, and which one matters depends on the recording. So the choice is yours per job, not ours once for everyone. These are measured runs on our own card, not projections.

Mercedes

default

7min per hour of audio

The best trade of the three on almost any recording.

The full-size listener paired with a shortened writer, weighing five readings of every line before it picks one. This is what the site was built around, and what most people should be using.

Run one on Mercedes →

3

Ferrari

min per hour of audio

Back in minutes, not tens of minutes.

Twice the word errors of Mercedes, mostly names and numbers.

Run one on Ferrari →
10

Volvo

min per hour of audio

Fewest errors on noise, crosstalk and heavy accents.

A third longer, and no better than Mercedes on clean audio.

Run one on Volvo →

On a deliberately poor recording — telephone bandwidth, hiss over the top — Ferrari made close to three times as many word errors as Volvo, and Mercedes about a quarter more. On a clean one, Mercedes and Volvo returned the same words. Every mode costs the same. The plan you are on buys hours, not quality, so a free transcript on Volvo is the same Volvo a paying customer gets.

What comes back

Finished text, not raw output.

Timestamps down to the word

Every line carries a start and an end, and every word inside it does too.

  • 00:00:07This program is brought to you by Stanford University.
  • 00:00:20I am honored to be with you today for your commencement
  • 00:00:27from one of the finest universities in the world.
  • 00:00:35Truth be told, I never graduated from college.

Subtitles or plain text

Export the same transcript three ways, no reformatting.

SRTVTTTXT

99 languages, detected on their own

Paste something in Swedish, Turkish or Portuguese without telling us which.

Punctuation that reads like writing

Sentences, commas and question marks, so the text is usable as it lands.

Files up to six hours

Lectures, podcasts and full recorded meetings, up to three gigabytes. Long files are sent and transcribed in parts, so the queue can tell you which part it is on and a dropped connection only costs the part that was in the air.

The engine runs on our own machine.

Your audio is transcribed on hardware we own and run. It is not forwarded to a third party AI provider, and it is not kept to train anything. The file you upload is deleted after seven days unless you ask us to hold on to it. The transcript stays until you delete it.

Read the privacy page before you upload anything.

It is short, and it is written from the code that actually runs — every retention period on it is a real setting, not a promise made in the abstract. It says where your audio is processed, what is kept, for how long, and what we can never see. If any of it does not suit you, it is better to know that now than after you have uploaded an interview.

Read the privacy page

Pricing

Cheaper than the going rate, on purpose.

Max

$9.99per month

or $99 a yearTwo months free

No monthly limit. Transcribe as many hours as you like, one file after another, for the same $9.99.

Light

$5.99per month

Ten hours of audio a month. Enough for a weekly podcast or a course. Unused hours do not roll over. If they run out, three shorter transcripts of up to half an hour each still go through before the month turns.

Free

Three transcripts

Full quality, full exports, no card. Then pick a plan or stop.

Something not working, or a question?

Write to us. The reply comes from the person who built it.

skribiasupport@gmail.com
Email us
Skribia — audio and video to text, with timestamps