Every word out of every video.

Drop a file or paste a link, pick how carefully we should listen, and a timestamped transcript comes back in a fraction of the runtime.

How carefully should we listen?

Three transcripts free. No card.

Paste a link from

Three engines

Choose the trade you actually want.

Speed and accuracy pull against each other, and which one matters depends on the recording. So the choice is yours per job, not ours once for everyone. These are measured runs on our own card, not projections.

On a deliberately poor recording — telephone bandwidth, hiss over the top — Ferrari made close to three times as many word errors as Volvo, and Mercedes about a quarter more. On a clean one, Mercedes and Volvo returned the same words.

Ferrari

3min per hour of audio

Strength
Back in minutes, not tens of minutes.
Cost
Twice the word errors of Mercedes, mostly names and numbers.

A tenth of the size, taking the first reading it hears rather than weighing several, and starting each line fresh instead of carrying the sentence over. On clean speech you will barely notice. On a noisy recording you will.

Run one on Ferrari →

Mercedes

default

7min per hour of audio

Strength
The best trade of the three on almost any recording.
Cost
A quarter more errors than Volvo when the audio is poor.

The full-size listener paired with a shortened writer, weighing five readings of every line before it picks one. This is what the site was built around, and what most people should be using.

Run one on Mercedes →

Volvo

10min per hour of audio

Strength
Fewest errors on noise, crosstalk and heavy accents.
Cost
A third longer, and no better than Mercedes on clean audio.

The same model as Mercedes, worked twice as hard: ten competing readings of every line instead of five, each carried further before one is chosen. On a clean recording it will hand back the same words. It earns its time on the difficult ones.

Run one on Volvo →

Every mode costs the same. The plan you are on buys hours, not quality, so a free transcript on Volvo is the same Volvo a paying customer gets.

What comes back

A transcript you can work with.

Timestamps down to the word

Every line carries a start and an end, and every word inside it does too.

  • 00:00:07This program is brought to you by Stanford University.
  • 00:00:20I am honored to be with you today for your commencement
  • 00:00:27from one of the finest universities in the world.
  • 00:00:35Truth be told, I never graduated from college.

Subtitles or plain text

Export the same transcript three ways, no reformatting.

SRTVTTTXT

99 languages, detected on their own

Paste something in Swedish, Turkish or Portuguese without telling us which.

Punctuation that reads like writing

Sentences, commas and question marks, so the text is usable as it lands.

Files up to six hours

Lectures, podcasts and full recorded meetings, up to three gigabytes. Long files are sent and transcribed in parts, so the queue can tell you which part it is on and a dropped connection only costs the part that was in the air.

The engine runs on our own machine.

Your audio is transcribed on hardware we own and run. It is not forwarded to a third party AI provider, and it is not kept to train anything. The file you upload is deleted after seven days unless you ask us to hold on to it. The transcript stays until you delete it.

Pricing

Cheaper than the going rate, on purpose.

Max

79 krper month

No monthly limit. Transcribe as many hours as you like, one file after another, for the same 79 kr.

Light

39 krper month

Ten hours of audio a month. Enough for a weekly podcast or a course. Unused hours do not roll over.

Free

Three transcripts

Full quality, full exports, no card. Then pick a plan or stop.

Skribia — audio and video to text, with timestamps