Free German Speech-to-Text
Upload your German audio or video and get an accurate, editable transcript in minutes — automatic, browser-based and free to start.
German is the most widely spoken native language in the European Union, with around 95 million native speakers. It is official in Germany, Austria, Switzerland, Liechtenstein, Luxembourg and parts of Belgium and Italy, which means German recordings turn up everywhere from Berlin newsrooms to Zurich boardrooms.
German also puts specific demands on a speech-recognition engine. Nouns are capitalised, so a transcript that gets this wrong is instantly recognisable as machine output — ConvertSpeech writes them correctly. Compound nouns are built by fusing words together rather than separating them, so a single spoken word can arrive as Rechtsschutzversicherung or Lebensmittelunverträglichkeit and must be written as one. And spoken numbers below a hundred run backwards: einundzwanzig is literally "one and twenty", which is why a naive system can turn 21 into 12.
Then there is the ß. It exists in Germany and Austria but not in Switzerland, where ss is used throughout — so the same sentence is spelled differently depending on where the speaker is from. Our engine follows the convention of the recording rather than forcing one standard, which is why the sample transcript below preserves the older spelling daß exactly as it was read.
Upload your audio or video and the engine writes out the spoken German as clean, searchable text with punctuation. It runs in your browser, needs no installation, and your first minutes are free with no account required.
What people transcribe German for
Journalists and interviews
Turn recorded German interviews into quotable text in minutes instead of typing them out by hand.
Researchers and students
Convert German lectures and research interviews into searchable notes you can quote and analyse.
Business and meetings
Keep written records of German calls and meetings so decisions and action points are never lost.
What a German transcript actually looks like
This is a real 45-second recording of a human reader, run through our AI engine exactly as you would run your own file. Nothing was corrected afterwards. Press play and read along.
AI engine, 24 seconds of processing
- 00:00Täufling, jetzt Abraham Petrowitsch Hannibal, einem Bruder derselben auszuliefern, der nach Petersburg gekommen war, um ihn auszulösen.
- 00:09Als der junge Hannibal achtzehn Jahre alt war, wurde er vom Zaren nach Frankreich geschickt, von wo er als Leutnant zurückkam.
- 00:17Seitdem war er von der Person seines kaiserlichen Herrn unzertrennlich.
- 00:21Er starb nach wechselnden Schicksalen als pensionierter General en chef unter der Kaiserin Katharina der zweiten im zweiundneunzigsten Lebensjahre.
- 00:32Puschkin selbst berichtet von ihm, daß seine heißen Leidenschaften und sein grenzenloser Leichtsinn ihn in große Verirrungen gestürzt, und wenn der Dichter selbst Bekenntnisse macht, so spielt er auch gern auf seiner eigenen Jugend.
Worth noticing: the Russian names Abraham Petrowitsch Hannibal and Katharina come out correctly, the French rank en chef is kept in French rather than being germanised, and the historic spelling daß is preserved instead of being modernised to dass. Achtzehn and zweiundneunzigsten stay written out as spoken.
This is careful, clearly articulated reading — close to a lecture or a dictated note. A noisy café interview with three people talking over each other will not come out this clean; nothing automatic will. Recording quality remains the single biggest factor in the result.
Recording: “Einleitung zu Die Russalka” from Ausgewählte Erzählungen by Alexander Pushkin, LibriVox (2015). Public Domain Mark 1.0. Excerpt from 1:00 to 1:45.
German accents and regional variants we handle
Standard German (Hochdeutsch) is what most broadcast, business and academic recordings use, and ConvertSpeech is tuned for it. It handles the standard spoken across Germany, Austria and Switzerland reliably.
German also has strong regional variation — Bavarian and Austrian German in the south, Swiss German (Schweizerdeutsch) in Switzerland, and Low German in the north. Swiss German in particular differs enough from standard German that heavy dialect is challenging for any automatic system: it is closer to a group of Alemannic dialects than to an accent, with its own vocabulary and vowel system. If your recording is strongly dialectal, expect the transcript to render it in standard German spelling rather than reproducing the dialect.
Austrian German sits much closer to the standard and transcribes well, with occasional slips on distinctly Austrian vocabulary — Jänner for Januar, Erdäpfel for Kartoffeln, Sackerl for Tüte. These are worth a quick scan when you proofread.
Tips for accurate German transcription
- Record with a good microphone and minimal background noise — audio quality is the single biggest factor.
- Keep the language set to German rather than auto-detect when you know the recording is German.
- One speaker at a time transcribes best; heavy crosstalk lowers accuracy.
- For interviews, enable speaker detection (a free registered feature) to label who said what.
- Proofread spoken numbers first. Because German says them back to front, digits are the most likely place to find a slip.
- Long compound nouns are usually right, but check technical or invented ones — they are the hardest words in any German recording.