Free conversions left: 1

Free Japanese Speech-to-Text

Upload your Japanese audio or video and get an accurate, editable transcript in minutes — automatic, browser-based and free to start.

Drop your Japanese file here

OR

Supported formats: MP3, WAV, MP4, FLAC, WEBM, M4A, OGG, OPUS, AAC, AMR, AIFF, AIF, WMA, MOV, MKV, AVI, MPG, MPEG, 3GP, WMV

Max. file size per upload: 100 MB

Free trial: we transcribe the first 3 minutes of your file · Register free for more

Japanese is spoken by around 125 million people, almost all of them in Japan, and it carries one of the world's largest media, business and research economies. Interviews, lectures, board meetings and video content in Japanese are recorded constantly, and most of it eventually needs to become searchable text.

Japanese also asks something of a speech-recognition engine that almost no other language does: it is written in three scripts at once, and they are interleaved inside a single sentence. Kanji carry the content words, hiragana carry the grammar and inflections, and katakana carry loanwords, foreign names and onomatopoeia. Choosing between them is not a formatting preference — it is the transcription. The same syllables spoken as "kami" are 紙 for paper, 神 for a god and 髪 for hair: identical in the air, three different characters on the page, and only the surrounding sentence decides which one is right. In the sample below you can watch this happen: ゴム, the word for rubber that Japanese borrowed from Dutch, is correctly written in katakana, while 西洋人 and 手拭 in the same passage are written in kanji and the particles around them in hiragana.

On top of that, written Japanese puts no spaces between words. An engine cannot simply write down sounds and let spacing do the rest — it has to decide where each word ends, because there is no visible boundary to fall back on. Punctuation is not spoken either, so the 、and 。 in a transcript are inferred from phrasing rather than heard. Spoken Japanese leans on pitch accent for some of the work: 箸 (chopsticks), 橋 (bridge) and 端 (edge) are all "hashi" and are told apart in Tokyo speech by where the pitch falls. Counting adds another layer, because Japanese counts things with a classifier that changes shape with the number — three long objects are 三本 (sanbon) but six are 六本 (roppon), while flat things take 枚 and small animals take 匹.

Politeness is grammatical rather than decorative. Moving from casual to keigo does not soften a verb, it replaces it: 見る becomes ご覧になる when the listener is doing the looking and 拝見する when the speaker is, and 言う becomes おっしゃる or 申す the same way. A transcript that gets this wrong does not just read oddly — it reports the wrong social register. Upload your audio or video and the engine writes out the spoken Japanese as clean, searchable text with the scripts mixed as a Japanese writer would mix them. It runs in your browser, needs no installation, and your first minutes are free with no account required.

Use cases

What people transcribe Japanese for

Business and meetings

Keep written minutes of Japanese meetings and calls so decisions and action points are captured in text you can search.

Media and creators

Generate Japanese subtitles and show notes straight from your video or podcast audio, in proper kanji and kana.

Researchers and students

Convert Japanese lectures and fieldwork interviews into searchable text you can quote, cite and study from.

Real example

What a Japanese transcript actually looks like

This is a real 45-second recording of a human reader, run through our AI engine exactly as you would run your own file. Nothing was corrected afterwards. Press play and read along.

AI engine, 20 seconds of processing

  1. 00:00私のじっとしている間にだいぶ多くの男が塩を浴びに出て来たが、いずれも胴と腕と股は出していなかった。
  2. 00:09女はことさら肉を隠しがちであった。
  3. 00:13大抵は頭にゴム線のずきんをかぶって海茶や昆夜藍の色を波間に浮していた。
  4. 00:20そういうあり様を目撃したばかりの私の目には――それまたひとつで済まして、みんなの前に立っているこの西洋人がいかにも珍しく見えた。
  5. 00:31彼はやがて自分の脇をかえりみて、そこに転んでいる日本人に一言二言何か言った。その日本人は砂の上に落ちた手拭を拾い上げているところであったがそれを取り上げるや否や腰に牽いて、

Worth noticing: the three scripts are mixed the way a Japanese writer would mix them, without anyone telling the engine to do so. ゴム — the Dutch loanword for rubber — is in katakana, while 西洋人, 手拭 and 日本人 in the same passage are in kanji and every particle and verb ending sits in hiragana. 一言二言 is written in kanji rather than spelled out phonetically, and the literary 否や survives intact. The commas and full stops were never spoken; they were placed from the reader's phrasing, and they land at the clause boundaries where a Japanese editor would put them.

It is not perfect, and the mistakes are instructive. Sōseki wrote 猿股一つ — the Westerner is standing there in nothing but a pair of trunks — and the engine heard それまたひとつ, a real-sounding string of kana that means nothing here. The list of colours 海老茶や紺や藍 came out as 海茶や昆夜藍, with the kanji chosen for sound rather than sense. That is what homophone pressure looks like in practice. And this is careful, clearly articulated reading, close to a lecture or a dictated note. A noisy izakaya interview with three people talking over each other will not come out this clean; nothing automatic will. Recording quality remains the single biggest factor in the result.

Recording: section 2 of こころ (Kokoro) by Natsume Sōseki, LibriVox (2014). Public Domain Mark 1.0. Excerpt from 1:30 to 2:15.

Variants and accents

How ConvertSpeech handles spoken Japanese

Standard Japanese — hyōjungo, built on the speech of Tokyo — is what broadcasting, business and academic recordings use, and ConvertSpeech is tuned for it. It is also the variety whose pitch accent the engine treats as the reference, which matters because pitch is what separates several common word pairs in speech before context gets a chance to.

Kansai-ben, spoken around Osaka, Kyoto and Kobe, is the strongest regional variety you are likely to record, and it differs in more than accent. The copula is や rather than だ, negatives are formed with -へん or -ん instead of -ない, and everyday words are simply different: おおきに for thank you, あかん for no good, ちゃう for that's not it. Clear Kansai speech transcribes usefully, but expect the engine to lean toward standard forms in places, and give dialect-heavy audio a closer proofread. Tōhoku and Kyushu varieties are further from the standard again; the more dialectal the recording, the more the transcript reads as standard Japanese wearing regional vocabulary.

Politeness is handled as spoken, in both directions. Keigo-heavy speech — a customer-service call, a formal presentation, an interview with someone senior — produces いらっしゃいます, させていただきます and ございます as said, and casual speech among friends produces plain forms and dropped particles as said. The engine does not normalise one into the other, which is the behaviour you want: in Japanese, the politeness level is part of the content, not a style setting. Where it does struggle is the same place humans do — humble and honorific forms of the same verb sound nothing like the dictionary form, so unusual keigo is worth a glance when you proofread.

Tips

Tips for accurate Japanese transcription

FAQ

Japanese speech-to-text — questions

Is Japanese speech-to-text really free?
Yes — you can transcribe Japanese audio for free with no sign-up. Create a free account if you need more transcription time.
Does the transcript come out in kanji and kana, or in romaji?
In kanji and kana — normal written Japanese, the way the speech would ordinarily be written down. You get a mix of kanji, hiragana and katakana with Japanese punctuation, not a romanised phonetic version. If you need romaji you would convert the finished text afterwards.
How does the engine decide between kanji, hiragana and katakana?
From context, the same way a writer does. Content words go into kanji, grammatical endings and particles into hiragana, and loanwords, foreign names and onomatopoeia into katakana. In the sample above the Dutch-derived word ゴム is correctly in katakana while 西洋人 in the same sentence is in kanji.
What happens with homophones like 機械 and 機会?
The engine picks the kanji that fits the surrounding sentence. That works well in ordinary speech, and it is also the most likely place to find a mistake — the sample transcript on this page contains exactly such a slip, left in deliberately. Homophone-dense technical vocabulary is worth a proofreading pass.
Does it handle Kansai-ben and other regional Japanese?
Clear Kansai-ben transcribes usefully, though the transcript may lean toward standard forms in places. Standard Japanese gives the most reliable results; the further a recording sits from it — strong Tōhoku or Kyushu speech, for instance — the more proofreading it needs, and that is true of any automatic system.
Does it handle keigo as well as casual speech?
Both, as spoken. Honorific and humble forms such as いらっしゃいます or 拝見する are transcribed as said rather than flattened into plain forms, and casual speech stays casual. Since politeness in Japanese is grammatical, this keeps the register of the original intact.
Can I transcribe Japanese video, and what formats can I upload?
Yes to video — MP4 and WEBM work, and we extract the spoken Japanese automatically. All common formats are supported, 20 in total: audio like MP3, WAV, M4A, OGG or FLAC, and video like MP4, MOV, MKV, AVI or WEBM.
How long does the transcription take?
Most files are ready within a couple of minutes, depending on the length of the recording. The 45-second sample above took 20 seconds.
Other languages

Transcribe another language