Free conversions left: 1

Free Spanish Speech-to-Text

Upload your Spanish audio or video and get an editable transcript in minutes — accents, ñ and inverted question marks included, browser-based and free to start.

Drop your Spanish file here

OR

Supported formats: MP3, WAV, MP4, FLAC, WEBM, M4A, OGG, OPUS, AAC, AMR, AIFF, AIF, WMA, MOV, MKV, AVI, MPG, MPEG, 3GP, WMV

Max. file size per upload: 100 MB

Free trial: we transcribe the first 3 minutes of your file · Register free for more

Spanish is spoken natively by around 500 million people, second only to Mandarin Chinese, and it is an official language in twenty countries — Spain, Equatorial Guinea and eighteen states across Latin America — as well as in Puerto Rico. That spread is why Spanish recordings arrive from Madrid newsrooms, Mexico City call centres, Buenos Aires podcast studios and Andean field research, all of them writing the same language and none of them saying it quite the same way.

Spanish spelling maps closely onto sound, which moves most of the difficulty somewhere else: into the marks. The tilde is not decoration, it carries meaning. Papá is dad and papa is a potato; sí is yes and si is if; término, termino and terminó are a noun, a first person and a third person separated by nothing but where the accent sits. The ñ is a letter in its own right, not an n with a hat — año is a year, ano is not. ConvertSpeech writes all of these, which is what separates readable Spanish from stripped-down machine output.

Then there are the inverted marks. Spanish opens a question with ¿ and an exclamation with ¡, and it opens them where the question actually begins, which is often in the middle of a sentence rather than at the start. This matters more in Spanish than the symmetry suggests: word order frequently does not distinguish a question from a statement — "tienes hambre" and "¿tienes hambre?" are the same words — so the marks are what tell the reader which one it is. The sample transcript below shows the engine placing a ¿ mid-sentence after a comma, exactly where a human editor would put it.

Pronunciation adds a last layer that no amount of clean audio removes. Speakers with seseo pronounce casa and caza identically, yeísmo makes cayó and calló the same sound, b and v are one sound in Spanish, and h is silent — so the spelling of tuvo versus tubo, or haya versus halla versus aya, can only be settled by context. Upload your audio or video and the engine writes out the spoken Spanish as clean, searchable text with punctuation. It runs in your browser, needs no installation, and your first minutes are free with no account required.

Use cases

What people transcribe Spanish for

Journalists and cross-border interviews

Turn recorded Spanish interviews into quotable text in minutes, whether the speaker is from Seville, Guadalajara or Montevideo — one project, several accents, one transcript.

Researchers and fieldwork

Convert Spanish lectures, oral histories and focus groups into searchable notes you can quote, code and analyse without typing them out first.

Podcasters and creators

Produce Spanish show notes, subtitles and articles straight from your episode audio, so your back catalogue becomes searchable text as well as sound.

Real example

What a Spanish transcript actually looks like

This is a real 45-second recording of a human reader, run through our AI engine exactly as you would run your own file. Nothing was corrected afterwards. Press play and read along.

AI engine, 24 seconds of processing

  1. 00:00Sacristán de la iglesia de Monserrate, le destinaba a seguir la carrera de derecho, porque se le había metido en la cabeza que el mocoso aquel llegaría a ser personaje, quizá orador célebre, ¿por qué no ministro?
  2. 00:13La futura celebridad habló así a su compañero: «¡Mia tú, Carso!
  3. 00:21si a mí me dieran esas chanzas, de la galleta que les pegaba les ponía la cara verde.
  4. 00:26Pero tú no tienes coraje.
  5. 00:28Yo digo que no se deben poner motes a las personas.
  6. 00:32¿Sabes tú quién tiene la culpa?
  7. 00:34Pues Posturitas, el de la casa de empeños; ayer fue contando que su mamá había dicho que a tu abuela y a tus tías las llaman las miau, porque tienen.

Worth noticing: the first line ends with ¿por qué no ministro? and the ¿ is opened mid-sentence, after a comma, exactly where the question starts rather than at the beginning of the sentence — this is the rule most automatic Spanish output gets wrong. The direct speech opens with «¡Mia tú, and the engine chose Spanish angular quotation marks and the inverted exclamation on its own. The accents land where they change the word: tú and quién in ¿Sabes tú quién tiene la culpa? are marked, while the unaccented tu in a tu abuela is not. Sacristán, había, llegaría, quizá, célebre, compañero, ponía, empeños, mamá and tías all come out correctly. It is not flawless — the boy is called Cadalso in the novel and the transcript writes Carso, the parish of Monserrat becomes Monserrate, and proper names are where a Spanish transcript slips first.

This is careful, clearly articulated reading — closer to a lecture or a dictated note than to real conversation. A noisy café interview with three people talking over each other, or a fast Caribbean speaker swallowing every syllable-final s, will not come out this clean; nothing automatic will. The last line also stops mid-sentence because that is simply where the 45 seconds end. Recording quality remains the single biggest factor in the result.

Recording: “Miau” by Benito Pérez Galdós, chapter 1, LibriVox (2026). Public Domain Mark 1.0. Excerpt from 3:29 to 4:14.

Variants and accents

Spanish accents and regional variants we handle

European (Castilian) Spanish keeps a distinction most of the Spanish-speaking world has lost: c before e or i, and z, are pronounced /θ/, the "th" of think, so casa and caza are different words to the ear as well as on paper. The technical name for this is distinción, not ceceo — ceceo is the opposite merger, where /s/ itself is realised as /θ/, and it is a regional feature of parts of Andalusia rather than the Spanish standard. Peninsular speech is also where you meet vosotros and its own verb endings — habláis, tenéis, sois, vosotros os — which simply do not occur in Latin American recordings.

Mexican Spanish is the largest single variety by number of speakers and one of the most straightforward to transcribe: syllable-final s stays clearly pronounced, and the rhythm is even. What needs a second look is vocabulary and place names of Nahuatl origin — elote, popote, chamaco, guajolote — and the letter x, which in México, Oaxaca and Xalapa is pronounced like a j, but as /ks/ in taxi and as "sh" in Xochimilco. Nothing about that is predictable from spelling, so proper nouns are worth scanning.

Rioplatense Spanish, around Buenos Aires and Montevideo, has two features that change the text itself. Voseo replaces tú with vos and takes its own verb forms — vos hablás, vos tenés, vos sos, and imperatives like vení, mirá, decime — and these are correct Spanish, not errors to be normalised into tú eres. The second is žeísmo: ll and y are pronounced as a "sh" or "zh" sound, so calle sounds like "cashe" and yo like "sho". The engine transcribes what the words are, not how they were pronounced, so the output stays in standard spelling.

Caribbean Spanish — Cuba, Puerto Rico, the Dominican Republic and the coasts of Venezuela and Colombia — is the hardest group for any automatic system, and for an honest reason. Syllable-final s is aspirated to an h or dropped entirely, so los dos becomes "loh doh" and las casas can sound identical to la casa. Since that s is what marks the plural, the engine has to recover grammatical information the speaker never pronounced, using context alone. Final consonants weaken, para contracts to pa, and the tempo is fast. Clear audio helps here more than anywhere else.

Andean Spanish, in the highlands of Peru, Bolivia and Ecuador, sits at the other end: it is conservative, syllable-final s survives intact, and in parts of the region the old ll/y distinction that yeísmo erased elsewhere is still alive. The complication is vocabulary rather than sound — Quechua and Aymara loanwords and place names run through everyday speech, and bilingual speakers may show vowel patterns from those languages. Chilean Spanish deserves its own warning: fast, heavily aspirated, with contracted verb endings like ¿cachái? and ¿estái?, it is one of the more demanding varieties to get right.

Tips

Tips for accurate Spanish transcription

FAQ

Spanish speech-to-text — questions

Is Spanish speech-to-text really free?
Yes — you can transcribe Spanish audio for free with no sign-up. Create a free account if you need more transcription time.
Does it work for both European and Latin American Spanish?
Yes. It handles Castilian Spanish, where c and z are pronounced /θ/, as well as the seseo of Latin America and the Canaries, where the same letters sound like an s. Because writing keeps the distinction that seseo speech loses, pairs like casa and caza are resolved from context — clear audio gives the most reliable result.
Does the transcript use the inverted ¿ and ¡?
Yes, and it places them where the question or exclamation actually begins, which in Spanish is frequently mid-sentence rather than at the start. The sample transcript above shows a ¿ opened after a comma inside a longer sentence. It is worth a quick scan when you proofread, since a missing opening mark is the most visible sign of machine-written Spanish.
Are accents and the letter ñ written correctly?
Yes. The tilde is applied where it belongs — tú versus tu, quién versus quien, terminó versus termino — and ñ is written as its own letter rather than being flattened to n. The transcript is normal Spanish text you can paste straight into a document.
Does it handle voseo — vos hablás, vos sos?
Yes. Voseo forms from Argentina, Uruguay, Paraguay, Central America and parts of Colombia are transcribed as spoken rather than being rewritten into tú forms. If a speaker says vos tenés, that is what the transcript says.
What about Caribbean Spanish, where the final s disappears?
It works, and it is the variety where recording quality matters most. Aspirating or dropping syllable-final s removes the plural marking from the sound, so the engine reconstructs it from context — los dos, las casas and similar phrases are the first place to check when you proofread a Cuban, Dominican or Puerto Rican recording.
Can I transcribe Spanish video, and what formats can I upload?
Yes to video — MP4, MOV, MKV, AVI and WEBM all work, and we extract the spoken Spanish automatically. Twenty formats are supported in total, including audio like MP3, WAV, M4A, OGG and FLAC.
How long does the transcription take?
Most files are ready within a couple of minutes, depending on the length of the recording. The 45-second sample above took 24 seconds.
Other languages

Transcribe another language