5 min read

Text-to-speech and speech-to-text in the browser (Web Speech API in 2026)

The Web Speech API turns any browser into a TTS engine and a live-transcription client. Here's what actually works, what's still broken, and when to pay for ElevenLabs.

By
Software Engineer ยท M.Sc. Mechanical Engineering ยท Ontario, Canada

Two APIs, wildly different maturity

The Web Speech API has two halves:

  • SpeechSynthesis (text โ†’ speech) - universally supported, 100+ voices, works offline in most browsers.
  • SpeechRecognition (speech โ†’ text) - supported by Chrome, Edge, Safari, Opera. Not supported by Firefox as of 2026.

Both are free, both are private-ish (details below), both are one function call.

Text-to-speech: what "100 voices" actually means

window.speechSynthesis.getVoices() returns whatever your OS + browser combination exposes. That's usually:

  • macOS Safari - ~140 voices via Apple's Speech framework. High quality neural voices for major languages, older robotic voices for the rest.
  • Windows Chrome / Edge - 60-80 voices via Microsoft's SAPI. Includes the newer natural-sounding "Neural" voices on Windows 11.
  • iOS Safari - 50+ voices, all Apple neural.
  • Android Chrome - 30-60 voices via Google's TTS engine.
  • Linux Chrome - 20-30 voices via espeak (robotic) unless a better engine is installed.

So "100 voices" is technically correct but ships wildly different quality depending on the client. A user on Windows 11 Edge hears studio-quality neural voices; a user on Ubuntu hears the 1990s espeak beep.

Text to Speech shows every voice available in the current browser, grouped by language. If a user complains about quality, switching to a different voice usually solves it - the good ones are hidden in the list.

Controls that matter

Three sliders on every voice:

  • Rate - 0.5 (slow) to 2.0 (fast). 1.0 is natural.
  • Pitch - 0 (deep) to 2 (high). 1.0 is natural.
  • Volume - 0 (silent) to 1 (loud).

Combined, they let you fake different speakers from a single voice - a lower-pitch, slower-rate variant of a female neural voice sounds different enough for basic dialogue mockups.

Downloading TTS as audio

The Web Speech API doesn't expose the raw audio buffer, so you can't easily save the output as an MP3. The workarounds:

  1. Record playback with Voice Recorder. Play the TTS through your speakers, record the mic. Quality is decent for non-critical use.
  2. Use a paid API for real audio. ElevenLabs (best quality in 2026), OpenAI TTS, Azure Speech, Google Cloud TTS. Costs pennies per minute; returns MP3/WAV directly.

For app prototyping, in-browser TTS is fine. For a published podcast intro or a YouTube voiceover, pay for real TTS.

Speech-to-text: what actually works

Every "SpeechRecognition supported" browser sends your mic audio to the browser vendor's cloud recognizer:

  • Chrome / Edge / Opera โ†’ Google's speech recognition.
  • Safari โ†’ Apple's on-device (mostly) recognition.

The audio doesn't touch a third-party site like 712tools - but it does touch Google or Apple's servers on most browsers. That's an important nuance for anyone treating the browser as a "local" transcription tool.

Accuracy in 2026: 90-95% on clean audio with a good mic. 70-80% with a bad mic in a noisy room. Accents and jargon reduce accuracy more than most people expect.

Continuous mode and interim results

Two flags matter:

  • recognition.continuous = true - keep listening after the first result. Without this, recognition stops after every pause.
  • recognition.interimResults = true - surface partial results as the recognizer works. Without this, you only see finalized text.

Both are on in Speech to Text - you see live transcription as you talk, with completed sentences appearing in the main transcript and the in-progress phrase shown italicized.

Language matters

Every speech recognition API needs the source language declared up front. en-US is not the same as en-GB (British), en-IN (Indian), or en-AU (Australian). Pick the closest match - accuracy drops significantly on the wrong regional variant.

Speech to Text lists 35+ languages and regional variants. Set it once and forget it, unless you're transcribing multilingual content.

When to pay for a real transcription service

The Web Speech API is fine for:

  • Live dictation in a note-taking app.
  • Voice-triggered UI (search, commands).
  • Rapid transcription of short clips.

It's not great for:

  • Transcribing existing audio files. The Web Speech API only takes live mic input - you'd have to play the file into the mic (or use a virtual audio device). Use Whisper (OpenAI's transcription API), AssemblyAI, Deepgram, or Rev instead.
  • Producing shipping-quality transcripts. Speaker diarization, punctuation, timestamps, edit-ready output - all things the pro APIs handle and Web Speech doesn't.
  • Privacy-sensitive content. Sending audio to Google / Apple's servers isn't compatible with HIPAA or many GDPR-sensitive use cases. Use a local Whisper model or a HIPAA-BAA-signed vendor.

Recording as a hedge

If TTS/STT quality is inconsistent for your use case, Voice Recorder captures raw mic audio to WebM (Opus). You can then:

  • Upload to a real transcription service (Whisper API, AssemblyAI).
  • Import into a DAW for cleanup and re-export.
  • Feed to a paid TTS with a voice-cloning feature (ElevenLabs Voice Design) to build a synthetic clone.

Local recording keeps the source under your control; you decide where it goes next.

The workflow

The three tools most people combine:

  • Text to Speech - draft narration, accessibility previews, quick prototypes.
  • Speech to Text - live dictation, quick transcription of your own voice.
  • Voice Recorder - capture raw audio when you need the file itself.

All three run 100% in the browser. TTS uses the OS's local speech engine (mostly offline); STT uses the browser vendor's cloud recognizer; recording is fully local.

Tools mentioned in this post

Written by Shan

Shan builds 712 Tools. He holds a Master's degree in Mechanical Engineering and now works as a Software Engineer, shipping browser-based developer utilities out of Ontario, Canada. Learn more ยท 712studiogames@gmail.com