Web Speech API: Complete Text-to-Speech Guide for Developers

The Web Speech API lets you add text-to-speech to any web app without external services or API keys. This guide covers SpeechSynthesis from basic usage to production-ready implementation, including voice selection, event handling, accessibility, and cross-browser quirks.

What the Web Speech API Does

The Web Speech API is a browser-native JavaScript API with two halves: SpeechSynthesis (text-to-speech) and SpeechRecognition (speech-to-text). This guide focuses on SpeechSynthesis — converting text into spoken audio using voices provided by the operating system.

No API keys. No external services. No npm packages. The browser talks using the same speech engine that powers your OS's built-in screen reader and voice assistant. All processing happens on the client device, which means zero latency, zero cost, and full privacy.

Browser Support

Browser SpeechSynthesis SpeechRecognition
Chrome 33+ Full support Full support (webkitSpeechRecognition)
Firefox 49+ Full support Limited (flag required)
Safari 7+ Full support 14.1+ (partial)
Edge 14+ Full support Full support (Chromium-based)
iOS Safari 7+ (with quirks) 14.5+ (partial)
Android Chrome 33+ (with quirks) Full support

Try the API right now with QTool's Text-to-Speech tool. Paste any text, select a voice, and hear it spoken — all in your browser.

Basic Text-to-Speech in 5 Lines

javascript — Minimal TTS
// Create an utterance (the thing to speak)
const utterance = new SpeechSynthesisUtterance('Hello, world!');

// Speak it
speechSynthesis.speak(utterance);

// That's it. The browser is now talking.

The SpeechSynthesisUtterance object represents a single piece of text to be spoken. The speechSynthesis object (available on window) is the controller that manages the speech queue, voice selection, and playback state.

javascript — Basic Controls
// Pause speech
speechSynthesis.pause();

// Resume speech
speechSynthesis.resume();

// Stop speech (clears the queue)
speechSynthesis.cancel();

// Check if currently speaking
console.log(speechSynthesis.speaking);  // true or false

// Check if paused
console.log(speechSynthesis.paused);    // true or false

// Check if there are pending utterances
console.log(speechSynthesis.pending);   // true or false

Voice Selection and Management

Every operating system comes with a different set of voices. macOS typically provides 60-80 voices across 40+ languages. Windows 10/11 has 20-40 voices. Mobile platforms have 30-50. Some browsers also offer cloud-based voices with higher quality (Chrome's Google voices, for example).

Loading Voices (The Async Trap)

On most browsers, speechSynthesis.getVoices() returns an empty array on the first call. The voice list is loaded asynchronously. You must listen for the voiceschanged event.

javascript — Reliable Voice Loading
function getVoices() {
  return new Promise((resolve) => {
    let voices = speechSynthesis.getVoices();

    if (voices.length > 0) {
      resolve(voices);
      return;
    }

    // Voices not loaded yet. Wait for the event.
    speechSynthesis.addEventListener('voiceschanged', () => {
      voices = speechSynthesis.getVoices();
      resolve(voices);
    }, { once: true });
  });
}

// Usage
const voices = await getVoices();
console.log(`Found ${voices.length} voices`);

// List all English voices
const englishVoices = voices.filter(v => v.lang.startsWith('en'));
englishVoices.forEach(v => {
  console.log(`${v.name} (${v.lang}) ${v.localService ? '[local]' : '[cloud]'}`);
});

Selecting a Voice

javascript — Choose a Specific Voice
const voices = await getVoices();

// Find a specific voice by name
const preferred = voices.find(v => v.name === 'Google UK English Female')
  || voices.find(v => v.name.includes('Samantha'))
  || voices.find(v => v.lang === 'en-US');

const utterance = new SpeechSynthesisUtterance('Hello from a specific voice.');
utterance.voice = preferred;
speechSynthesis.speak(utterance);

Voice Object Properties

Property Type Description
name string Human-readable voice name (e.g., "Google UK English Female")
lang string BCP 47 language tag (e.g., "en-US", "de-DE", "ja-JP")
localService boolean true = runs locally (offline-capable), false = cloud-based
voiceURI string Unique identifier for the voice
default boolean true if this is the browser's default voice for the language

Need to analyze the text before speaking it? QTool's Text Analyzer gives you word count, character count, reading time, and readability scores instantly.

Controlling Rate, Pitch, and Volume

javascript — Speech Parameters
const utterance = new SpeechSynthesisUtterance('This demonstrates rate, pitch, and volume control.');

// Rate: 0.1 (very slow) to 10 (very fast). Default: 1
utterance.rate = 1.3;  // Slightly faster than normal

// Pitch: 0 (lowest) to 2 (highest). Default: 1
utterance.pitch = 0.9; // Slightly deeper voice

// Volume: 0 (silent) to 1 (loudest). Default: 1
utterance.volume = 0.8;

// Language (overrides voice's default language)
utterance.lang = 'en-US';

speechSynthesis.speak(utterance);

Practical Rate Settings

Rate Use Case
0.7 - 0.9 Accessibility: users who need slower speech
1.0 Default: natural conversational speed
1.2 - 1.5 Speed reading: experienced listeners, podcast-style
1.5 - 2.0 Power users: familiar with accelerated speech
2.0+ Screen reader users: trained to process fast speech

Event Handling and Progress Tracking

SpeechSynthesisUtterance fires events at key moments during speech. These are essential for building a proper UI with progress indicators, word highlighting, and state management.

javascript — Utterance Events
const utterance = new SpeechSynthesisUtterance('The quick brown fox jumps over the lazy dog.');

// Fired when speech starts
utterance.onstart = (event) => {
  console.log('Started speaking');
};

// Fired when speech ends normally
utterance.onend = (event) => {
  console.log(`Finished. Elapsed: ${event.elapsedTime}ms`);
};

// Fired when speech is paused
utterance.onpause = (event) => {
  console.log('Paused at character:', event.charIndex);
};

// Fired when speech resumes after pause
utterance.onresume = (event) => {
  console.log('Resumed');
};

// Fired at word boundaries (not all browsers)
utterance.onboundary = (event) => {
  if (event.name === 'word') {
    const word = utterance.text.substring(
      event.charIndex,
      event.charIndex + event.charLength
    );
    console.log(`Word: "${word}" at char ${event.charIndex}`);
    // Use this for word-by-word highlighting
  }
};

// Fired on error
utterance.onerror = (event) => {
  console.error('Speech error:', event.error);
  // Common errors: 'canceled', 'interrupted', 'network', 'synthesis-failed'
};

speechSynthesis.speak(utterance);

Word Highlighting

The boundary event provides the character index and length of each word as it is spoken. You can use this to highlight the current word in the UI.

javascript — Highlight Words as They Are Spoken
function speakWithHighlighting(text, containerEl) {
  // Wrap each word in a span
  const words = text.split(/(\s+)/);
  containerEl.innerHTML = words.map((word, i) =>
    word.trim() ? `<span data-index="${i}">${word}</span>` : word
  ).join('');

  const utterance = new SpeechSynthesisUtterance(text);

  utterance.onboundary = (event) => {
    if (event.name !== 'word') return;

    // Remove previous highlight
    containerEl.querySelectorAll('.speaking').forEach(el =>
      el.classList.remove('speaking')
    );

    // Find the span containing this character index
    const charIndex = event.charIndex;
    let currentIndex = 0;
    for (const span of containerEl.querySelectorAll('span[data-index]')) {
      const end = currentIndex + span.textContent.length;
      if (charIndex >= currentIndex && charIndex < end) {
        span.classList.add('speaking');
        span.scrollIntoView({ behavior: 'smooth', block: 'center' });
        break;
      }
      currentIndex = end + 1; // +1 for the space
    }
  };

  utterance.onend = () => {
    containerEl.querySelectorAll('.speaking').forEach(el =>
      el.classList.remove('speaking')
    );
  };

  speechSynthesis.speak(utterance);
}

// CSS: .speaking { background: rgba(0, 212, 255, 0.3); border-radius: 3px; }
Boundary Events Are Not Universal

The boundary event fires reliably in Chrome and Edge (Chromium) but is less consistent in Firefox and Safari. Always design your UI to work without word highlighting as a fallback, and test on your target browsers.

Handling Long Text and Chunking

Chrome has a known issue: speech stops after approximately 200-300 characters (or about 15 seconds) of continuous speech. This is a bug that has persisted for years. The workaround is to split long text into chunks and speak them sequentially.

javascript — Chunked Speech for Long Text
function speakLongText(text, options = {}) {
  const { voice, rate = 1, pitch = 1, volume = 1 } = options;

  // Split on sentence boundaries
  const chunks = text.match(/[^.!?]+[.!?]+/g) || [text];

  let currentIndex = 0;
  let isPaused = false;
  let isCancelled = false;

  function speakNext() {
    if (isCancelled || currentIndex >= chunks.length) return;

    const utterance = new SpeechSynthesisUtterance(chunks[currentIndex].trim());
    if (voice) utterance.voice = voice;
    utterance.rate = rate;
    utterance.pitch = pitch;
    utterance.volume = volume;

    utterance.onend = () => {
      currentIndex++;
      if (!isPaused && !isCancelled) {
        speakNext();
      }
    };

    utterance.onerror = (event) => {
      if (event.error !== 'canceled') {
        console.error('Chunk error:', event.error);
        currentIndex++;
        speakNext(); // Skip failed chunk
      }
    };

    speechSynthesis.speak(utterance);
  }

  speakNext();

  // Return control object
  return {
    pause() {
      isPaused = true;
      speechSynthesis.pause();
    },
    resume() {
      isPaused = false;
      speechSynthesis.resume();
    },
    cancel() {
      isCancelled = true;
      speechSynthesis.cancel();
    },
    get progress() {
      return currentIndex / chunks.length;
    }
  };
}

SSML and Advanced Speech Control

SSML (Speech Synthesis Markup Language) is an XML-based language that provides fine-grained control over speech: pauses, emphasis, pronunciation, prosody changes, phonetic spelling, and more.

SSML Is Not Supported in the Web Speech API

Despite being mentioned in the W3C specification, no browser implements SSML parsing for the SpeechSynthesis API as of February 2026. If you pass SSML tags in the utterance text, the browser will speak the tags literally as text.

Simulating SSML Features

You can approximate some SSML features by creating multiple utterances with different parameters.

javascript — Simulating Pauses and Emphasis
// Simulate a pause by inserting a short silent utterance
function speakWithPause(before, after, pauseMs = 500) {
  const u1 = new SpeechSynthesisUtterance(before);
  const pause = new SpeechSynthesisUtterance('');
  const u2 = new SpeechSynthesisUtterance(after);

  // The pause duration is approximate
  u1.onend = () => {
    setTimeout(() => speechSynthesis.speak(u2), pauseMs);
  };

  speechSynthesis.speak(u1);
}

// Simulate emphasis by changing rate and pitch
function speakWithEmphasis(normal, emphasized, afterEmphasis) {
  const u1 = new SpeechSynthesisUtterance(normal);

  const u2 = new SpeechSynthesisUtterance(emphasized);
  u2.rate = 0.85;   // Slightly slower
  u2.pitch = 1.15;  // Slightly higher pitch
  u2.volume = 1;    // Full volume

  const u3 = new SpeechSynthesisUtterance(afterEmphasis);

  speechSynthesis.speak(u1);
  speechSynthesis.speak(u2);
  speechSynthesis.speak(u3);
}

// Example: "Please do NOT press the red button"
speakWithEmphasis('Please do ', 'NOT', ' press the red button.');

Cloud TTS with SSML Support

If you need full SSML support, use a cloud-based TTS API. These accept SSML input and return audio files.

Service SSML Support Free Tier
Google Cloud TTS Full SSML + custom lexicons 1M characters/month
Amazon Polly Full SSML + Neural voices 5M characters/month (12 months)
Azure Cognitive Services Full SSML + Viseme data 500K characters/month
ElevenLabs Partial (prosody only) 10K characters/month

Accessibility Use Cases

The SpeechSynthesis API is a powerful accessibility tool when used correctly. Here are the most impactful use cases and the patterns to implement them well.

1. Read Aloud for Articles

A "Read aloud" button on blog posts and documentation helps users with visual impairments, reading difficulties (dyslexia), or those who simply prefer audio. It also enables multitasking — users can listen while doing something else.

javascript — Article Read-Aloud Button
class ArticleReader {
  constructor(articleSelector) {
    this.article = document.querySelector(articleSelector);
    this.controller = null;
    this.isPlaying = false;
  }

  getArticleText() {
    // Get text content, skip code blocks and nav elements
    const clone = this.article.cloneNode(true);
    clone.querySelectorAll('pre, code, nav, .toc, .cta-box').forEach(el => el.remove());
    return clone.textContent.replace(/\s+/g, ' ').trim();
  }

  play() {
    if (this.isPlaying) {
      speechSynthesis.resume();
      return;
    }

    const text = this.getArticleText();
    this.controller = speakLongText(text, {
      rate: parseFloat(localStorage.getItem('tts-rate') || '1'),
      voice: this.getPreferredVoice()
    });
    this.isPlaying = true;
  }

  pause() {
    if (this.controller) this.controller.pause();
  }

  stop() {
    if (this.controller) this.controller.cancel();
    this.isPlaying = false;
  }

  getPreferredVoice() {
    const savedName = localStorage.getItem('tts-voice');
    if (!savedName) return null;
    return speechSynthesis.getVoices().find(v => v.name === savedName);
  }
}

2. Form Validation Feedback

javascript — Speak Validation Errors
function announceError(fieldName, message) {
  // Visual feedback first
  showVisualError(fieldName, message);

  // Audio feedback (only if user has opted in)
  if (localStorage.getItem('tts-feedback') === 'true') {
    const utterance = new SpeechSynthesisUtterance(
      `Error in ${fieldName}: ${message}`
    );
    utterance.rate = 1.1;
    utterance.volume = 0.8;
    speechSynthesis.speak(utterance);
  }
}

// Usage
announceError('email', 'Please enter a valid email address');

3. Live Content Announcements

javascript — Announce Dynamic Updates
// For notifications, chat messages, stock tickers, etc.
function announceUpdate(message, priority = 'polite') {
  // ARIA live region (works with screen readers)
  const liveRegion = document.getElementById('live-announcements');
  liveRegion.setAttribute('aria-live', priority);
  liveRegion.textContent = message;

  // Speech synthesis (for users who have enabled it)
  if (userPrefersAudioFeedback()) {
    const utterance = new SpeechSynthesisUtterance(message);
    utterance.rate = 1.2;
    speechSynthesis.speak(utterance);
  }
}
Accessibility Best Practices

Always make speech features opt-in. Never auto-play speech on page load. Provide visible controls to pause, resume, and stop. Let users choose their preferred voice and speed. Store preferences in localStorage. Ensure the feature works alongside screen readers rather than replacing them.

Count words and estimate reading time for your content with QTool's Word Counter — useful for calculating speech duration before users press play.

Cross-Browser Quirks and Fixes

The Web Speech API has excellent browser coverage but inconsistent behavior. Here are the issues you will encounter and how to work around them.

1. Chrome: Speech Stops After ~15 Seconds

Chrome cancels utterances that run too long. Use the chunking approach from the Handling Long Text section to split text into sentence-length utterances.

2. iOS Safari: Requires User Gesture

iOS Safari blocks speechSynthesis.speak() unless triggered by a user tap. Additionally, you need to "unlock" the speech engine with an initial call.

javascript — iOS Safari Unlock
// Call this on the first user interaction to unlock iOS speech
function unlockSpeechSynthesis() {
  const utterance = new SpeechSynthesisUtterance('');
  speechSynthesis.speak(utterance);
}

// Add to first click/tap handler
document.addEventListener('click', unlockSpeechSynthesis, { once: true });

3. Firefox: No Boundary Events

Firefox does not fire boundary events consistently. If you rely on word highlighting, provide a visual-only fallback (e.g., a progress bar based on elapsed time) for Firefox users.

4. Safari: Voice List Timing

Safari sometimes does not fire the voiceschanged event. Use a polling fallback.

javascript — Reliable Voice Loading (All Browsers)
function getVoicesReliable(timeout = 3000) {
  return new Promise((resolve) => {
    let voices = speechSynthesis.getVoices();
    if (voices.length > 0) {
      resolve(voices);
      return;
    }

    // Listen for the event
    const handler = () => {
      voices = speechSynthesis.getVoices();
      if (voices.length > 0) {
        clearInterval(pollId);
        clearTimeout(timeoutId);
        resolve(voices);
      }
    };

    speechSynthesis.addEventListener('voiceschanged', handler, { once: true });

    // Poll as fallback (Safari)
    const pollId = setInterval(() => {
      voices = speechSynthesis.getVoices();
      if (voices.length > 0) {
        clearInterval(pollId);
        clearTimeout(timeoutId);
        speechSynthesis.removeEventListener('voiceschanged', handler);
        resolve(voices);
      }
    }, 100);

    // Timeout fallback
    const timeoutId = setTimeout(() => {
      clearInterval(pollId);
      speechSynthesis.removeEventListener('voiceschanged', handler);
      resolve(speechSynthesis.getVoices()); // Return whatever we have
    }, timeout);
  });
}

5. All Browsers: Queue Management

Speech utterances are queued. If you call speak() multiple times without cancel(), all utterances play sequentially. Always cancel before speaking new content if you want to interrupt.

javascript — Interrupt Current Speech
function speakImmediate(text, options = {}) {
  // Cancel anything currently speaking
  speechSynthesis.cancel();

  const utterance = new SpeechSynthesisUtterance(text);
  Object.assign(utterance, options);
  speechSynthesis.speak(utterance);
}

Try Text-to-Speech in Your Browser

QTool's free Text-to-Speech tool lets you paste any text, choose a voice, adjust speed and pitch, and listen instantly. No signup needed.

Open Text-to-Speech Speech to Text

Complete Read-Aloud Component

Here is a production-ready read-aloud component that handles all the cross-browser quirks, chunking, and user preferences.

javascript — Production Read-Aloud Component
class ReadAloud {
  #state = 'idle'; // idle | playing | paused
  #chunks = [];
  #currentChunk = 0;
  #voice = null;
  #rate = 1;
  #pitch = 1;

  constructor(options = {}) {
    this.onStateChange = options.onStateChange || (() => {});
    this.onProgress = options.onProgress || (() => {});
    this.onWord = options.onWord || (() => {});

    // Load saved preferences
    this.#rate = parseFloat(localStorage.getItem('ra-rate') || '1');
    this.#pitch = parseFloat(localStorage.getItem('ra-pitch') || '1');

    // Preload voices
    getVoicesReliable().then(voices => {
      const savedVoice = localStorage.getItem('ra-voice');
      this.#voice = voices.find(v => v.name === savedVoice) || null;
    });
  }

  speak(text) {
    this.stop();
    // Split on sentence boundaries, keep punctuation
    this.#chunks = text.match(/[^.!?\n]+[.!?\n]+|[^.!?\n]+$/g) || [text];
    this.#currentChunk = 0;
    this.#state = 'playing';
    this.onStateChange(this.#state);
    this.#speakNext();
  }

  #speakNext() {
    if (this.#state !== 'playing' || this.#currentChunk >= this.#chunks.length) {
      if (this.#currentChunk >= this.#chunks.length) {
        this.#state = 'idle';
        this.onStateChange(this.#state);
      }
      return;
    }

    const text = this.#chunks[this.#currentChunk].trim();
    if (!text) {
      this.#currentChunk++;
      this.#speakNext();
      return;
    }

    const utterance = new SpeechSynthesisUtterance(text);
    if (this.#voice) utterance.voice = this.#voice;
    utterance.rate = this.#rate;
    utterance.pitch = this.#pitch;

    utterance.onboundary = (e) => {
      if (e.name === 'word') this.onWord(e.charIndex, e.charLength);
    };

    utterance.onend = () => {
      this.#currentChunk++;
      this.onProgress(this.#currentChunk / this.#chunks.length);
      this.#speakNext();
    };

    utterance.onerror = (e) => {
      if (e.error !== 'canceled') {
        this.#currentChunk++;
        this.#speakNext();
      }
    };

    speechSynthesis.speak(utterance);
  }

  pause() {
    speechSynthesis.pause();
    this.#state = 'paused';
    this.onStateChange(this.#state);
  }

  resume() {
    speechSynthesis.resume();
    this.#state = 'playing';
    this.onStateChange(this.#state);
  }

  stop() {
    speechSynthesis.cancel();
    this.#chunks = [];
    this.#currentChunk = 0;
    this.#state = 'idle';
    this.onStateChange(this.#state);
  }

  setRate(rate) {
    this.#rate = Math.max(0.1, Math.min(10, rate));
    localStorage.setItem('ra-rate', String(this.#rate));
  }

  setVoice(voiceName) {
    const voices = speechSynthesis.getVoices();
    this.#voice = voices.find(v => v.name === voiceName) || null;
    if (this.#voice) localStorage.setItem('ra-voice', this.#voice.name);
  }

  get state() { return this.#state; }
  get progress() { return this.#chunks.length ? this.#currentChunk / this.#chunks.length : 0; }
}

Speech and Text Tools


Frequently Asked Questions

The Web Speech API is a browser-native JavaScript API that provides two capabilities: SpeechSynthesis (text-to-speech) and SpeechRecognition (speech-to-text). The SpeechSynthesis API converts text into spoken audio using voices provided by the operating system or browser. It is supported in Chrome 33+, Firefox 49+, Safari 7+, Edge 14+, and all Chromium-based browsers including Opera and Brave. Mobile support includes iOS Safari 7+ and Android Chrome 33+. No plugins, downloads, or API keys are required.

Use window.speechSynthesis.getVoices() to get an array of SpeechSynthesisVoice objects. Important: on many browsers (especially Chrome), the voices list is loaded asynchronously and getVoices() returns an empty array on the first call. You must listen for the voiceschanged event. Each voice object has properties: name (display name), lang (BCP 47 language tag like "en-US"), localService (boolean, true if the voice runs locally), and voiceURI (unique identifier). The number of available voices depends on the operating system: macOS typically has 60-80, Windows 10+ has 20-40.

Yes. The SpeechSynthesisUtterance object has three properties for controlling speech output: rate (speed, range 0.1 to 10, default 1), pitch (range 0 to 2, default 1), and volume (range 0 to 1, default 1). A rate of 1.5 to 2.0 is commonly used for speed reading features. Pitch values below 1 produce a deeper voice, above 1 a higher voice. These properties work consistently across all browsers that support the SpeechSynthesis API, though the exact audio quality varies by voice and platform.

The Web Speech API can enhance accessibility by reading page content aloud for users with visual impairments or reading difficulties, providing audio feedback for form validation errors, announcing dynamic content changes, and offering a "read aloud" button for long-form articles. When implementing, always make speech features opt-in (never auto-play speech), provide visible controls to pause, resume, and stop, respect the user's preferred voice and rate settings, and ensure the feature works alongside screen readers rather than replacing them.

SSML (Speech Synthesis Markup Language) is an XML-based markup language that gives fine-grained control over speech output, including pauses, emphasis, pronunciation, and prosody. The Web Speech API's SpeechSynthesis does NOT support SSML natively in any browser as of 2026. If you need SSML support in the browser, use a cloud-based TTS service (Google Cloud Text-to-Speech, Amazon Polly, Azure Cognitive Services) that accepts SSML input and returns audio, or simulate basic features by splitting text and creating multiple utterances with different parameters.

Mobile browsers require a user gesture (tap, click) before allowing speech synthesis. This is a security measure to prevent websites from auto-playing audio. If you call speechSynthesis.speak() without a preceding user interaction, it will be silently blocked. The fix is to always trigger speech from a click or tap event handler. On iOS Safari specifically, you must call speechSynthesis.speak() at least once (even with an empty utterance) in a user gesture handler before subsequent calls will work. A common pattern is to "unlock" the audio context on the first user interaction.

NT

Christian Bucher

We build free developer tools including text-to-speech, speech-to-text, and text analysis utilities. 269 tool pages, all browser-based, no signup required.

269 Developer Tools, One Place

Browse 269 indexed tool pages with no QTool account required, and inspect the source on GitHub.

Open Source — Free Forever Try Free Tools

Related Tools

CSS Box Shadow Generator · Emoji Picker & Search · Free JSON to YAML Converter

Related Tools

Free JSON to YAML Converter · YAML to JSON Converter · Free API Mock Server

Related Articles

Built by Miguel

Need a custom tool or website?

From . Delivered in 24-48h. You own the code.

View Services →