What the Web Speech API Does
The Web Speech API is a browser-native JavaScript API with two halves: SpeechSynthesis (text-to-speech) and SpeechRecognition (speech-to-text). This guide focuses on SpeechSynthesis — converting text into spoken audio using voices provided by the operating system.
No API keys. No external services. No npm packages. The browser talks using the same speech engine that powers your OS's built-in screen reader and voice assistant. All processing happens on the client device, which means zero latency, zero cost, and full privacy.
Browser Support
| Browser | SpeechSynthesis | SpeechRecognition |
|---|---|---|
| Chrome 33+ | Full support | Full support (webkitSpeechRecognition) |
| Firefox 49+ | Full support | Limited (flag required) |
| Safari 7+ | Full support | 14.1+ (partial) |
| Edge 14+ | Full support | Full support (Chromium-based) |
| iOS Safari | 7+ (with quirks) | 14.5+ (partial) |
| Android Chrome | 33+ (with quirks) | Full support |
Try the API right now with QTool's Text-to-Speech tool. Paste any text, select a voice, and hear it spoken — all in your browser.
Basic Text-to-Speech in 5 Lines
// Create an utterance (the thing to speak)
const utterance = new SpeechSynthesisUtterance('Hello, world!');
// Speak it
speechSynthesis.speak(utterance);
// That's it. The browser is now talking.
The SpeechSynthesisUtterance object represents a single piece of text to be spoken. The speechSynthesis object (available on window) is the controller that manages the speech queue, voice selection, and playback state.
// Pause speech
speechSynthesis.pause();
// Resume speech
speechSynthesis.resume();
// Stop speech (clears the queue)
speechSynthesis.cancel();
// Check if currently speaking
console.log(speechSynthesis.speaking); // true or false
// Check if paused
console.log(speechSynthesis.paused); // true or false
// Check if there are pending utterances
console.log(speechSynthesis.pending); // true or false
Voice Selection and Management
Every operating system comes with a different set of voices. macOS typically provides 60-80 voices across 40+ languages. Windows 10/11 has 20-40 voices. Mobile platforms have 30-50. Some browsers also offer cloud-based voices with higher quality (Chrome's Google voices, for example).
Loading Voices (The Async Trap)
On most browsers, speechSynthesis.getVoices() returns an empty array on the first call. The voice list is loaded asynchronously. You must listen for the voiceschanged event.
function getVoices() {
return new Promise((resolve) => {
let voices = speechSynthesis.getVoices();
if (voices.length > 0) {
resolve(voices);
return;
}
// Voices not loaded yet. Wait for the event.
speechSynthesis.addEventListener('voiceschanged', () => {
voices = speechSynthesis.getVoices();
resolve(voices);
}, { once: true });
});
}
// Usage
const voices = await getVoices();
console.log(`Found ${voices.length} voices`);
// List all English voices
const englishVoices = voices.filter(v => v.lang.startsWith('en'));
englishVoices.forEach(v => {
console.log(`${v.name} (${v.lang}) ${v.localService ? '[local]' : '[cloud]'}`);
});
Selecting a Voice
const voices = await getVoices();
// Find a specific voice by name
const preferred = voices.find(v => v.name === 'Google UK English Female')
|| voices.find(v => v.name.includes('Samantha'))
|| voices.find(v => v.lang === 'en-US');
const utterance = new SpeechSynthesisUtterance('Hello from a specific voice.');
utterance.voice = preferred;
speechSynthesis.speak(utterance);
Voice Object Properties
| Property | Type | Description |
|---|---|---|
name |
string | Human-readable voice name (e.g., "Google UK English Female") |
lang |
string | BCP 47 language tag (e.g., "en-US", "de-DE", "ja-JP") |
localService |
boolean | true = runs locally (offline-capable), false = cloud-based |
voiceURI |
string | Unique identifier for the voice |
default |
boolean | true if this is the browser's default voice for the language |
Need to analyze the text before speaking it? QTool's Text Analyzer gives you word count, character count, reading time, and readability scores instantly.
Controlling Rate, Pitch, and Volume
const utterance = new SpeechSynthesisUtterance('This demonstrates rate, pitch, and volume control.');
// Rate: 0.1 (very slow) to 10 (very fast). Default: 1
utterance.rate = 1.3; // Slightly faster than normal
// Pitch: 0 (lowest) to 2 (highest). Default: 1
utterance.pitch = 0.9; // Slightly deeper voice
// Volume: 0 (silent) to 1 (loudest). Default: 1
utterance.volume = 0.8;
// Language (overrides voice's default language)
utterance.lang = 'en-US';
speechSynthesis.speak(utterance);
Practical Rate Settings
| Rate | Use Case |
|---|---|
| 0.7 - 0.9 | Accessibility: users who need slower speech |
| 1.0 | Default: natural conversational speed |
| 1.2 - 1.5 | Speed reading: experienced listeners, podcast-style |
| 1.5 - 2.0 | Power users: familiar with accelerated speech |
| 2.0+ | Screen reader users: trained to process fast speech |
Event Handling and Progress Tracking
SpeechSynthesisUtterance fires events at key moments during speech. These are essential for building a proper UI with progress indicators, word highlighting, and state management.
const utterance = new SpeechSynthesisUtterance('The quick brown fox jumps over the lazy dog.');
// Fired when speech starts
utterance.onstart = (event) => {
console.log('Started speaking');
};
// Fired when speech ends normally
utterance.onend = (event) => {
console.log(`Finished. Elapsed: ${event.elapsedTime}ms`);
};
// Fired when speech is paused
utterance.onpause = (event) => {
console.log('Paused at character:', event.charIndex);
};
// Fired when speech resumes after pause
utterance.onresume = (event) => {
console.log('Resumed');
};
// Fired at word boundaries (not all browsers)
utterance.onboundary = (event) => {
if (event.name === 'word') {
const word = utterance.text.substring(
event.charIndex,
event.charIndex + event.charLength
);
console.log(`Word: "${word}" at char ${event.charIndex}`);
// Use this for word-by-word highlighting
}
};
// Fired on error
utterance.onerror = (event) => {
console.error('Speech error:', event.error);
// Common errors: 'canceled', 'interrupted', 'network', 'synthesis-failed'
};
speechSynthesis.speak(utterance);
Word Highlighting
The boundary event provides the character index and length of each word as it is spoken. You can use this to highlight the current word in the UI.
function speakWithHighlighting(text, containerEl) {
// Wrap each word in a span
const words = text.split(/(\s+)/);
containerEl.innerHTML = words.map((word, i) =>
word.trim() ? `<span data-index="${i}">${word}</span>` : word
).join('');
const utterance = new SpeechSynthesisUtterance(text);
utterance.onboundary = (event) => {
if (event.name !== 'word') return;
// Remove previous highlight
containerEl.querySelectorAll('.speaking').forEach(el =>
el.classList.remove('speaking')
);
// Find the span containing this character index
const charIndex = event.charIndex;
let currentIndex = 0;
for (const span of containerEl.querySelectorAll('span[data-index]')) {
const end = currentIndex + span.textContent.length;
if (charIndex >= currentIndex && charIndex < end) {
span.classList.add('speaking');
span.scrollIntoView({ behavior: 'smooth', block: 'center' });
break;
}
currentIndex = end + 1; // +1 for the space
}
};
utterance.onend = () => {
containerEl.querySelectorAll('.speaking').forEach(el =>
el.classList.remove('speaking')
);
};
speechSynthesis.speak(utterance);
}
// CSS: .speaking { background: rgba(0, 212, 255, 0.3); border-radius: 3px; }
The boundary event fires reliably in Chrome and Edge (Chromium) but is less consistent in Firefox and Safari. Always design your UI to work without word highlighting as a fallback, and test on your target browsers.
Handling Long Text and Chunking
Chrome has a known issue: speech stops after approximately 200-300 characters (or about 15 seconds) of continuous speech. This is a bug that has persisted for years. The workaround is to split long text into chunks and speak them sequentially.
function speakLongText(text, options = {}) {
const { voice, rate = 1, pitch = 1, volume = 1 } = options;
// Split on sentence boundaries
const chunks = text.match(/[^.!?]+[.!?]+/g) || [text];
let currentIndex = 0;
let isPaused = false;
let isCancelled = false;
function speakNext() {
if (isCancelled || currentIndex >= chunks.length) return;
const utterance = new SpeechSynthesisUtterance(chunks[currentIndex].trim());
if (voice) utterance.voice = voice;
utterance.rate = rate;
utterance.pitch = pitch;
utterance.volume = volume;
utterance.onend = () => {
currentIndex++;
if (!isPaused && !isCancelled) {
speakNext();
}
};
utterance.onerror = (event) => {
if (event.error !== 'canceled') {
console.error('Chunk error:', event.error);
currentIndex++;
speakNext(); // Skip failed chunk
}
};
speechSynthesis.speak(utterance);
}
speakNext();
// Return control object
return {
pause() {
isPaused = true;
speechSynthesis.pause();
},
resume() {
isPaused = false;
speechSynthesis.resume();
},
cancel() {
isCancelled = true;
speechSynthesis.cancel();
},
get progress() {
return currentIndex / chunks.length;
}
};
}
SSML and Advanced Speech Control
SSML (Speech Synthesis Markup Language) is an XML-based language that provides fine-grained control over speech: pauses, emphasis, pronunciation, prosody changes, phonetic spelling, and more.
Despite being mentioned in the W3C specification, no browser implements SSML parsing for the SpeechSynthesis API as of February 2026. If you pass SSML tags in the utterance text, the browser will speak the tags literally as text.
Simulating SSML Features
You can approximate some SSML features by creating multiple utterances with different parameters.
// Simulate a pause by inserting a short silent utterance
function speakWithPause(before, after, pauseMs = 500) {
const u1 = new SpeechSynthesisUtterance(before);
const pause = new SpeechSynthesisUtterance('');
const u2 = new SpeechSynthesisUtterance(after);
// The pause duration is approximate
u1.onend = () => {
setTimeout(() => speechSynthesis.speak(u2), pauseMs);
};
speechSynthesis.speak(u1);
}
// Simulate emphasis by changing rate and pitch
function speakWithEmphasis(normal, emphasized, afterEmphasis) {
const u1 = new SpeechSynthesisUtterance(normal);
const u2 = new SpeechSynthesisUtterance(emphasized);
u2.rate = 0.85; // Slightly slower
u2.pitch = 1.15; // Slightly higher pitch
u2.volume = 1; // Full volume
const u3 = new SpeechSynthesisUtterance(afterEmphasis);
speechSynthesis.speak(u1);
speechSynthesis.speak(u2);
speechSynthesis.speak(u3);
}
// Example: "Please do NOT press the red button"
speakWithEmphasis('Please do ', 'NOT', ' press the red button.');
Cloud TTS with SSML Support
If you need full SSML support, use a cloud-based TTS API. These accept SSML input and return audio files.
| Service | SSML Support | Free Tier |
|---|---|---|
| Google Cloud TTS | Full SSML + custom lexicons | 1M characters/month |
| Amazon Polly | Full SSML + Neural voices | 5M characters/month (12 months) |
| Azure Cognitive Services | Full SSML + Viseme data | 500K characters/month |
| ElevenLabs | Partial (prosody only) | 10K characters/month |
Accessibility Use Cases
The SpeechSynthesis API is a powerful accessibility tool when used correctly. Here are the most impactful use cases and the patterns to implement them well.
1. Read Aloud for Articles
A "Read aloud" button on blog posts and documentation helps users with visual impairments, reading difficulties (dyslexia), or those who simply prefer audio. It also enables multitasking — users can listen while doing something else.
class ArticleReader {
constructor(articleSelector) {
this.article = document.querySelector(articleSelector);
this.controller = null;
this.isPlaying = false;
}
getArticleText() {
// Get text content, skip code blocks and nav elements
const clone = this.article.cloneNode(true);
clone.querySelectorAll('pre, code, nav, .toc, .cta-box').forEach(el => el.remove());
return clone.textContent.replace(/\s+/g, ' ').trim();
}
play() {
if (this.isPlaying) {
speechSynthesis.resume();
return;
}
const text = this.getArticleText();
this.controller = speakLongText(text, {
rate: parseFloat(localStorage.getItem('tts-rate') || '1'),
voice: this.getPreferredVoice()
});
this.isPlaying = true;
}
pause() {
if (this.controller) this.controller.pause();
}
stop() {
if (this.controller) this.controller.cancel();
this.isPlaying = false;
}
getPreferredVoice() {
const savedName = localStorage.getItem('tts-voice');
if (!savedName) return null;
return speechSynthesis.getVoices().find(v => v.name === savedName);
}
}
2. Form Validation Feedback
function announceError(fieldName, message) {
// Visual feedback first
showVisualError(fieldName, message);
// Audio feedback (only if user has opted in)
if (localStorage.getItem('tts-feedback') === 'true') {
const utterance = new SpeechSynthesisUtterance(
`Error in ${fieldName}: ${message}`
);
utterance.rate = 1.1;
utterance.volume = 0.8;
speechSynthesis.speak(utterance);
}
}
// Usage
announceError('email', 'Please enter a valid email address');
3. Live Content Announcements
// For notifications, chat messages, stock tickers, etc.
function announceUpdate(message, priority = 'polite') {
// ARIA live region (works with screen readers)
const liveRegion = document.getElementById('live-announcements');
liveRegion.setAttribute('aria-live', priority);
liveRegion.textContent = message;
// Speech synthesis (for users who have enabled it)
if (userPrefersAudioFeedback()) {
const utterance = new SpeechSynthesisUtterance(message);
utterance.rate = 1.2;
speechSynthesis.speak(utterance);
}
}
Always make speech features opt-in. Never auto-play speech on page load. Provide visible controls to pause, resume, and stop. Let users choose their preferred voice and speed. Store preferences in localStorage. Ensure the feature works alongside screen readers rather than replacing them.
Count words and estimate reading time for your content with QTool's Word Counter — useful for calculating speech duration before users press play.
Cross-Browser Quirks and Fixes
The Web Speech API has excellent browser coverage but inconsistent behavior. Here are the issues you will encounter and how to work around them.
1. Chrome: Speech Stops After ~15 Seconds
Chrome cancels utterances that run too long. Use the chunking approach from the Handling Long Text section to split text into sentence-length utterances.
2. iOS Safari: Requires User Gesture
iOS Safari blocks speechSynthesis.speak() unless triggered by a user tap. Additionally, you need to "unlock" the speech engine with an initial call.
// Call this on the first user interaction to unlock iOS speech
function unlockSpeechSynthesis() {
const utterance = new SpeechSynthesisUtterance('');
speechSynthesis.speak(utterance);
}
// Add to first click/tap handler
document.addEventListener('click', unlockSpeechSynthesis, { once: true });
3. Firefox: No Boundary Events
Firefox does not fire boundary events consistently. If you rely on word highlighting, provide a visual-only fallback (e.g., a progress bar based on elapsed time) for Firefox users.
4. Safari: Voice List Timing
Safari sometimes does not fire the voiceschanged event. Use a polling fallback.
function getVoicesReliable(timeout = 3000) {
return new Promise((resolve) => {
let voices = speechSynthesis.getVoices();
if (voices.length > 0) {
resolve(voices);
return;
}
// Listen for the event
const handler = () => {
voices = speechSynthesis.getVoices();
if (voices.length > 0) {
clearInterval(pollId);
clearTimeout(timeoutId);
resolve(voices);
}
};
speechSynthesis.addEventListener('voiceschanged', handler, { once: true });
// Poll as fallback (Safari)
const pollId = setInterval(() => {
voices = speechSynthesis.getVoices();
if (voices.length > 0) {
clearInterval(pollId);
clearTimeout(timeoutId);
speechSynthesis.removeEventListener('voiceschanged', handler);
resolve(voices);
}
}, 100);
// Timeout fallback
const timeoutId = setTimeout(() => {
clearInterval(pollId);
speechSynthesis.removeEventListener('voiceschanged', handler);
resolve(speechSynthesis.getVoices()); // Return whatever we have
}, timeout);
});
}
5. All Browsers: Queue Management
Speech utterances are queued. If you call speak() multiple times without cancel(), all utterances play sequentially. Always cancel before speaking new content if you want to interrupt.
function speakImmediate(text, options = {}) {
// Cancel anything currently speaking
speechSynthesis.cancel();
const utterance = new SpeechSynthesisUtterance(text);
Object.assign(utterance, options);
speechSynthesis.speak(utterance);
}
Try Text-to-Speech in Your Browser
QTool's free Text-to-Speech tool lets you paste any text, choose a voice, adjust speed and pitch, and listen instantly. No signup needed.
Open Text-to-Speech Speech to TextComplete Read-Aloud Component
Here is a production-ready read-aloud component that handles all the cross-browser quirks, chunking, and user preferences.
class ReadAloud {
#state = 'idle'; // idle | playing | paused
#chunks = [];
#currentChunk = 0;
#voice = null;
#rate = 1;
#pitch = 1;
constructor(options = {}) {
this.onStateChange = options.onStateChange || (() => {});
this.onProgress = options.onProgress || (() => {});
this.onWord = options.onWord || (() => {});
// Load saved preferences
this.#rate = parseFloat(localStorage.getItem('ra-rate') || '1');
this.#pitch = parseFloat(localStorage.getItem('ra-pitch') || '1');
// Preload voices
getVoicesReliable().then(voices => {
const savedVoice = localStorage.getItem('ra-voice');
this.#voice = voices.find(v => v.name === savedVoice) || null;
});
}
speak(text) {
this.stop();
// Split on sentence boundaries, keep punctuation
this.#chunks = text.match(/[^.!?\n]+[.!?\n]+|[^.!?\n]+$/g) || [text];
this.#currentChunk = 0;
this.#state = 'playing';
this.onStateChange(this.#state);
this.#speakNext();
}
#speakNext() {
if (this.#state !== 'playing' || this.#currentChunk >= this.#chunks.length) {
if (this.#currentChunk >= this.#chunks.length) {
this.#state = 'idle';
this.onStateChange(this.#state);
}
return;
}
const text = this.#chunks[this.#currentChunk].trim();
if (!text) {
this.#currentChunk++;
this.#speakNext();
return;
}
const utterance = new SpeechSynthesisUtterance(text);
if (this.#voice) utterance.voice = this.#voice;
utterance.rate = this.#rate;
utterance.pitch = this.#pitch;
utterance.onboundary = (e) => {
if (e.name === 'word') this.onWord(e.charIndex, e.charLength);
};
utterance.onend = () => {
this.#currentChunk++;
this.onProgress(this.#currentChunk / this.#chunks.length);
this.#speakNext();
};
utterance.onerror = (e) => {
if (e.error !== 'canceled') {
this.#currentChunk++;
this.#speakNext();
}
};
speechSynthesis.speak(utterance);
}
pause() {
speechSynthesis.pause();
this.#state = 'paused';
this.onStateChange(this.#state);
}
resume() {
speechSynthesis.resume();
this.#state = 'playing';
this.onStateChange(this.#state);
}
stop() {
speechSynthesis.cancel();
this.#chunks = [];
this.#currentChunk = 0;
this.#state = 'idle';
this.onStateChange(this.#state);
}
setRate(rate) {
this.#rate = Math.max(0.1, Math.min(10, rate));
localStorage.setItem('ra-rate', String(this.#rate));
}
setVoice(voiceName) {
const voices = speechSynthesis.getVoices();
this.#voice = voices.find(v => v.name === voiceName) || null;
if (this.#voice) localStorage.setItem('ra-voice', this.#voice.name);
}
get state() { return this.#state; }
get progress() { return this.#chunks.length ? this.#currentChunk / this.#chunks.length : 0; }
}
Speech and Text Tools
Frequently Asked Questions
The Web Speech API is a browser-native JavaScript API that provides two capabilities: SpeechSynthesis (text-to-speech) and SpeechRecognition (speech-to-text). The SpeechSynthesis API converts text into spoken audio using voices provided by the operating system or browser. It is supported in Chrome 33+, Firefox 49+, Safari 7+, Edge 14+, and all Chromium-based browsers including Opera and Brave. Mobile support includes iOS Safari 7+ and Android Chrome 33+. No plugins, downloads, or API keys are required.
Use window.speechSynthesis.getVoices() to get an array of SpeechSynthesisVoice objects. Important: on many browsers (especially Chrome), the voices list is loaded asynchronously and getVoices() returns an empty array on the first call. You must listen for the voiceschanged event. Each voice object has properties: name (display name), lang (BCP 47 language tag like "en-US"), localService (boolean, true if the voice runs locally), and voiceURI (unique identifier). The number of available voices depends on the operating system: macOS typically has 60-80, Windows 10+ has 20-40.
Yes. The SpeechSynthesisUtterance object has three properties for controlling speech output: rate (speed, range 0.1 to 10, default 1), pitch (range 0 to 2, default 1), and volume (range 0 to 1, default 1). A rate of 1.5 to 2.0 is commonly used for speed reading features. Pitch values below 1 produce a deeper voice, above 1 a higher voice. These properties work consistently across all browsers that support the SpeechSynthesis API, though the exact audio quality varies by voice and platform.
The Web Speech API can enhance accessibility by reading page content aloud for users with visual impairments or reading difficulties, providing audio feedback for form validation errors, announcing dynamic content changes, and offering a "read aloud" button for long-form articles. When implementing, always make speech features opt-in (never auto-play speech), provide visible controls to pause, resume, and stop, respect the user's preferred voice and rate settings, and ensure the feature works alongside screen readers rather than replacing them.
SSML (Speech Synthesis Markup Language) is an XML-based markup language that gives fine-grained control over speech output, including pauses, emphasis, pronunciation, and prosody. The Web Speech API's SpeechSynthesis does NOT support SSML natively in any browser as of 2026. If you need SSML support in the browser, use a cloud-based TTS service (Google Cloud Text-to-Speech, Amazon Polly, Azure Cognitive Services) that accepts SSML input and returns audio, or simulate basic features by splitting text and creating multiple utterances with different parameters.
Mobile browsers require a user gesture (tap, click) before allowing speech synthesis. This is a security measure to prevent websites from auto-playing audio. If you call speechSynthesis.speak() without a preceding user interaction, it will be silently blocked. The fix is to always trigger speech from a click or tap event handler. On iOS Safari specifically, you must call speechSynthesis.speak() at least once (even with an empty utterance) in a user gesture handler before subsequent calls will work. A common pattern is to "unlock" the audio context on the first user interaction.