SpeechSynthesisUtterance: Speech synthesis with native JavaScript | Complete tutorial

- Andrés Cruz - ES En español

SpeechSynthesisUtterance: Speech synthesis with native JavaScript | Complete tutorial

See demo

Ever since assistants like Siri, Google Assistant, or Cortana became part of our daily lives, speech synthesis stopped being a curiosity and turned into a real, expected feature in many web applications. What's interesting is that, thanks to HTML5's Web Speech API, JavaScript provides us with a native API to work with speech synthesis (text-to-speech or TTS) easily, without needing external libraries or third-party dependencies.

The native speech synthesizer in JavaScript allows any text to be played directly in the browser, with control over language, pitch, and speed. What surprised me the most when I used it for the first time was how little it takes to have a working example: barely two lines of code.

How to Get Started with the Speech Synthesis API in JavaScript

Create Your First SpeechSynthesisUtterance Object

The SpeechSynthesisUtterance class is the core of all this: it captures text and prepares it to be converted into audio. Creating an object and starting it is as straightforward as this:

var speechMessage = new SpeechSynthesisUtterance('Hola, mundo!');
window.speechSynthesis.speak(speechMessage);

This minimal example works right away in modern browsers. First, we create an instance of SpeechSynthesisUtterance passing the text as a parameter, and then we process it through the window.speechSynthesis interface, which actually gives voice to the browser by invoking the speak() method.

Play Dynamic Text: Basic Example

To read text coming from a variable instead of a fixed literal, simply assign it when creating the instance:

var texto = "Bienvenido a mi aplicación web";
var speechMessage = new SpeechSynthesisUtterance(texto);
window.speechSynthesis.speak(speechMessage);

This opens the door to interactive applications where content can read personalized messages to the user, such as notifications, form results, or any text generated in real-time.

Essential Properties of SpeechSynthesisUtterance

The SpeechSynthesisUtterance class exposes a set of properties that allow you to customize exactly how the browser will speak. Among the most used are:

SpeechSynthesisUtterance.text

It is the most important property of all. It allows setting or getting the text that the browser will read aloud. You can also assign it directly in the constructor or modify it later:

speechMessage.text = "Texto actualizado dinámicamente";

SpeechSynthesisUtterance.lang

Allows specifying the language of the synthesis using BCP 47 standard language tags. In my tests, changing the language noticeably improves the naturalness of pronunciation:

speechMessage.lang = 'es-ES'; // Spanish from Spain
speechMessage.lang = 'es-MX'; // Spanish from Mexico
speechMessage.lang = 'en-US'; // American English

SpeechSynthesisUtterance.volume

Controls the voice volume. Accepts a float value between 0 (silence) and 1 (maximum). Useful for balancing audio with other sound effects on the page:

speechMessage.volume = 0.8; // 80% of maximum volume

SpeechSynthesisUtterance.pitch

Sets the pitch of the voice. The value is a float between 0 (lowest) and 2 (highest). The default value is 1:

speechMessage.pitch = 1.2; // slightly higher pitch than normal

SpeechSynthesisUtterance.rate

Controls the speech speed. Accepts values between 0.1 (very slow) and 10 (very fast), with 1 being normal speed. Lowering it slightly to 0.9 can be more comfortable for educational or tutorial content:

speechMessage.rate = 0.9; // slightly slower speed

SpeechSynthesisUtterance.voice

Allows assigning a specific voice from those available in the browser. Voices are fetched asynchronously via window.speechSynthesis.getVoices(). It is important to listen to the onvoiceschanged event because the voice list might not be immediately available:

window.speechSynthesis.onvoiceschanged = function() {
    var voices = window.speechSynthesis.getVoices();
    speechMessage.voice = voices[0]; // selects the first available voice
};

Events and Speech Synthesis Flow Control

The API also exposes events that let you react to different moments in the playback lifecycle. The most common are onstart and onend, useful for syncing animations, visual indicators, or any effect with the voice:

speechMessage.onstart = function(e) {
    console.log('Playback started...');
};

speechMessage.onend = function(e) {
    console.log('Playback finished.');
};

Personally, I used them in an interactive tutorial project to trigger an animated icon while the browser was speaking, and the result was great in terms of user experience.

Other Useful Events for Advanced Projects

In addition to onstart and onend, the API exposes other events that allow more granular control over the playback flow:

  • onerror: fires if an error occurs during synthesis. Ideal for handling failures without breaking the user experience.
  • onpause and onresume: detect when playback is paused or resumed, enabling the implementation of play/pause controls in the UI.
  • onboundary: fires at each word or sentence boundary. Very useful for highlighting the text currently being read in real-time (karaoke effect).

You can check the full list of events in the official Web Speech API documentation.

The Other Half of Web Speech API: Speech Recognition

In a previous post, we talked about the other side of the Web Speech API: speech recognition with SpeechRecognition(), which allows your applications to listen to the user through the device's microphone and transcribe what they say according to the configured language:

If speech recognition is the browser's "ear," speech synthesis with SpeechSynthesisUtterance and window.speechSynthesis is its "mouth." Together, they make up the complete Web Speech API suite.

Support in Modern Browsers

The Web Speech API has good support across major modern browsers: Chrome, Firefox, Edge, and Safari implement it, although available voices may vary depending on the operating system and installed language. Before using it, it's a good idea to check support with a simple check:

if ('speechSynthesis' in window) {
    // Browser supports Speech Synthesis API
    console.log('Ready to speak!');
} else {
    console.log('API is not supported in this browser');
}

It is good practice to always include this check before initializing any TTS functionality to avoid errors in browsers that don't implement it yet or have it disabled.

Tips to Improve User Experience

  • Give the user the ability to select from available voices using window.speechSynthesis.getVoices().
  • Adjust volume and rate according to the content type: slower for educational text, faster for short notifications.
  • Avoid overly long texts in a single call to speak(). Some browsers have character limits; split the content into chunks if necessary.
  • Use the onboundary event to visually highlight the text currently being read and improve accessibility.

Complete Speech Synthesis Example in JavaScript

An example encapsulated in a reusable function is the cleanest way to integrate TTS into any project:

function hablar(texto, idioma, velocidad, tono) {
    if (!('speechSynthesis' in window)) {
        console.warn('Speech Synthesis API is not available.');
        return;
    }

    var speechMessage = new SpeechSynthesisUtterance(texto);
    speechMessage.lang   = idioma    || 'es-ES';
    speechMessage.rate   = velocidad || 1;
    speechMessage.pitch  = tono      || 1;
    speechMessage.volume = 1;

    speechMessage.onstart = function() { console.log('Speaking...'); };
    speechMessage.onend   = function() { console.log('Done.'); };

    window.speechSynthesis.speak(speechMessage);
}

// Usage:
hablar("¡Hola! Esto es un ejemplo dinámico.", 'es-ES', 0.9, 1.1);

In my experience, allowing the user to adjust pitch and rate directly from the UI significantly increases accessibility and engagement in educational or interactive tutorial projects.

See live demo

Frequently Asked Questions About Speech Synthesis in JavaScript

  1. Which browsers support the Speech Synthesis API?
    1. Chrome, Firefox, Edge, and Safari support window.speechSynthesis, although available voices and quality may vary depending on the operating system. Chrome on Android and desktop has the broadest support.
  2. How can I change the default language or voice?
    1. Use the lang property of SpeechSynthesisUtterance for the language (e.g., 'es-ES' or 'en-US'), and the voice property with an object obtained from window.speechSynthesis.getVoices() to select a specific voice.
  3. Can I control the speed, volume, and pitch of the voice?
    1. Yes. Use rate for speed (between 0.1 and 10), volume for volume (between 0 and 1), and pitch for pitch (between 0 and 2).
  4. Is it possible to play dynamic text from a variable in JavaScript?
    1. Yes, simply assign the variable to the SpeechSynthesisUtterance constructor or its text property, and call window.speechSynthesis.speak() with the object.
  5. How do I stop or pause speech synthesis in JavaScript?
    1. Use window.speechSynthesis.cancel() to completely stop playback, window.speechSynthesis.pause() to pause it, and window.speechSynthesis.resume() to resume it.

Learn how to use the JavaScript SpeechSynthesisUtterance API to convert text to speech without libraries. Control lang, pitch, rate, volume, and voice with ready-to-use code examples.


Únete a la comunidad de desarrolladores que han decidido dejar de picar código y empezar a construir productos reales. Recibe mis mejores trucos de arquitectura cada semana:

I agree to receive announcements of interest about this Blog.