Ever since assistants like Siri, Google Assistant, or Cortana became part of our daily lives, speech synthesis stopped being a curiosity and turned into a real, expected feature in many web applications. What's interesting is that, thanks to HTML5's Web Speech API, JavaScript provides us with a native API to work with speech synthesis (text-to-speech or TTS) easily, without needing external libraries or third-party dependencies.
The native speech synthesizer in JavaScript allows any text to be played directly in the browser, with control over language, pitch, and speed. What surprised me the most when I used it for the first time was how little it takes to have a working example: barely two lines of code.
How to Get Started with the Speech Synthesis API in JavaScript
Create Your First SpeechSynthesisUtterance Object
The SpeechSynthesisUtterance class is the core of all this: it captures text and prepares it to be converted into audio. Creating an object and starting it is as straightforward as this:
var speechMessage = new SpeechSynthesisUtterance('Hola, mundo!');
window.speechSynthesis.speak(speechMessage);This minimal example works right away in modern browsers. First, we create an instance of SpeechSynthesisUtterance passing the text as a parameter, and then we process it through the window.speechSynthesis interface, which actually gives voice to the browser by invoking the speak() method.
Play Dynamic Text: Basic Example
To read text coming from a variable instead of a fixed literal, simply assign it when creating the instance:
var texto = "Bienvenido a mi aplicación web";
var speechMessage = new SpeechSynthesisUtterance(texto);
window.speechSynthesis.speak(speechMessage);This opens the door to interactive applications where content can read personalized messages to the user, such as notifications, form results, or any text generated in real-time.
Essential Properties of SpeechSynthesisUtterance
The SpeechSynthesisUtterance class exposes a set of properties that allow you to customize exactly how the browser will speak. Among the most used are:
SpeechSynthesisUtterance.text
It is the most important property of all. It allows setting or getting the text that the browser will read aloud. You can also assign it directly in the constructor or modify it later:
speechMessage.text = "Texto actualizado dinámicamente";SpeechSynthesisUtterance.lang
Allows specifying the language of the synthesis using BCP 47 standard language tags. In my tests, changing the language noticeably improves the naturalness of pronunciation:
speechMessage.lang = 'es-ES'; // Spanish from Spain
speechMessage.lang = 'es-MX'; // Spanish from Mexico
speechMessage.lang = 'en-US'; // American EnglishSpeechSynthesisUtterance.volume
Controls the voice volume. Accepts a float value between 0 (silence) and 1 (maximum). Useful for balancing audio with other sound effects on the page:
speechMessage.volume = 0.8; // 80% of maximum volumeSpeechSynthesisUtterance.pitch
Sets the pitch of the voice. The value is a float between 0 (lowest) and 2 (highest). The default value is 1:
speechMessage.pitch = 1.2; // slightly higher pitch than normalSpeechSynthesisUtterance.rate
Controls the speech speed. Accepts values between 0.1 (very slow) and 10 (very fast), with 1 being normal speed. Lowering it slightly to 0.9 can be more comfortable for educational or tutorial content:
speechMessage.rate = 0.9; // slightly slower speedSpeechSynthesisUtterance.voice
Allows assigning a specific voice from those available in the browser. Voices are fetched asynchronously via window.speechSynthesis.getVoices(). It is important to listen to the onvoiceschanged event because the voice list might not be immediately available:
window.speechSynthesis.onvoiceschanged = function() {
var voices = window.speechSynthesis.getVoices();
speechMessage.voice = voices[0]; // selects the first available voice
};Events and Speech Synthesis Flow Control
The API also exposes events that let you react to different moments in the playback lifecycle. The most common are onstart and onend, useful for syncing animations, visual indicators, or any effect with the voice:
speechMessage.onstart = function(e) {
console.log('Playback started...');
};
speechMessage.onend = function(e) {
console.log('Playback finished.');
};Personally, I used them in an interactive tutorial project to trigger an animated icon while the browser was speaking, and the result was great in terms of user experience.
Other Useful Events for Advanced Projects
In addition to onstart and onend, the API exposes other events that allow more granular control over the playback flow:
onerror: fires if an error occurs during synthesis. Ideal for handling failures without breaking the user experience.onpauseandonresume: detect when playback is paused or resumed, enabling the implementation of play/pause controls in the UI.onboundary: fires at each word or sentence boundary. Very useful for highlighting the text currently being read in real-time (karaoke effect).
You can check the full list of events in the official Web Speech API documentation.
The Other Half of Web Speech API: Speech Recognition
In a previous post, we talked about the other side of the Web Speech API: speech recognition with SpeechRecognition(), which allows your applications to listen to the user through the device's microphone and transcribe what they say according to the configured language:
If speech recognition is the browser's "ear," speech synthesis with SpeechSynthesisUtterance and window.speechSynthesis is its "mouth." Together, they make up the complete Web Speech API suite.
Support in Modern Browsers
The Web Speech API has good support across major modern browsers: Chrome, Firefox, Edge, and Safari implement it, although available voices may vary depending on the operating system and installed language. Before using it, it's a good idea to check support with a simple check:
if ('speechSynthesis' in window) {
// Browser supports Speech Synthesis API
console.log('Ready to speak!');
} else {
console.log('API is not supported in this browser');
}It is good practice to always include this check before initializing any TTS functionality to avoid errors in browsers that don't implement it yet or have it disabled.
Tips to Improve User Experience
- Give the user the ability to select from available voices using
window.speechSynthesis.getVoices(). - Adjust
volumeandrateaccording to the content type: slower for educational text, faster for short notifications. - Avoid overly long texts in a single call to
speak(). Some browsers have character limits; split the content into chunks if necessary. - Use the
onboundaryevent to visually highlight the text currently being read and improve accessibility.
Complete Speech Synthesis Example in JavaScript
An example encapsulated in a reusable function is the cleanest way to integrate TTS into any project:
function hablar(texto, idioma, velocidad, tono) {
if (!('speechSynthesis' in window)) {
console.warn('Speech Synthesis API is not available.');
return;
}
var speechMessage = new SpeechSynthesisUtterance(texto);
speechMessage.lang = idioma || 'es-ES';
speechMessage.rate = velocidad || 1;
speechMessage.pitch = tono || 1;
speechMessage.volume = 1;
speechMessage.onstart = function() { console.log('Speaking...'); };
speechMessage.onend = function() { console.log('Done.'); };
window.speechSynthesis.speak(speechMessage);
}
// Usage:
hablar("¡Hola! Esto es un ejemplo dinámico.", 'es-ES', 0.9, 1.1);In my experience, allowing the user to adjust pitch and rate directly from the UI significantly increases accessibility and engagement in educational or interactive tutorial projects.
Frequently Asked Questions About Speech Synthesis in JavaScript
- Which browsers support the Speech Synthesis API?
- Chrome, Firefox, Edge, and Safari support
window.speechSynthesis, although available voices and quality may vary depending on the operating system. Chrome on Android and desktop has the broadest support.
- Chrome, Firefox, Edge, and Safari support
- How can I change the default language or voice?
- Use the
langproperty ofSpeechSynthesisUtterancefor the language (e.g.,'es-ES'or'en-US'), and thevoiceproperty with an object obtained fromwindow.speechSynthesis.getVoices()to select a specific voice.
- Use the
- Can I control the speed, volume, and pitch of the voice?
- Yes. Use
ratefor speed (between0.1and10),volumefor volume (between0and1), andpitchfor pitch (between0and2).
- Yes. Use
- Is it possible to play dynamic text from a variable in JavaScript?
- Yes, simply assign the variable to the
SpeechSynthesisUtteranceconstructor or itstextproperty, and callwindow.speechSynthesis.speak()with the object.
- Yes, simply assign the variable to the
- How do I stop or pause speech synthesis in JavaScript?
- Use
window.speechSynthesis.cancel()to completely stop playback,window.speechSynthesis.pause()to pause it, andwindow.speechSynthesis.resume()to resume it.
- Use