A synthetic voice announcing an arriving train in Sweden.

Problems playing this file? See media help.

Speech synthesis is the artificial production of human speech. A computer system used for this purpose is called a speech synthesizer, and can be implemented in software or hardware products. A text-to-speech (TTS) system converts normal language text into speech; other systems render symbolic linguistic representations like phonetic transcriptions into speech.^[1] The reverse process is speech recognition.

Synthesized speech can be created by concatenating pieces of recorded speech that are stored in a database. Systems differ in the size of the stored speech units; a system that stores phones or diphones provides the largest output range, but may lack clarity. For specific usage domains, the storage of entire words or sentences allows for high-quality output. Alternatively, a synthesizer can incorporate a model of the vocal tract and other human voice characteristics to create a completely "synthetic" voice output.^[2]

The quality of a speech synthesizer is judged by its similarity to the human voice and by its ability to be understood clearly. An intelligible text-to-speech program allows people with visual impairments or reading disabilities to listen to written words on a home computer. Many computer operating systems have included speech synthesizers since the early 1990s.

A text-to-speech system (or "engine") is composed of two parts:^[3] a front-end and a back-end. The front-end has two major tasks. First, it converts raw text containing symbols like numbers and abbreviations into the equivalent of written-out words. This process is often called text normalization, pre-processing, or tokenization. The front-end then assigns phonetic transcriptions to each word, and divides and marks the text into prosodic units, like phrases, clauses, and sentences. The process of assigning phonetic transcriptions to words is called text-to-phoneme or grapheme-to-phoneme conversion. Phonetic transcriptions and prosody information together make up the symbolic linguistic representation that is output by the front-end. The back-end—often referred to as the synthesizer—then converts the symbolic linguistic representation into sound. In certain systems, this part includes the computation of the target prosody (pitch contour, phoneme durations),^[4] which is then imposed on the output speech.

v t e Sound synthesis types
Frequency modulation Linear arithmetic Phase distortion Scanned Subtractive Additive Distortion
Sample-based or Sampler	Wavetable Granular Vector Concatenative
Physical modelling	Banded waveguide Digital waveguide Direct digital Formant Karplus–Strong string
Analog synthesizer	Graphical sound Modular
Digital synthesizer	Analog modeling Scanned synthesis Software synthesizer

Authority control databases
National	Israel Japan Czech Republic
Other	MusicBrainz instrument

History

Electronic devices

Synthesizer technologies

Concatenation synthesis

Unit selection synthesis

Diphone synthesis

Domain-specific synthesis

Formant synthesis

Articulatory synthesis

HMM-based synthesis

Sinewave synthesis

Deep learning-based synthesis

Audio deepfakes

Challenges

Text normalization challenges

Text-to-phoneme challenges

Evaluation challenges

Prosodics and emotional content

Dedicated hardware

Hardware and software systems

Texas Instruments

Mattel

SAM

Atari

Apple

Amazon

AmigaOS

Microsoft Windows

Votrax

Text-to-speech systems

Android

Internet

Open source

Others

Digital sound-alikes

Speech synthesis markup languages

Applications

Singing synthesis

See also

References

External links