
Episode 8 of 14
For the full video series, click here: https://aka.ms/AI-901onYouTube
How do applications understand and generate human speech? In this video, we explore the fundamentals of speech-enabled AI, focusing on the two core capabilities: speech-to-text and text-to-speech. You’ll discover how spoken language is converted into text, how text becomes natural-sounding audio, and what happens behind the scenes in each process. Whether it’s powering voice assistants, transcribing meetings, or improving accessibility, these technologies are transforming the way we interact with applications.
Immerse yourself in our rich, interactive materials at your own pace with self-directed learning: https://aka.ms/AI-901onLearn
00:00 Video Start
01:45 Speech-enabled solutions
05:16 Speech recognition
08:47 Speech synthesis
12:29 Demo: Explore AI speech











