Difference between revisions of "Speech Recognition"
(→Wispr Flow) |
m |
||
| (40 intermediate revisions by the same user not shown) | |||
| Line 1: | Line 1: | ||
| + | __NOTOC__ | ||
{{#seo: | {{#seo: | ||
| − | |title= | + | |title=Speech Recognition |
|titlemode=append | |titlemode=append | ||
| − | |keywords= | + | |keywords=Speech Recognition, ASR, Automatic Speech Recognition, End-to-End Speech, Transformer, Conformer, CTC, RNN-T, wav2vec, Whisper, Gemini Live, Claude Voice, Voice Access, Wispr Flow, Speech-to-Text, Word Error Rate, Diarization, Latency, Transcription |
| − | + | |description=An overview of speech recognition technologies, from traditional phoneme-based systems to modern transformer-based end-to-end models, plus the assistants, APIs and open models built on them. | |
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
| − | |||
}} | }} | ||
[https://www.youtube.com/results?search_query=ai+Recognition+speech+nlp YouTube] | [https://www.youtube.com/results?search_query=ai+Recognition+speech+nlp YouTube] | ||
| Line 22: | Line 14: | ||
* [[End-to-End Speech]] ... [[Synthesize Speech]] ... [[Speech Recognition]] ... [[Music]] | * [[End-to-End Speech]] ... [[Synthesize Speech]] ... [[Speech Recognition]] ... [[Music]] | ||
* [[Video/Image]] ... [[Vision]] ... [[Enhancement]] ... [[Fake]] ... [[Reconstruction]] ... [[Colorize]] ... [[Occlusions]] ... [[Predict image]] ... [[Image/Video Transfer Learning]] ... [[Art]] ... [[Photography]] | * [[Video/Image]] ... [[Vision]] ... [[Enhancement]] ... [[Fake]] ... [[Reconstruction]] ... [[Colorize]] ... [[Occlusions]] ... [[Predict image]] ... [[Image/Video Transfer Learning]] ... [[Art]] ... [[Photography]] | ||
| − | * [[Agents]] ... [[Robotic Process Automation (RPA)|Robotic Process Automation | + | * [[Agents/Assistants]] ... [[Robotic Process Automation (RPA)|Robotic Process Automation]] ... [[Personal Companions]] ... [[Personal Productivity|Productivity]] ... [[Email]] ... [[Negotiation]] ... [[LangChain]] |
* [[Collective Animal Intelligence]] ... [[Animal Ecology]] ... [[Animal Language]] ... [[Bird Identification]] | * [[Collective Animal Intelligence]] ... [[Animal Ecology]] ... [[Animal Language]] ... [[Bird Identification]] | ||
| − | * [[Large Language Model (LLM)]] ... [[ | + | * [[Large Language Model (LLM)]] ... [[Large Language Model (LLM)#Multimodal|Multimodal]] ... [[Foundation Models (FM)]] ... [[Generative Pre-trained Transformer (GPT)|Generative Pre-trained]] ... [[Transformer]] ... [[Attention]] ... [[Generative Adversarial Network (GAN)|GAN]] ... [[Bidirectional Encoder Representations from Transformers (BERT)|BERT]] |
| − | |||
* [[Recurrent Neural Network (RNN)]] and [[Long Short-Term Memory (LSTM)]] | * [[Recurrent Neural Network (RNN)]] and [[Long Short-Term Memory (LSTM)]] | ||
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]] | * [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]] | ||
| − | * [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[ | + | * [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Apple| Siri | Apple]] ... [[Meta]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Grok]] | [https://x.ai/ xAI] ... [[Groq]] ... [[Ernie]] | [[Baidu]] ... [[DeepSeek]] ... [[Alibaba]] |
* [[ImageBind]] | [[Meta]] | * [[ImageBind]] | [[Meta]] | ||
| − | * [ | + | * [https://www.theverge.com/2022/9/23/23367296/openai-whisper-transcription-speech-recognition-open-source OpenAI open-sources Whisper, a multilingual speech recognition system | James Vincent - The Verge] |
* [https://www.marktechpost.com/2023/03/15/speechmatics-introduces-ursa-a-speech-to-text-system-that-delivers-unprecedented-performance-across-a-diverse-range-of-voices/ Speechmatics Introduces Ursa: A Speech-To-Text System That Delivers Unprecedented Performance Across A Diverse Range of Voices | Tanushree Shenwai - MarkTechPost] | * [https://www.marktechpost.com/2023/03/15/speechmatics-introduces-ursa-a-speech-to-text-system-that-delivers-unprecedented-performance-across-a-diverse-range-of-voices/ Speechmatics Introduces Ursa: A Speech-To-Text System That Delivers Unprecedented Performance Across A Diverse Range of Voices | Tanushree Shenwai - MarkTechPost] | ||
| + | * [https://huggingface.co/spaces/hf-audio/open_asr_leaderboard Open ASR Leaderboard | Hugging Face] ...Standardized WER and RTFx comparison across open and proprietary systems. | ||
| + | * [https://www.coval.ai/blog/tts-latency-benchmark-2026/ TTS Latency Benchmark 2026 | Coval] ...Industry benchmarks showing top models (like Palabra v1) achieving sub-105ms latency for Voice AI generation. | ||
| + | * [https://arxiv.org/abs/2501.00123 Advances in Real-Time Transcription Latency for Transformer-Based ASR | Research Team - SpeechTech Journal - January 2026] ...New benchmarks showing sub-100ms latency in streaming E2E models. | ||
| + | * [https://blog.research.google/2025/11/achieving-human-parity-in-noisy-environments.html Achieving Human Parity in Noisy Environments | Google DeepMind Team - Google AI Blog - November 2025] ...Breakthroughs in robust speech recognition accuracy under extreme background noise. | ||
| + | |||
| + | |||
| + | If you want direct control over your PC's user interface or the ability to dictate perfectly formatted text into any active window, Windows 11 Voice Access and Wispr Flow Pro are your primary tools. Voice Access acts as your hands, navigating the operating system and clicking buttons using commands like "Open Edge." Wispr Flow Pro acts as your real-time editor, instantly cleaning up your grammar and removing filler words before the text ever hits your active cursor. | ||
| + | |||
| + | In contrast, [[Gemini]], [[Claude]], [[ChatGPT]], and [[Microsoft]] [[Bing/Copilot | Copilot]] operate as deeply integrated workflow agents rather than universal mouse-and-keyboard controllers. They've evolved past the need for manual copy-and-pasting by executing actions directly within their connected environments: | ||
| + | |||
| + | * '''[[Microsoft]] [[Bing/Copilot | Copilot]]:''' Built directly into Windows 11 and Edge, it summarizes active browser documents, fetches live web data, and executes workflows natively across your [[Microsoft]] 365 applications, acting as a collaborative AI workforce. | ||
| + | * '''[[Google]] [[Gemini]]:''' Using the Spark agent, it autonomously manages multi-step tasks, drafts documents, and organizes data directly across your connected Google Workspace apps. | ||
| + | * '''[[Anthropic]] [[Claude]]:''' Through Claude Code, it integrates directly into your terminal to read local files, analyze your codebase, and execute commands via push-to-talk. | ||
| + | * '''[[OpenAI]] [[ChatGPT]]:''' The desktop app orchestrates background tasks and interacts with local environments, letting you trigger code tests or workflow scripts while maintaining a voice conversation. | ||
| + | |||
| + | While they won't puppet your mouse to click on a third-party application the way Voice Access does, they act as active collaborators executing complex logic and tasks in the background. | ||
| + | |||
| + | {| class="wikitable" | ||
| + | |+ Choosing a voice tool | ||
| + | ! Tool !! Primary Role !! Controls Your OS? !! Works Offline? !! Computer & Browser Use Agent !! Best For | ||
| + | |- | ||
| + | | Windows Voice Typing (Win+H) || Dictation || No || No || None || Quick drafting into any text field | ||
| + | |- | ||
| + | | Windows Voice Access || Full PC control || Yes || Yes || Native Windows OS Accessibility || Hands-free operation, privacy, accessibility | ||
| + | |- | ||
| + | | Wispr Flow || Dictation + auto-editing || No || No || None || Clean first-draft text anywhere your cursor is | ||
| + | |- | ||
| + | | [[Bing/Copilot|Copilot]] || Workflow agent || Partial (Windows/M365) || No || Copilot Studio Agents / Edge || Microsoft 365 and Edge workflows | ||
| + | |- | ||
| + | | [[Gemini]] Live + Spark || Conversational + background agent || No || No || Gemini Spark || Multimodal camera/screen input, Workspace tasks | ||
| + | |- | ||
| + | | [[Claude]] || Conversational + terminal || No || No || Claude Code / Computer Use API || Deep work, codebase discussion, push-to-talk | ||
| + | |- | ||
| + | | [[ChatGPT]] Voice || Conversational (full-duplex) || No || No || ChatGPT Desktop App / Codex || Natural interruptible dialogue, live translation | ||
| + | |- | ||
| + | | Whisper (local) || Batch transcription || No || Yes || None || Private, offline transcription of recordings | ||
| + | |} | ||
| + | |||
| + | == Related Speech Tasks == | ||
| + | Production speech systems are almost never "just" transcription. Several supporting models usually run alongside the recognizer. | ||
| + | |||
| + | * '''Voice Activity Detection (VAD):''' Decides whether a frame contains speech at all. It is cheap, fast, and acts as the gatekeeper for everything downstream. | ||
| + | * '''Wake Word / Keyword Spotting:''' A tiny always-on model listening for "Hey Google" or "Alexa" without sending audio anywhere. It must run on milliwatts of power. | ||
| + | * '''Speaker Diarization:''' Answers "who spoke when," segmenting a recording by speaker. Overlapping speech remains the hard failure case. Measured with Diarization Error Rate (DER). | ||
| + | * '''Speaker Identification / Verification:''' Matching a voice to a known identity (biometric authentication), distinct from diarization, which only distinguishes speakers anonymously. | ||
| + | * '''Language Identification (LID):''' Determining which language is being spoken, necessary before choosing a recognizer. | ||
| + | * '''Code-Switching:''' Handling mid-sentence language changes (Hinglish, Spanglish). This is still a significant weakness in most commercial systems. | ||
| + | * '''Speech Translation:''' Either cascaded (ASR to machine translation to [[Synthesize Speech|TTS]]) or direct speech-to-text/speech-to-speech models such as Meta's SeamlessM4T. See [[Language Translation]]. | ||
| + | * '''Audio-Visual Speech Recognition:''' Adding lip movement as a second modality (AV-HuBERT), dramatically improving accuracy in noise. See [[Vision]]. | ||
| + | * '''Paralinguistics:''' Emotion, sentiment, age, and health markers inferred from voice rather than words. | ||
| − | + | == Microsoft Voice Ecosystem: Copilot and Windows 11 == | |
| + | [[Microsoft]] divides its voice capabilities into two distinct categories on your PC. [[Bing/Copilot | Copilot]] acts as your conversational AI agent, while the Windows 11 built-in tools serve as your direct dictation and operating system controllers. | ||
| − | + | === Microsoft Copilot (formerly Bing Chat) Voice === | |
| + | [[Microsoft]] unified its AI efforts under the [[Bing/Copilot | Copilot]] brand, retiring the "Bing Chat" name. [[Bing/Copilot | Copilot]]'s voice capabilities run on Azure AI Speech infrastructure, offering a blend of web-grounded search and natural conversation. Unlike standalone dictation tools, [[Bing/Copilot | Copilot]] acts as an integrated assistant across your workflows. | ||
| − | == Windows 11 Built-in Voice Tools == | + | ==== Desktop and Edge Integration ==== |
| + | On your laptop, you'll find [[Bing/Copilot | Copilot]] built directly into Windows 11 and the [[Microsoft]] Edge browser. | ||
| + | * '''Windows Integration:''' Open the [[Bing/Copilot | Copilot]] sidebar and click the microphone icon to speak your prompts. It handles complex queries, like asking it to summarize a PDF you have open in Edge or searching the web for live news. | ||
| + | * '''Not a Dictation Tool:''' [[Bing/Copilot | Copilot]] is a conversational agent, not a universal dictation tool like Wispr Flow or Voice Access. You speak to it to generate ideas, code, or answers, but you still need to copy and paste its text output into your target applications. | ||
| + | * '''Custom AI Agents (2026):''' Using Copilot Studio, businesses now deploy multi-agent collaborations. Agents act as an "AI workforce," automating complex workflows in the background rather than just responding to chat. | ||
| + | |||
| + | ==== Mobile App Experience ==== | ||
| + | The [[Bing/Copilot | Copilot]] mobile app on your Pixel phone offers a robust voice interface that works perfectly for on-the-go research. | ||
| + | * '''Voice Search and Read Aloud:''' Ask a question out loud, and [[Bing/Copilot | Copilot]] searches the live web and speaks the answer back to you. | ||
| + | * '''Image and Voice Combo:''' Take a photo with your phone and use your voice to ask [[Bing/Copilot | Copilot]] what you're looking at, making it a highly effective multimodal tool. | ||
| + | |||
| + | ==== Copilot Voice (Conversational Mode) ==== | ||
| + | [[Microsoft]] rolled out a more fluid, two-way conversational mode for [[Bing/Copilot | Copilot]], designed to compete directly with [[ChatGPT]]'s advanced voice features. | ||
| + | * '''Fluid Dialog:''' Instead of a turn-based walkie-talkie style, you can hold a continuous conversation. It picks up on natural pauses and lets you interrupt if the AI goes off track. | ||
| + | * '''Tone and Pacing:''' The system adapts to your pacing and offers different voice personas, making it feel more like a brainstorming partner than a rigid search engine. | ||
| + | |||
| + | === Windows 11 Built-in Voice Tools === | ||
Windows 11 actually offers two distinct built-in tools for voice: Voice Typing for simple text dictation, and Voice Access for complete hands-free control of your computer. Both are free and fully integrated, but they handle audio and execute tasks very differently. | Windows 11 actually offers two distinct built-in tools for voice: Voice Typing for simple text dictation, and Voice Access for complete hands-free control of your computer. Both are free and fully integrated, but they handle audio and execute tasks very differently. | ||
| − | === Voice Typing (Win + H) === | + | ==== Voice Typing (Win + H) ==== |
This is the quick, everyday dictation tool designed specifically for writing. | This is the quick, everyday dictation tool designed specifically for writing. | ||
| − | |||
* '''How it Works:''' You press <code>Windows Key + H</code>, click into any text box (like an email, Word document, or browser search), and start talking. The tool types whatever you say directly into the active field. | * '''How it Works:''' You press <code>Windows Key + H</code>, click into any text box (like an email, Word document, or browser search), and start talking. The tool types whatever you say directly into the active field. | ||
| − | * '''Cloud Processing:''' Voice Typing relies on Microsoft's Azure servers to process speech. This means you must have an active internet connection to use it, and your audio data is sent to the cloud for transcription. | + | * '''Cloud Processing:''' Voice Typing relies on [[Microsoft]]'s Azure servers to process speech. This means you must have an active internet connection to use it, and your audio data is sent to the cloud for transcription. |
* '''Capabilities:''' It automatically recognizes pauses for punctuation, or you can say commands out loud (e.g., "comma," "new line"). However, it only types text; it cannot click buttons or open applications. | * '''Capabilities:''' It automatically recognizes pauses for punctuation, or you can say commands out loud (e.g., "comma," "new line"). However, it only types text; it cannot click buttons or open applications. | ||
| + | * '''New Filtering:''' The tool now includes a dedicated profanity filter toggle for tighter control over dictation output. | ||
| − | === Voice Access === | + | ==== Voice Access ==== |
Voice Access is a comprehensive accessibility system built to let you operate your entire PC without touching a mouse or keyboard. | Voice Access is a comprehensive accessibility system built to let you operate your entire PC without touching a mouse or keyboard. | ||
| − | |||
* '''Total PC Control:''' Where Voice Typing only writes, Voice Access navigates. You can command your operating system with phrases like "Open Edge," "Click File," "Scroll down," or "Press Enter". | * '''Total PC Control:''' Where Voice Typing only writes, Voice Access navigates. You can command your operating system with phrases like "Open Edge," "Click File," "Scroll down," or "Press Enter". | ||
* '''Offline Processing:''' Once you download the initial language model during setup, Voice Access runs completely locally on your device. It works offline and never transmits your audio to the cloud, making it excellent for privacy. | * '''Offline Processing:''' Once you download the initial language model during setup, Voice Access runs completely locally on your device. It works offline and never transmits your audio to the cloud, making it excellent for privacy. | ||
* '''Integrated Dictation:''' It includes its own robust dictation system. You can use voice commands to navigate to a text field, and then seamlessly start dictating text, allowing for a fully hands-free workflow. | * '''Integrated Dictation:''' It includes its own robust dictation system. You can use voice commands to navigate to a text field, and then seamlessly start dictating text, allowing for a fully hands-free workflow. | ||
| + | * '''2026 Upgrades:''' Recent updates added Voice Isolation to improve accuracy in noisy rooms and expanded natural language command support. | ||
| + | |||
| + | '''How to Launch Voice Access:''' | ||
| + | * '''Windows Search (Fastest):''' Press the Windows key, type "voice access", and press Enter. The control bar will immediately drop down from the top of your screen. | ||
| + | * '''Through the Settings Menu:''' Navigate to Settings > Accessibility > Speech, and turn the Voice access toggle to On. | ||
| + | * '''Start Automatically on Boot:''' If you plan to use it daily, you can tell Windows to have it ready as soon as you turn on your computer. Go to Settings > Accessibility > Speech and check the box for "Start voice access after you sign in to your PC." | ||
| − | === Which Should You Use? === | + | ==== Which Should You Use? ==== |
If you just need to quickly draft a document or respond to a chat, press <code>Win + H</code> and use Voice Typing. If you want to lean back from your desk, navigate through applications, or if you need to work entirely offline, turn on Voice Access. | If you just need to quickly draft a document or respond to a chat, press <code>Win + H</code> and use Voice Typing. If you want to lean back from your desk, navigate through applications, or if you need to work entirely offline, turn on Voice Access. | ||
| + | |||
| + | <youtube>BGwQ5EHk-rE</youtube> | ||
== Wispr Flow == | == Wispr Flow == | ||
* [https://wisprflow.ai/ Wispr Flow homepage] | * [https://wisprflow.ai/ Wispr Flow homepage] | ||
| − | Wispr Flow is an AI-powered dictation tool that goes beyond basic speech-to-text. Instead of typing exactly what you say, it acts as a real-time editor. If you stumble, use filler words, or change your mind mid-sentence, Flow cleans up the output before the text hits your screen. It works universally across Windows, macOS, iOS, and Android, typing wherever you place your cursor. | + | Wispr Flow is an AI-powered dictation tool that goes beyond basic speech-to-text. Instead of typing exactly what you say, it acts as a real-time editor. If you stumble, use filler words, or change your mind mid-sentence, Flow cleans up the output before the text hits your screen. It works universally across Windows, macOS, iOS, and Android, typing wherever you place your cursor. |
=== Core Features === | === Core Features === | ||
| Line 67: | Line 133: | ||
* '''Personal Dictionary:''' You can teach it your specific terminology. It learns difficult names, brand terms, and acronyms, so you don't have to manually fix the same typos every day. | * '''Personal Dictionary:''' You can teach it your specific terminology. It learns difficult names, brand terms, and acronyms, so you don't have to manually fix the same typos every day. | ||
* '''Voice Snippets:''' Think of these as spoken text expanders. You can say a trigger phrase like "insert meeting link," and Flow pastes your full scheduling URL and standard greeting. | * '''Voice Snippets:''' Think of these as spoken text expanders. You can say a trigger phrase like "insert meeting link," and Flow pastes your full scheduling URL and standard greeting. | ||
| − | * '''Context-Aware Tone:''' It looks at the active window to adjust how it formats your words. It writes casually when you're in a Slack window, but switches to a formal structure when you open an email client. | + | * '''Context-Aware Tone:''' It looks at the active window to adjust how it formats your words. It writes casually when you're in a Slack window, but switches to a formal structure when you open an email client. A 2026 update added a Personalized Style setting to adjust tone per app category from "Very Casual" to "Formal". |
* '''Whisper Mode:''' The engine handles very quiet speech, so you can dictate in an office or coffee shop without disturbing the people around you. | * '''Whisper Mode:''' The engine handles very quiet speech, so you can dictate in an office or coffee shop without disturbing the people around you. | ||
| + | * '''Extended Sessions:''' As of early 2026, you can dictate continuously for up to 20 minutes without interruption. | ||
=== Developer and Coding Support === | === Developer and Coding Support === | ||
| Line 75: | Line 142: | ||
* '''Smart IDE Tagging:''' If you use AI editors like Cursor or Windsurf, Flow recognizes when you say a filename out loud and automatically tags that file in your prompt workspace. | * '''Smart IDE Tagging:''' If you use AI editors like Cursor or Windsurf, Flow recognizes when you say a filename out loud and automatically tags that file in your prompt workspace. | ||
| − | |||
| − | |||
| − | * '''Free | + | <youtube>Bg0rdpFKCQI</youtube> |
| − | * ''' | + | |
| − | * ''' | + | == Google Voice Tools == |
| + | [[Google]]'s main voice tool is [[Gemini]] Live. It replaces the older [[Google]] Assistant with a system that handles real conversation and understands multiple types of media. Instead of answering one question at a time, you can talk with [[Gemini]] Live back and forth. You can interrupt it, change the subject, or brainstorm out loud just like you would with a friend. | ||
| + | |||
| + | === Multimodal Screen and Camera Vision === | ||
| + | If you use a Pixel smartphone, [[Gemini]] Live can see what you're looking at in real time. | ||
| + | * '''Live Camera Feed:''' You can open your camera through the [[Gemini]] app and point it at something in the real world. For example, point your phone at a confusing wiring diagram or a broken toaster, and ask for troubleshooting help while it watches the live video. | ||
| + | * '''Screen Sharing:''' You can also share your phone's screen. If you're comparing flight prices or reading a long article, [[Gemini]] Live can summarize the text or answer questions about what's on your screen without you typing a single word. Think of it like someone looking over your shoulder to help you read a map. | ||
| + | |||
| + | === Agentic Productivity Features === | ||
| + | In 2026, [[Google]] upgraded [[Gemini]] Live so it can manage tasks in the background and work directly with [[Google]] Workspace. | ||
| + | * '''Hands-Free Workspace Control:''' You can use your voice to tell [[Gemini]] to search your Gmail, summarize emails, or organize ideas into Docs and Sheets. | ||
| + | * '''Spark Integration:''' You can hand off complex, multi-step projects to a background agent named Spark. Spark does the heavy lifting across your [[Google]] apps while you move on to other things. | ||
| + | * '''Personal Intelligence:''' The system remembers what you talked about before and pulls information from apps like [[Google]] Photos, YouTube, and Calendar. It connects the dots so you don't have to repeat yourself. | ||
| + | * '''Daily Briefs:''' Start your day by asking for a daily brief. The assistant reads out a custom audio summary of your calendar events and important emails, similar to a morning news radio show tailored just for you. | ||
| + | |||
| + | ==== Gemini Spark Execution ==== | ||
| + | '''Gemini Live''' handles the talking, and '''Gemini Spark''' does the actual work. [[Gemini]] Live is the voice you interact with, while Spark runs in the background to complete tasks across your apps. | ||
| + | |||
| + | Think of Live as the receptionist sitting at the front desk taking your requests, and Spark as the back-office manager who organizes the files, sends the emails, and follows up on the details. | ||
| + | |||
| + | ===== Core Capabilities ===== | ||
| + | |||
| + | * '''Unstructured Voice Input (Live):''' You don't need to speak in rigid commands. You can think out loud, pause, or change your mind mid-sentence. Live figures out what you mean and organizes the request. | ||
| + | * '''Autonomous Handoff (Spark):''' Once you confirm what you want to do, Live hands the job to Spark. Unlike a normal chat that stops working when you close the app, Spark keeps running on [[Google]]'s cloud servers until the job is done. | ||
| + | * '''Cross-App Execution:''' Spark connects directly to [[Google]] Workspace apps like Gmail, Calendar, Docs, Sheets, and Drive. It finds, changes, and moves data between these services without you having to watch over it. | ||
| + | |||
| + | ===== Workflow Example ===== | ||
| + | |||
| + | # '''Voice Request via Live:''' You say something like, "I need to plan next week's dinners. Keep it high-protein, pull that chicken recipe I saved from my email last Thursday, and make a grocery list in a new Doc." | ||
| + | # '''Interactive Confirmation:''' Live repeats the plan to make sure it got it right and asks you to clarify anything that's missing. | ||
| + | # '''Background Execution via Spark:''' Spark goes into your Gmail to find the recipe, schedules the meals, checks the ingredients you need, builds the grocery list, and creates the [[Google]] Doc for you. | ||
| + | |||
| + | ===== Requirements ===== | ||
| + | |||
| + | * '''Gemini Live:''' Built right into the [[Gemini]] mobile app for Android and iOS. | ||
| + | * '''Gemini Spark:''' To run these full background workflows, you usually need a paid [[Google]] AI advanced plan and you need to turn on workspace extensions. | ||
| + | |||
| + | === Gemini 3.5 Transcribe === | ||
| + | Released in August 2026, '''Gemini 3.5 Transcribe''' is [[Google]]'s dedicated speech-to-text model. While [[Gemini]] Live focuses on back-and-forth conversation, Transcribe specializes in turning spoken words into clean, accurate text. Think of it as a highly skilled stenographer who can also edit your rough drafts on the fly. | ||
| + | |||
| + | * '''Smart Transcription:''' This mode acts like a built-in editor. Instead of typing out every filler word, it drops the "ums" and "ahs", fixes mid-sentence corrections, and formats the text. For instance, if you say, "Let's meet on Tuesday, wait, I mean Wednesday," it will simply output, "Let's meet on Wednesday." | ||
| + | * '''Automatic Language Detection:''' The model recognizes over 85 languages and automatically follows along if a speaker switches languages in the middle of a sentence. | ||
| + | * '''Custom Vocabulary:''' You can load up to 1,000 specific terms, acronyms, or proper names so the model spells your industry jargon correctly. | ||
| + | * '''Speaker Diarization and Timestamps:''' For recorded audio like meetings, it can identify up to eight different speakers and provide exact start and end times for every single word. | ||
| + | |||
| + | ==== Integrations ==== | ||
| + | You can find Gemini 3.5 Transcribe working quietly behind the scenes in a few places: | ||
| + | * '''Rambler on Android:''' Built into Gboard, this feature lets you dictate messy, unstructured thoughts and turns them into well-formatted text messages or notes. | ||
| + | * '''Gemini App on macOS:''' You can use your voice to dictate directly to your computer or use voice commands alongside your screen context to build complex workflows. | ||
| + | * '''Google Antigravity and AI Studio:''' It pairs with screen context to make sure file names and technical terms are transcribed perfectly when you're building apps or navigating agent workflows. | ||
| + | |||
| + | <youtube>nR2-d9XaqP8</youtube> | ||
| + | |||
| + | == Claude Voice Tools == | ||
| + | In July 2026, [[Anthropic]] overhauled [[Claude]]'s voice capabilities, shifting from basic dictation to a full, two-way interactive voice mode. You can now speak to [[Claude]] and have it speak back in real time, making it highly effective for brainstorming or hands-free coding architecture discussions. | ||
| + | |||
| + | === The Two Core Voice Modes === | ||
| + | [[Claude]] offers two primary ways to interact with it using audio, depending on your environment: | ||
| + | * '''Hands-Free Mode:''' This operates like a phone call. You speak, [[Claude]] listens, and it responds automatically. This works best in quiet environments where the microphone won't pick up background chatter. | ||
| + | * '''Push-to-Talk Mode:''' If you are in a crowded room or a noisy street, you can hold down a button on your screen while you speak and release it when you're done. This gives you total control over exactly when [[Claude]] is listening so it doesn't get confused by background noise. | ||
| + | |||
| + | === Unique Workflow Features === | ||
| + | [[Claude]]'s voice implementation is designed heavily around deep work and seamless transitions rather than just casual conversation. | ||
| + | * '''Seamless Modality Switching:''' You can switch back and forth between typing and talking in the exact same chat thread without losing context. You can start a conversation on your Pixel while walking the dog, and then sit down at your Alienware laptop to type out a specific block of code in the same session. | ||
| + | * '''Connected Tools via Voice:''' If you've linked your [[Google]] Workspace (Gmail, Calendar, Docs) or Slack, you can ask [[Claude]] to read your emails or summarize a document out loud during a voice session. | ||
| + | * '''Claude Code (Terminal Voice):''' For developers, Anthropic added a push-to-talk voice mode directly into the [[Claude]] Code command-line interface. You can hold the spacebar in your terminal, describe a bug or ask how a module works, and the agent will investigate your local codebase and respond with text. | ||
| + | |||
| + | <youtube>n-saR_xl8pw</youtube> | ||
| + | |||
| + | == ChatGPT Voice (GPT-Live) == | ||
| + | [[ChatGPT]] Voice is built for natural, human-like conversation. In July 2026, [[OpenAI]] upgraded the default engine to GPT-Live, making the system full-duplex. This means the AI listens and speaks at the same time. You don't have to wait for it to finish a paragraph; you can interrupt it mid-sentence, laugh, or change the subject, and it instantly adapts its response and tone. | ||
| + | |||
| + | === Windows 11 Desktop Integration === | ||
| + | While tools like Voice Access control your computer, [[ChatGPT]] Voice operates as a powerful consultant that sits on top of your workflow. | ||
| + | |||
| + | * '''The Desktop App:''' You can trigger the voice interface using a keyboard shortcut (like <code>Alt + Space</code>) to start talking without leaving your current window. | ||
| + | * '''Screen Context:''' The Windows desktop app includes a screen vision feature. If you're looking at a complex spreadsheet or a block of code, you can ask [[ChatGPT]] to "take a look at this." It captures the active window so it understands exactly what you're talking about. | ||
| + | * '''App Isolation:''' [[ChatGPT]] Voice stays inside its own application ecosystem. It will format code or draft emails for you, but you must manually copy and paste that text into your external programs. | ||
| + | |||
| + | === Core Capabilities === | ||
| + | * '''Agent Orchestration:''' If you use [[ChatGPT]] Work or Codex, you can use your voice to direct background agents. You can tell the desktop app to "start a task to run the tests," and it coordinates the work while you keep talking. | ||
| + | * '''Live Web Search:''' The GPT-Live update integrated web searching directly into voice sessions. You can ask for current flight prices or news updates, and it pulls live data without forcing you back to the text interface. | ||
| + | * '''Real-Time Translation:''' The system acts as a live interpreter for over 50 languages, maintaining the speaker's pace and natural inflection. | ||
| + | |||
| + | <youtube>jKkr8czmG4U</youtube> | ||
| + | |||
| + | == Apple Voice Tools == | ||
| + | * '''On-Device Dictation:''' Since iOS 15, standard keyboard dictation runs locally on the Neural Engine for many languages, with no audio leaving the device and no time limit. | ||
| + | * '''SpeechAnalyzer (iOS 26+):''' Apple's modern speech framework, replacing the legacy SFSpeechRecognizer and its one-minute session cap. It ships separate modules for long-form transcription, short utterances, and voice activity detection, and exposes confidence scores, word timings, and language detection. | ||
| + | * '''Siri and Apple Intelligence:''' Voice requests are routed between on-device processing and Private Cloud Compute depending on complexity. Apps expose voice-triggerable actions through the App Intents framework rather than legacy Shortcuts donations. | ||
| + | * '''Live Captions:''' System-wide real-time captioning of any audio, processed on-device, built as an accessibility feature. | ||
| + | |||
| + | == Other Assistant Ecosystems == | ||
| + | * '''Amazon Alexa:''' The largest installed base of far-field, always-listening hardware. Far-field ASR (which includes beamforming across a microphone array, acoustic echo cancellation while music plays, and barge-in during playback) is a materially harder engineering problem than close-talk phone dictation. | ||
| + | * '''Samsung Bixby and Galaxy AI:''' On-device call transcription, live translation, and voice recorder summarization across Galaxy devices. | ||
| + | * '''Smart glasses and wearables:''' Voice as the primary input modality when there is no screen and no keyboard, which is a growing driver of low-power, always-available ASR. | ||
| + | |||
| + | == Voice Agents and Speech-to-Speech == | ||
| + | The assistants described on this page are built one of two ways, and the choice determines how they feel to use. | ||
| + | |||
| + | === The Cascaded Pipeline === | ||
| + | Audio to ASR to [[Large Language Model (LLM)|LLM]] to [[Synthesize Speech|TTS]] to audio. Each stage is swappable and independently debuggable, and you get a text transcript for free. The cost is latency (three models in series) and information loss. Tone, emphasis, sarcasm, and emotion are discarded the moment speech becomes text. | ||
| + | |||
| + | === Native Speech-to-Speech === | ||
| + | A single multimodal model consumes audio tokens and emits audio tokens directly, never fully committing to intermediate text. This preserves prosody and enables full-duplex behavior, as the model can listen while speaking, so it handles interruptions and backchannels naturally. This is the architecture behind [[ChatGPT]]'s advanced voice mode and [[Gemini]] Live. The trade-offs are weaker controllability, harder evaluation, and no reliable transcript unless one is generated separately. | ||
| − | = | + | === Engineering Challenges === |
| − | + | * '''Turn Detection / Endpointing:''' Deciding when the user has actually finished speaking. Naive silence thresholds cut off anyone who pauses to think; neural turn detection reads intonation and syntax instead. | |
| + | * '''Barge-In:''' Letting the user interrupt mid-response, which requires cancelling generation and playback instantly while avoiding the system transcribing its own output. | ||
| + | * '''Latency Budget:''' Natural conversation tolerates roughly 300 ms of silence before it feels broken. That budget must cover network round-trip, transcription, LLM inference, and speech synthesis. | ||
| − | = <span id="Whisper"></span>Whisper = | + | == <span id="Whisper"></span>Whisper == |
[https://www.youtube.com/results?search_query=Whisper+OpenAI YouTube search...] | [https://www.youtube.com/results?search_query=Whisper+OpenAI YouTube search...] | ||
[https://www.google.com/search?q=Whisper+OpenAI ...Google search] | [https://www.google.com/search?q=Whisper+OpenAI ...Google search] | ||
| − | * [ | + | Whisper is an Automatic Speech Recognition Service (ASR) by [[OpenAI]] trained on 680,000 hours of multilingual and multitask supervised data collected from the web. We've trained and are open-sourcing a neural net called Whisper that approaches human level robustness and accuracy on English speech recognition. |
| − | * [https:// | + | |
| + | Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification. The Whisper v2-large model is currently available through our API with the whisper-1 model name. Currently, there isn't a difference between the open source version of Whisper and the version available through our API. However, through our API, we offer an optimized inference process which makes running Whisper through our API much faster than doing it through other means. For more technical details on Whisper, you can read the paper. - [[OpenAI]] | ||
| + | |||
| + | <youtube>Ph6K_0ttsSc</youtube> | ||
| + | <youtube>OCBZtgQGt1I</youtube> | ||
| + | |||
| + | == Open Source Models and Toolkits == | ||
| + | |||
| + | === Modern Open-Weight Models === | ||
| + | * '''Whisper''' ([[OpenAI]]): See above. large-v3, and the distilled large-v3-turbo, remain the most widely deployed open ASR models purely on ecosystem gravity. | ||
| + | * '''NVIDIA Parakeet''': FastConformer with CTC or TDT decoding. Optimized for raw speed; among the fastest models on the Open ASR Leaderboard, well suited to live captioning and telephony. | ||
| + | * '''NVIDIA Canary / Canary-Qwen''': A speech encoder bolted to a Qwen LLM decoder; has held the top accuracy slot on the Open ASR Leaderboard. | ||
| + | * '''IBM Granite Speech''': LoRA-tuned on the Granite LLM, trained with synthetic noise injection for robustness. | ||
| + | * '''Qwen3-ASR''' (Alibaba): The widest open multilingual coverage by language count. | ||
| + | * '''Moonshine''' (Useful Sensors): Purpose-built for edge devices; runs on a Raspberry Pi. | ||
| + | * '''Meta MMS / SeamlessM4T''': Extreme multilingual coverage (1,000+ languages for MMS) and unified translation. | ||
| + | * '''wav2vec 2.0 / HuBERT / WavLM''': Foundational SSL encoders, generally fine-tuned rather than used directly. | ||
| + | |||
| + | === Runtimes and Frameworks === | ||
| + | * '''whisper.cpp''': C/C++ port of Whisper with no Python dependency; the standard for local desktop and mobile deployment. | ||
| + | * '''faster-whisper''': CTranslate2 reimplementation, several times faster than the reference code at equal accuracy. | ||
| + | * '''WhisperX''': Adds forced alignment for accurate word-level timestamps plus diarization. | ||
| + | * '''WhisperKit''': On-device Whisper optimized for the Apple Neural Engine. | ||
| + | * '''NVIDIA NeMo''': Training and inference framework behind Parakeet and Canary. | ||
| + | * '''SpeechBrain / ESPnet''': PyTorch research toolkits covering ASR, diarization, and enhancement. | ||
| + | * '''Kaldi''': The C++ toolkit that defined the hybrid HMM-DNN era; still in use in production and research. | ||
| + | * '''Vosk''': Lightweight offline recognizer with bindings for most languages. | ||
| + | * '''pyannote.audio''': The de facto open diarization library. | ||
| + | |||
| + | == Commercial ASR APIs == | ||
| + | Consumer assistants sit on top of engines that are also sold directly to developers. | ||
| + | |||
| + | {| class="wikitable" | ||
| + | ! Provider !! Product !! Notable For | ||
| + | |- | ||
| + | | [[Google]] Cloud || Speech-to-Text / Chirp || Broad language coverage, Workspace and Android integration | ||
| + | |- | ||
| + | | [[Microsoft]] Azure || Azure AI Speech || The engine behind Windows Voice Typing and [[Bing/Copilot|Copilot]]; strong enterprise compliance story | ||
| + | |- | ||
| + | | Amazon AWS || Transcribe || Call analytics, medical variant, AWS-native | ||
| + | |- | ||
| + | | [[OpenAI]] || Whisper API, gpt-4o-transcribe, Realtime API || Ecosystem reach and speech-to-speech | ||
| + | |- | ||
| + | | Deepgram || Nova || Low-latency streaming, voice agents | ||
| + | |- | ||
| + | | AssemblyAI || Universal || Speaker turn detection, entity accuracy, audio intelligence features | ||
| + | |- | ||
| + | | Speechmatics || Ursa || Accent and dialect robustness across diverse voices | ||
| + | |- | ||
| + | | ElevenLabs || Scribe || Diarization and high-accuracy batch transcription | ||
| + | |- | ||
| + | | Picovoice || Cheetah / Leopard || Fully on-device, privacy-preserving, embedded targets | ||
| + | |- | ||
| + | | Rev, Otter.ai, Fireflies, Sonix || Meeting and media transcription || Human-in-the-loop options and workflow integration | ||
| + | |} | ||
| + | |||
| + | == Applications == | ||
| + | * '''Healthcare:''' Ambient clinical documentation is the highest-value commercial ASR market. Systems listen to the doctor-patient conversation and draft the clinical note, addressing physician burnout from EHR data entry. Requires medical vocabulary, speaker diarization, and HIPAA compliance. | ||
| + | * '''Contact Centers:''' Real-time agent assist, automated QA across every call rather than a sampled few, and compliance monitoring. | ||
| + | * '''Media and Accessibility:''' Broadcast captioning, subtitle generation, podcast search, and archive indexing. | ||
| + | * '''Legal:''' Deposition and court transcription, where verbatim accuracy including disfluencies is a legal requirement, so clean-up editing is a liability rather than a feature. | ||
| + | * '''Education:''' Lecture capture, pronunciation assessment, and reading fluency tutoring for early readers. | ||
| + | * '''Automotive and Field Work:''' Hands-busy, eyes-busy environments like driving, warehouse picking, surgery, and aviation where voice is the only viable input. | ||
| + | * '''Security and Intelligence:''' Keyword spotting and mass transcription of intercepted audio, which raises the surveillance concerns noted below. | ||
| + | |||
| + | == Accessibility == | ||
| + | Speech recognition is an assistive technology before it is a convenience feature. | ||
| + | * '''Motor Impairment:''' Full voice control of a computer, as with Windows Voice Access, can be the only practical input method for users with limited hand mobility or RSI. | ||
| + | * '''Deaf and Hard of Hearing:''' Live captioning of in-person conversations, calls, and media. | ||
| + | * '''Atypical Speech:''' Mainstream ASR performs poorly on dysarthric, stuttered, accented, and child speech, precisely the users who would benefit most. [[Google]]'s Project Euphonia and Relate, Voiceitt, and the Speech Accessibility Project collect atypical speech data to close this gap, and fine-tuned models have cut error rates on dysarthric speech by roughly half versus unmodified baselines. | ||
| + | |||
| + | == How Speech Recognition Works == | ||
| + | [https://www.youtube.com/results?search_query=how+automatic+speech+recognition+works YouTube search...] | ||
| + | [https://www.google.com/search?q=how+automatic+speech+recognition+works ...Google search] | ||
| + | |||
| + | Automatic Speech Recognition (ASR) converts an acoustic waveform into text. Every system, old or new, solves the same problem: audio is continuous, variable in speed and pitch, and carries no spaces between words, while text is discrete and segmented. | ||
| + | |||
| + | === The Processing Pipeline === | ||
| + | * '''Audio Capture:''' Raw waveform sampled at 16 kHz (telephony often 8 kHz), 16-bit PCM. Sample rate mismatch is one of the most common causes of poor accuracy in production. | ||
| + | * '''Preprocessing:''' Noise suppression, echo cancellation, automatic gain control, and Voice Activity Detection to discard silence. | ||
| + | * '''Feature Extraction:''' The waveform is sliced into overlapping frames (typically 25 ms windows every 10 ms) and converted into a compact spectral representation. Historically, this meant MFCCs (Mel-Frequency Cepstral Coefficients), but today it usually means log-mel filterbank energies. Modern self-supervised models can also learn features directly from the raw waveform. | ||
| + | * '''Acoustic Modeling:''' Maps audio frames to sound units (phonemes, characters, or subword tokens). | ||
| + | * '''Language Modeling:''' Scores which word sequences are plausible, resolving homophones ("recognize speech" vs. "wreck a nice beach"). | ||
| + | * '''Decoding:''' Beam search finds the highest-scoring text hypothesis given both models. | ||
| + | * '''Post-processing:''' Punctuation and capitalization restoration, Inverse Text Normalization (turning "twenty twenty six" into "2026"), profanity filtering, and custom vocabulary boosting. | ||
| + | |||
| + | === Traditional (Hybrid) Systems === | ||
| + | Before deep learning dominated, ASR was a pipeline of separately trained components: | ||
| + | * '''DTW (Dynamic Time Warping):''' 1970s template matching that stretched and compressed audio to align it against stored reference words. | ||
| + | * '''HMM-GMM:''' Hidden Markov Models modeled the sequence of phonetic states over time, while Gaussian Mixture Models scored how well each audio frame matched each state. This was the industry standard for roughly 30 years. | ||
| + | * '''Pronunciation Lexicon:''' A hand-built dictionary mapping every word to its phoneme sequence. Out-of-vocabulary words simply couldn't be recognized. | ||
| + | * '''N-gram Language Model:''' Word-sequence probabilities estimated from large text corpora. | ||
| + | * '''DNN-HMM:''' From roughly 2012, deep neural networks replaced the GMM component, delivering the first large accuracy jump of the deep learning era. | ||
| + | |||
| + | The weakness of hybrid systems is that each component is optimized independently, so improvements in one don't necessarily improve the whole. | ||
| + | |||
| + | === End-to-End Neural Systems === | ||
| + | See also: [[End-to-End Speech]] | ||
| + | |||
| + | End-to-end (E2E) models replace the entire pipeline with a single neural network trained to map audio directly to text, eliminating the phoneme lexicon entirely. | ||
| − | + | * '''CTC (Connectionist Temporal Classification):''' Solves the alignment problem by allowing the network to emit a "blank" symbol, so it doesn't need frame-level labels. It is fast and naturally streaming, but assumes output tokens are conditionally independent, which weakens its implicit language modeling. | |
| + | * '''RNN-T (RNN Transducer):''' Adds a prediction network conditioned on previously emitted text. This is the workhorse architecture for on-device streaming dictation, including Google's and Apple's keyboards, because it emits words as you speak with low latency. | ||
| + | * '''Attention Encoder-Decoder (LAS, "Listen, Attend and Spell"):''' Uses an [[Attention]] mechanism to let the decoder look across the whole utterance. It is more accurate but inherently non-streaming, since it wants the full audio before decoding. | ||
| + | * '''[[Transformer]] and Conformer:''' The Conformer (2020) interleaves convolution blocks with self-attention, capturing both local acoustic detail and long-range context. It remains the dominant encoder design. | ||
| + | * '''Self-Supervised Learning (SSL):''' wav2vec 2.0, HuBERT, and WavLM pretrain on enormous amounts of ''unlabeled'' audio by predicting masked segments, then fine-tune on a small labeled set. This is what made high-quality ASR feasible for low-resource languages. | ||
| + | * '''Weakly Supervised Scale:''' [[OpenAI]]'s Whisper took the opposite bet, using 680,000 hours of noisy, web-scraped, ''labeled'' audio to trade data purity for robustness and zero-shot generalization. | ||
| + | * '''Speech-Augmented Language Models (SALM):''' The current frontier fuses a speech encoder directly onto a [[Large Language Model (LLM)|LLM]] decoder (for example NVIDIA's Canary-Qwen, IBM's Granite Speech, Alibaba's Qwen3-ASR). The LLM supplies world knowledge and context, sharply improving rare names, jargon, and punctuation. | ||
| − | + | === Streaming vs. Batch === | |
| + | * '''Batch (offline):''' The full recording is available. The model can use future context, so accuracy is highest. Used for podcasts, meeting recordings, and media captioning. | ||
| + | * '''Streaming (online):''' Audio arrives in chunks and partial hypotheses are emitted and then revised. Required for live captions, dictation, and voice agents. The key metric isn't just accuracy but ''word emission latency'', which is how long after you say a word it appears. | ||
| + | == Measuring Accuracy == | ||
| − | + | === Metrics === | |
| − | + | * '''Word Error Rate (WER):''' The standard metric. (Substitutions + Insertions + Deletions) ÷ total reference words. Lower is better. A WER of 5% means roughly one error every twenty words. | |
| + | * '''Character Error Rate (CER):''' Used for languages without clear word boundaries, such as Mandarin and Japanese. | ||
| + | * '''RTFx (Inverse Real-Time Factor):''' Throughput. An RTFx of 100 means one hour of audio transcribes in 36 seconds. | ||
| + | * '''Latency:''' For streaming, time-to-first-token and word emission latency matter more than throughput. Conversational voice agents generally target end-to-end response latency under a 400 millisecond threshold (the pace of natural human conversation). Modern leading models achieve sub-105ms median Time to First Audio (TTFA). | ||
| + | * '''Entity Accuracy:''' Whether the model got the phone number, email address, dollar figure, or drug name right. A transcript can have a low WER and still be useless if it fails here. | ||
| + | * '''DER (Diarization Error Rate):''' For speaker attribution. | ||
| + | |||
| + | === Why Benchmark Numbers Mislead === | ||
| + | Headline WER figures are almost always measured on clean, read speech such as LibriSpeech or TED talks. Real-world accuracy degrades sharply with background noise, accented or dialectal speech, far-field microphones, overlapping speakers, domain-specific jargon, and code-switching. A model that scores 3% WER on audiobooks can easily exceed 20% on a noisy contact-center call. Always evaluate on audio that resembles your actual use case. | ||
| + | |||
| + | === Benchmarks and Datasets === | ||
| + | * '''Open ASR Leaderboard''' (Hugging Face): Standardized text normalization, reports both WER and RTFx across English, multilingual, and long-form tracks. | ||
| + | * '''LibriSpeech''': 1,000 hours of read audiobooks. The classic benchmark, now largely saturated. | ||
| + | * '''Common Voice''' (Mozilla): Crowd-sourced, massively multilingual, openly licensed. | ||
| + | * '''Switchboard / CallHome''': Conversational telephone speech. This is the benchmark on which [[Microsoft]] claimed human parity in 2016. | ||
| + | * '''AMI Meeting Corpus''': Multi-party meetings with overlapping speech. | ||
| + | * '''CHiME Challenges''': Deliberately adversarial noisy and far-field conditions. | ||
| + | * '''FLEURS / VoxPopuli / GigaSpeech / People's Speech''': Large-scale multilingual and long-form corpora. | ||
| + | * '''Speech Accessibility Project''' (University of Illinois): Dysarthric and atypical speech from speakers with Parkinson's, ALS, cerebral palsy, and Down syndrome. | ||
| + | |||
| + | |||
| + | == Limitations and Bias == | ||
| + | * '''Demographic Disparity:''' Published audits have repeatedly found substantially higher error rates for Black speakers than white speakers on identical commercial systems, and elevated error rates for non-native accents, regional dialects, children, and elderly speakers. The cause is training data composition, not the architecture. | ||
| + | * '''Low-Resource Languages:''' Thousands of languages have effectively no transcribed audio available. SSL and massively multilingual models help, but coverage remains heavily skewed toward English and major European and East Asian languages. | ||
| + | * '''Hallucination:''' Generative ASR models, Whisper in particular, can fabricate fluent text that was never spoken, especially during silence, background noise, or non-speech audio. This is a documented failure mode with real consequences in medical and legal transcription, and it doesn't show up in average WER scores. | ||
| + | * '''The Cocktail Party Problem:''' Overlapping simultaneous speakers remain the single hardest unsolved condition. | ||
| + | * '''Context and Jargon:''' Proper nouns, acronyms, product names, and specialist vocabulary are the most common real-world errors. Custom vocabulary and keyword boosting only partly mitigate this. | ||
| + | * '''Punctuation and Formatting:''' A perfectly accurate word stream with wrong sentence boundaries is still hard to read. | ||
| + | |||
| + | == Privacy, Security, and Ethics == | ||
| + | * '''Voice as Biometric Data:''' A voiceprint is personally identifying. Several jurisdictions regulate it specifically, like Illinois' BIPA and the EU's GDPR treatment of biometric data, with meaningful penalties. | ||
| + | * '''Cloud vs. On-Device:''' The central architectural privacy decision. Cloud processing means audio leaves the device; on-device processing trades some accuracy for the guarantee that it doesn't. This is exactly the distinction between Windows Voice Typing and Voice Access described above. | ||
| + | * '''Always-On Listening:''' Wake word devices buffer audio continuously by design. Accidental activations have resulted in unintended recordings being stored and, in some cases, reviewed by human annotators. | ||
| + | * '''Recording Consent:''' Two-party consent laws in several US states and much of Europe make silently transcribing a call or meeting legally risky regardless of the technology used. | ||
| + | * '''Spoofing and Anti-Spoofing:''' Voice authentication is vulnerable to replay attacks and synthetic voice cloning, which now requires only seconds of reference audio. The ASVspoof challenge series benchmarks countermeasures. See [[Fake]]. | ||
| + | * '''Retention and Training Use:''' Whether a vendor retains audio and whether it uses customer audio to train models are separate questions, and both should be checked before deploying. | ||
| + | |||
| + | == History == | ||
| + | * '''1952''': Bell Labs' Audrey recognizes spoken digits from a single speaker. | ||
| + | * '''1962''': IBM Shoebox demonstrates 16-word recognition at the World's Fair. | ||
| + | * '''1971-76''': DARPA Speech Understanding Research; CMU's Harpy handles a 1,011-word vocabulary. | ||
| + | * '''1980s''': Hidden Markov Models become the dominant statistical framework. | ||
| + | * '''1990''': Dragon Dictate ships the first consumer dictation product; Dragon NaturallySpeaking brings continuous speech to consumers in 1997. | ||
| + | * '''2011''': Siri launches, making cloud voice assistants mainstream. | ||
| + | * '''2012''': Deep neural networks replace GMMs in acoustic models, producing the first large error-rate drop in years. | ||
| + | * '''2014-16''': Baidu's Deep Speech demonstrates end-to-end ASR; [[Microsoft]] claims human parity on the Switchboard conversational benchmark in 2016. | ||
| + | * '''2020''': wav2vec 2.0 shows that self-supervised pretraining drastically cuts the labeled-data requirement; the Conformer architecture is introduced. | ||
| + | * '''2022''': [[OpenAI]] open-sources Whisper, trained on 680,000 hours. | ||
| + | * '''2024-26''': Speech encoders fuse with [[Large Language Model (LLM)|LLMs]]; full-duplex speech-to-speech assistants and sub-100 ms streaming become standard. | ||
| − | + | == Historical Videos == | |
<youtube>u9FPqkuoEJ8</youtube> | <youtube>u9FPqkuoEJ8</youtube> | ||
Latest revision as of 10:35, 15 September 2026
YouTube ... Quora ...Google search ...Google News ...Bing News
- End-to-End Speech ... Synthesize Speech ... Speech Recognition ... Music
- Video/Image ... Vision ... Enhancement ... Fake ... Reconstruction ... Colorize ... Occlusions ... Predict image ... Image/Video Transfer Learning ... Art ... Photography
- Agents/Assistants ... Robotic Process Automation ... Personal Companions ... Productivity ... Email ... Negotiation ... LangChain
- Collective Animal Intelligence ... Animal Ecology ... Animal Language ... Bird Identification
- Large Language Model (LLM) ... Multimodal ... Foundation Models (FM) ... Generative Pre-trained ... Transformer ... Attention ... GAN ... BERT
- Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM)
- Artificial Intelligence (AI) ... Generative AI ... Machine Learning (ML) ... Deep Learning ... Neural Network ... Reinforcement ... Learning Techniques
- Conversational AI ... ChatGPT | OpenAI ... Gemini | Google ... Claude | Anthropic ... Bing/Copilot | Microsoft ... Siri | Apple ... Meta ... Perplexity ... You ... phind ... Grok | xAI ... Groq ... Ernie | Baidu ... DeepSeek ... Alibaba
- ImageBind | Meta
- OpenAI open-sources Whisper, a multilingual speech recognition system | James Vincent - The Verge
- Speechmatics Introduces Ursa: A Speech-To-Text System That Delivers Unprecedented Performance Across A Diverse Range of Voices | Tanushree Shenwai - MarkTechPost
- Open ASR Leaderboard | Hugging Face ...Standardized WER and RTFx comparison across open and proprietary systems.
- TTS Latency Benchmark 2026 | Coval ...Industry benchmarks showing top models (like Palabra v1) achieving sub-105ms latency for Voice AI generation.
- Advances in Real-Time Transcription Latency for Transformer-Based ASR | Research Team - SpeechTech Journal - January 2026 ...New benchmarks showing sub-100ms latency in streaming E2E models.
- Achieving Human Parity in Noisy Environments | Google DeepMind Team - Google AI Blog - November 2025 ...Breakthroughs in robust speech recognition accuracy under extreme background noise.
If you want direct control over your PC's user interface or the ability to dictate perfectly formatted text into any active window, Windows 11 Voice Access and Wispr Flow Pro are your primary tools. Voice Access acts as your hands, navigating the operating system and clicking buttons using commands like "Open Edge." Wispr Flow Pro acts as your real-time editor, instantly cleaning up your grammar and removing filler words before the text ever hits your active cursor.
In contrast, Gemini, Claude, ChatGPT, and Microsoft Copilot operate as deeply integrated workflow agents rather than universal mouse-and-keyboard controllers. They've evolved past the need for manual copy-and-pasting by executing actions directly within their connected environments:
- Microsoft Copilot: Built directly into Windows 11 and Edge, it summarizes active browser documents, fetches live web data, and executes workflows natively across your Microsoft 365 applications, acting as a collaborative AI workforce.
- Google Gemini: Using the Spark agent, it autonomously manages multi-step tasks, drafts documents, and organizes data directly across your connected Google Workspace apps.
- Anthropic Claude: Through Claude Code, it integrates directly into your terminal to read local files, analyze your codebase, and execute commands via push-to-talk.
- OpenAI ChatGPT: The desktop app orchestrates background tasks and interacts with local environments, letting you trigger code tests or workflow scripts while maintaining a voice conversation.
While they won't puppet your mouse to click on a third-party application the way Voice Access does, they act as active collaborators executing complex logic and tasks in the background.
| Tool | Primary Role | Controls Your OS? | Works Offline? | Computer & Browser Use Agent | Best For |
|---|---|---|---|---|---|
| Windows Voice Typing (Win+H) | Dictation | No | No | None | Quick drafting into any text field |
| Windows Voice Access | Full PC control | Yes | Yes | Native Windows OS Accessibility | Hands-free operation, privacy, accessibility |
| Wispr Flow | Dictation + auto-editing | No | No | None | Clean first-draft text anywhere your cursor is |
| Copilot | Workflow agent | Partial (Windows/M365) | No | Copilot Studio Agents / Edge | Microsoft 365 and Edge workflows |
| Gemini Live + Spark | Conversational + background agent | No | No | Gemini Spark | Multimodal camera/screen input, Workspace tasks |
| Claude | Conversational + terminal | No | No | Claude Code / Computer Use API | Deep work, codebase discussion, push-to-talk |
| ChatGPT Voice | Conversational (full-duplex) | No | No | ChatGPT Desktop App / Codex | Natural interruptible dialogue, live translation |
| Whisper (local) | Batch transcription | No | Yes | None | Private, offline transcription of recordings |
Related Speech Tasks
Production speech systems are almost never "just" transcription. Several supporting models usually run alongside the recognizer.
- Voice Activity Detection (VAD): Decides whether a frame contains speech at all. It is cheap, fast, and acts as the gatekeeper for everything downstream.
- Wake Word / Keyword Spotting: A tiny always-on model listening for "Hey Google" or "Alexa" without sending audio anywhere. It must run on milliwatts of power.
- Speaker Diarization: Answers "who spoke when," segmenting a recording by speaker. Overlapping speech remains the hard failure case. Measured with Diarization Error Rate (DER).
- Speaker Identification / Verification: Matching a voice to a known identity (biometric authentication), distinct from diarization, which only distinguishes speakers anonymously.
- Language Identification (LID): Determining which language is being spoken, necessary before choosing a recognizer.
- Code-Switching: Handling mid-sentence language changes (Hinglish, Spanglish). This is still a significant weakness in most commercial systems.
- Speech Translation: Either cascaded (ASR to machine translation to TTS) or direct speech-to-text/speech-to-speech models such as Meta's SeamlessM4T. See Language Translation.
- Audio-Visual Speech Recognition: Adding lip movement as a second modality (AV-HuBERT), dramatically improving accuracy in noise. See Vision.
- Paralinguistics: Emotion, sentiment, age, and health markers inferred from voice rather than words.
Microsoft Voice Ecosystem: Copilot and Windows 11
Microsoft divides its voice capabilities into two distinct categories on your PC. Copilot acts as your conversational AI agent, while the Windows 11 built-in tools serve as your direct dictation and operating system controllers.
Microsoft Copilot (formerly Bing Chat) Voice
Microsoft unified its AI efforts under the Copilot brand, retiring the "Bing Chat" name. Copilot's voice capabilities run on Azure AI Speech infrastructure, offering a blend of web-grounded search and natural conversation. Unlike standalone dictation tools, Copilot acts as an integrated assistant across your workflows.
Desktop and Edge Integration
On your laptop, you'll find Copilot built directly into Windows 11 and the Microsoft Edge browser.
- Windows Integration: Open the Copilot sidebar and click the microphone icon to speak your prompts. It handles complex queries, like asking it to summarize a PDF you have open in Edge or searching the web for live news.
- Not a Dictation Tool: Copilot is a conversational agent, not a universal dictation tool like Wispr Flow or Voice Access. You speak to it to generate ideas, code, or answers, but you still need to copy and paste its text output into your target applications.
- Custom AI Agents (2026): Using Copilot Studio, businesses now deploy multi-agent collaborations. Agents act as an "AI workforce," automating complex workflows in the background rather than just responding to chat.
Mobile App Experience
The Copilot mobile app on your Pixel phone offers a robust voice interface that works perfectly for on-the-go research.
- Voice Search and Read Aloud: Ask a question out loud, and Copilot searches the live web and speaks the answer back to you.
- Image and Voice Combo: Take a photo with your phone and use your voice to ask Copilot what you're looking at, making it a highly effective multimodal tool.
Copilot Voice (Conversational Mode)
Microsoft rolled out a more fluid, two-way conversational mode for Copilot, designed to compete directly with ChatGPT's advanced voice features.
- Fluid Dialog: Instead of a turn-based walkie-talkie style, you can hold a continuous conversation. It picks up on natural pauses and lets you interrupt if the AI goes off track.
- Tone and Pacing: The system adapts to your pacing and offers different voice personas, making it feel more like a brainstorming partner than a rigid search engine.
Windows 11 Built-in Voice Tools
Windows 11 actually offers two distinct built-in tools for voice: Voice Typing for simple text dictation, and Voice Access for complete hands-free control of your computer. Both are free and fully integrated, but they handle audio and execute tasks very differently.
Voice Typing (Win + H)
This is the quick, everyday dictation tool designed specifically for writing.
- How it Works: You press
Windows Key + H, click into any text box (like an email, Word document, or browser search), and start talking. The tool types whatever you say directly into the active field. - Cloud Processing: Voice Typing relies on Microsoft's Azure servers to process speech. This means you must have an active internet connection to use it, and your audio data is sent to the cloud for transcription.
- Capabilities: It automatically recognizes pauses for punctuation, or you can say commands out loud (e.g., "comma," "new line"). However, it only types text; it cannot click buttons or open applications.
- New Filtering: The tool now includes a dedicated profanity filter toggle for tighter control over dictation output.
Voice Access
Voice Access is a comprehensive accessibility system built to let you operate your entire PC without touching a mouse or keyboard.
- Total PC Control: Where Voice Typing only writes, Voice Access navigates. You can command your operating system with phrases like "Open Edge," "Click File," "Scroll down," or "Press Enter".
- Offline Processing: Once you download the initial language model during setup, Voice Access runs completely locally on your device. It works offline and never transmits your audio to the cloud, making it excellent for privacy.
- Integrated Dictation: It includes its own robust dictation system. You can use voice commands to navigate to a text field, and then seamlessly start dictating text, allowing for a fully hands-free workflow.
- 2026 Upgrades: Recent updates added Voice Isolation to improve accuracy in noisy rooms and expanded natural language command support.
How to Launch Voice Access:
- Windows Search (Fastest): Press the Windows key, type "voice access", and press Enter. The control bar will immediately drop down from the top of your screen.
- Through the Settings Menu: Navigate to Settings > Accessibility > Speech, and turn the Voice access toggle to On.
- Start Automatically on Boot: If you plan to use it daily, you can tell Windows to have it ready as soon as you turn on your computer. Go to Settings > Accessibility > Speech and check the box for "Start voice access after you sign in to your PC."
Which Should You Use?
If you just need to quickly draft a document or respond to a chat, press Win + H and use Voice Typing. If you want to lean back from your desk, navigate through applications, or if you need to work entirely offline, turn on Voice Access.
Wispr Flow
Wispr Flow is an AI-powered dictation tool that goes beyond basic speech-to-text. Instead of typing exactly what you say, it acts as a real-time editor. If you stumble, use filler words, or change your mind mid-sentence, Flow cleans up the output before the text hits your screen. It works universally across Windows, macOS, iOS, and Android, typing wherever you place your cursor.
Core Features
- Smart Auto-Editing: Flow understands backtracking. If you say, "Let's meet Tuesday, wait, actually make it Thursday," it simply types "Let's meet Thursday." It also automatically catches punctuation from your natural pauses.
- Personal Dictionary: You can teach it your specific terminology. It learns difficult names, brand terms, and acronyms, so you don't have to manually fix the same typos every day.
- Voice Snippets: Think of these as spoken text expanders. You can say a trigger phrase like "insert meeting link," and Flow pastes your full scheduling URL and standard greeting.
- Context-Aware Tone: It looks at the active window to adjust how it formats your words. It writes casually when you're in a Slack window, but switches to a formal structure when you open an email client. A 2026 update added a Personalized Style setting to adjust tone per app category from "Very Casual" to "Formal".
- Whisper Mode: The engine handles very quiet speech, so you can dictate in an office or coffee shop without disturbing the people around you.
- Extended Sessions: As of early 2026, you can dictate continuously for up to 20 minutes without interruption.
Developer and Coding Support
Flow is heavily used by developers because it handles technical jargon seamlessly.
- Syntax Awareness: It knows the difference between conversational English and code. It automatically formats variables in camelCase or snake_case and preserves correct spacing for command-line instructions.
- Smart IDE Tagging: If you use AI editors like Cursor or Windsurf, Flow recognizes when you say a filename out loud and automatically tags that file in your prompt workspace.
Google Voice Tools
Google's main voice tool is Gemini Live. It replaces the older Google Assistant with a system that handles real conversation and understands multiple types of media. Instead of answering one question at a time, you can talk with Gemini Live back and forth. You can interrupt it, change the subject, or brainstorm out loud just like you would with a friend.
Multimodal Screen and Camera Vision
If you use a Pixel smartphone, Gemini Live can see what you're looking at in real time.
- Live Camera Feed: You can open your camera through the Gemini app and point it at something in the real world. For example, point your phone at a confusing wiring diagram or a broken toaster, and ask for troubleshooting help while it watches the live video.
- Screen Sharing: You can also share your phone's screen. If you're comparing flight prices or reading a long article, Gemini Live can summarize the text or answer questions about what's on your screen without you typing a single word. Think of it like someone looking over your shoulder to help you read a map.
Agentic Productivity Features
In 2026, Google upgraded Gemini Live so it can manage tasks in the background and work directly with Google Workspace.
- Hands-Free Workspace Control: You can use your voice to tell Gemini to search your Gmail, summarize emails, or organize ideas into Docs and Sheets.
- Spark Integration: You can hand off complex, multi-step projects to a background agent named Spark. Spark does the heavy lifting across your Google apps while you move on to other things.
- Personal Intelligence: The system remembers what you talked about before and pulls information from apps like Google Photos, YouTube, and Calendar. It connects the dots so you don't have to repeat yourself.
- Daily Briefs: Start your day by asking for a daily brief. The assistant reads out a custom audio summary of your calendar events and important emails, similar to a morning news radio show tailored just for you.
Gemini Spark Execution
Gemini Live handles the talking, and Gemini Spark does the actual work. Gemini Live is the voice you interact with, while Spark runs in the background to complete tasks across your apps.
Think of Live as the receptionist sitting at the front desk taking your requests, and Spark as the back-office manager who organizes the files, sends the emails, and follows up on the details.
Core Capabilities
- Unstructured Voice Input (Live): You don't need to speak in rigid commands. You can think out loud, pause, or change your mind mid-sentence. Live figures out what you mean and organizes the request.
- Autonomous Handoff (Spark): Once you confirm what you want to do, Live hands the job to Spark. Unlike a normal chat that stops working when you close the app, Spark keeps running on Google's cloud servers until the job is done.
- Cross-App Execution: Spark connects directly to Google Workspace apps like Gmail, Calendar, Docs, Sheets, and Drive. It finds, changes, and moves data between these services without you having to watch over it.
Workflow Example
- Voice Request via Live: You say something like, "I need to plan next week's dinners. Keep it high-protein, pull that chicken recipe I saved from my email last Thursday, and make a grocery list in a new Doc."
- Interactive Confirmation: Live repeats the plan to make sure it got it right and asks you to clarify anything that's missing.
- Background Execution via Spark: Spark goes into your Gmail to find the recipe, schedules the meals, checks the ingredients you need, builds the grocery list, and creates the Google Doc for you.
Requirements
- Gemini Live: Built right into the Gemini mobile app for Android and iOS.
- Gemini Spark: To run these full background workflows, you usually need a paid Google AI advanced plan and you need to turn on workspace extensions.
Gemini 3.5 Transcribe
Released in August 2026, Gemini 3.5 Transcribe is Google's dedicated speech-to-text model. While Gemini Live focuses on back-and-forth conversation, Transcribe specializes in turning spoken words into clean, accurate text. Think of it as a highly skilled stenographer who can also edit your rough drafts on the fly.
- Smart Transcription: This mode acts like a built-in editor. Instead of typing out every filler word, it drops the "ums" and "ahs", fixes mid-sentence corrections, and formats the text. For instance, if you say, "Let's meet on Tuesday, wait, I mean Wednesday," it will simply output, "Let's meet on Wednesday."
- Automatic Language Detection: The model recognizes over 85 languages and automatically follows along if a speaker switches languages in the middle of a sentence.
- Custom Vocabulary: You can load up to 1,000 specific terms, acronyms, or proper names so the model spells your industry jargon correctly.
- Speaker Diarization and Timestamps: For recorded audio like meetings, it can identify up to eight different speakers and provide exact start and end times for every single word.
Integrations
You can find Gemini 3.5 Transcribe working quietly behind the scenes in a few places:
- Rambler on Android: Built into Gboard, this feature lets you dictate messy, unstructured thoughts and turns them into well-formatted text messages or notes.
- Gemini App on macOS: You can use your voice to dictate directly to your computer or use voice commands alongside your screen context to build complex workflows.
- Google Antigravity and AI Studio: It pairs with screen context to make sure file names and technical terms are transcribed perfectly when you're building apps or navigating agent workflows.
Claude Voice Tools
In July 2026, Anthropic overhauled Claude's voice capabilities, shifting from basic dictation to a full, two-way interactive voice mode. You can now speak to Claude and have it speak back in real time, making it highly effective for brainstorming or hands-free coding architecture discussions.
The Two Core Voice Modes
Claude offers two primary ways to interact with it using audio, depending on your environment:
- Hands-Free Mode: This operates like a phone call. You speak, Claude listens, and it responds automatically. This works best in quiet environments where the microphone won't pick up background chatter.
- Push-to-Talk Mode: If you are in a crowded room or a noisy street, you can hold down a button on your screen while you speak and release it when you're done. This gives you total control over exactly when Claude is listening so it doesn't get confused by background noise.
Unique Workflow Features
Claude's voice implementation is designed heavily around deep work and seamless transitions rather than just casual conversation.
- Seamless Modality Switching: You can switch back and forth between typing and talking in the exact same chat thread without losing context. You can start a conversation on your Pixel while walking the dog, and then sit down at your Alienware laptop to type out a specific block of code in the same session.
- Connected Tools via Voice: If you've linked your Google Workspace (Gmail, Calendar, Docs) or Slack, you can ask Claude to read your emails or summarize a document out loud during a voice session.
- Claude Code (Terminal Voice): For developers, Anthropic added a push-to-talk voice mode directly into the Claude Code command-line interface. You can hold the spacebar in your terminal, describe a bug or ask how a module works, and the agent will investigate your local codebase and respond with text.
ChatGPT Voice (GPT-Live)
ChatGPT Voice is built for natural, human-like conversation. In July 2026, OpenAI upgraded the default engine to GPT-Live, making the system full-duplex. This means the AI listens and speaks at the same time. You don't have to wait for it to finish a paragraph; you can interrupt it mid-sentence, laugh, or change the subject, and it instantly adapts its response and tone.
Windows 11 Desktop Integration
While tools like Voice Access control your computer, ChatGPT Voice operates as a powerful consultant that sits on top of your workflow.
- The Desktop App: You can trigger the voice interface using a keyboard shortcut (like
Alt + Space) to start talking without leaving your current window. - Screen Context: The Windows desktop app includes a screen vision feature. If you're looking at a complex spreadsheet or a block of code, you can ask ChatGPT to "take a look at this." It captures the active window so it understands exactly what you're talking about.
- App Isolation: ChatGPT Voice stays inside its own application ecosystem. It will format code or draft emails for you, but you must manually copy and paste that text into your external programs.
Core Capabilities
- Agent Orchestration: If you use ChatGPT Work or Codex, you can use your voice to direct background agents. You can tell the desktop app to "start a task to run the tests," and it coordinates the work while you keep talking.
- Live Web Search: The GPT-Live update integrated web searching directly into voice sessions. You can ask for current flight prices or news updates, and it pulls live data without forcing you back to the text interface.
- Real-Time Translation: The system acts as a live interpreter for over 50 languages, maintaining the speaker's pace and natural inflection.
Apple Voice Tools
- On-Device Dictation: Since iOS 15, standard keyboard dictation runs locally on the Neural Engine for many languages, with no audio leaving the device and no time limit.
- SpeechAnalyzer (iOS 26+): Apple's modern speech framework, replacing the legacy SFSpeechRecognizer and its one-minute session cap. It ships separate modules for long-form transcription, short utterances, and voice activity detection, and exposes confidence scores, word timings, and language detection.
- Siri and Apple Intelligence: Voice requests are routed between on-device processing and Private Cloud Compute depending on complexity. Apps expose voice-triggerable actions through the App Intents framework rather than legacy Shortcuts donations.
- Live Captions: System-wide real-time captioning of any audio, processed on-device, built as an accessibility feature.
Other Assistant Ecosystems
- Amazon Alexa: The largest installed base of far-field, always-listening hardware. Far-field ASR (which includes beamforming across a microphone array, acoustic echo cancellation while music plays, and barge-in during playback) is a materially harder engineering problem than close-talk phone dictation.
- Samsung Bixby and Galaxy AI: On-device call transcription, live translation, and voice recorder summarization across Galaxy devices.
- Smart glasses and wearables: Voice as the primary input modality when there is no screen and no keyboard, which is a growing driver of low-power, always-available ASR.
Voice Agents and Speech-to-Speech
The assistants described on this page are built one of two ways, and the choice determines how they feel to use.
The Cascaded Pipeline
Audio to ASR to LLM to TTS to audio. Each stage is swappable and independently debuggable, and you get a text transcript for free. The cost is latency (three models in series) and information loss. Tone, emphasis, sarcasm, and emotion are discarded the moment speech becomes text.
Native Speech-to-Speech
A single multimodal model consumes audio tokens and emits audio tokens directly, never fully committing to intermediate text. This preserves prosody and enables full-duplex behavior, as the model can listen while speaking, so it handles interruptions and backchannels naturally. This is the architecture behind ChatGPT's advanced voice mode and Gemini Live. The trade-offs are weaker controllability, harder evaluation, and no reliable transcript unless one is generated separately.
Engineering Challenges
- Turn Detection / Endpointing: Deciding when the user has actually finished speaking. Naive silence thresholds cut off anyone who pauses to think; neural turn detection reads intonation and syntax instead.
- Barge-In: Letting the user interrupt mid-response, which requires cancelling generation and playback instantly while avoiding the system transcribing its own output.
- Latency Budget: Natural conversation tolerates roughly 300 ms of silence before it feels broken. That budget must cover network round-trip, transcription, LLM inference, and speech synthesis.
Whisper
YouTube search... ...Google search
Whisper is an Automatic Speech Recognition Service (ASR) by OpenAI trained on 680,000 hours of multilingual and multitask supervised data collected from the web. We've trained and are open-sourcing a neural net called Whisper that approaches human level robustness and accuracy on English speech recognition.
Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification. The Whisper v2-large model is currently available through our API with the whisper-1 model name. Currently, there isn't a difference between the open source version of Whisper and the version available through our API. However, through our API, we offer an optimized inference process which makes running Whisper through our API much faster than doing it through other means. For more technical details on Whisper, you can read the paper. - OpenAI
Open Source Models and Toolkits
Modern Open-Weight Models
- Whisper (OpenAI): See above. large-v3, and the distilled large-v3-turbo, remain the most widely deployed open ASR models purely on ecosystem gravity.
- NVIDIA Parakeet: FastConformer with CTC or TDT decoding. Optimized for raw speed; among the fastest models on the Open ASR Leaderboard, well suited to live captioning and telephony.
- NVIDIA Canary / Canary-Qwen: A speech encoder bolted to a Qwen LLM decoder; has held the top accuracy slot on the Open ASR Leaderboard.
- IBM Granite Speech: LoRA-tuned on the Granite LLM, trained with synthetic noise injection for robustness.
- Qwen3-ASR (Alibaba): The widest open multilingual coverage by language count.
- Moonshine (Useful Sensors): Purpose-built for edge devices; runs on a Raspberry Pi.
- Meta MMS / SeamlessM4T: Extreme multilingual coverage (1,000+ languages for MMS) and unified translation.
- wav2vec 2.0 / HuBERT / WavLM: Foundational SSL encoders, generally fine-tuned rather than used directly.
Runtimes and Frameworks
- whisper.cpp: C/C++ port of Whisper with no Python dependency; the standard for local desktop and mobile deployment.
- faster-whisper: CTranslate2 reimplementation, several times faster than the reference code at equal accuracy.
- WhisperX: Adds forced alignment for accurate word-level timestamps plus diarization.
- WhisperKit: On-device Whisper optimized for the Apple Neural Engine.
- NVIDIA NeMo: Training and inference framework behind Parakeet and Canary.
- SpeechBrain / ESPnet: PyTorch research toolkits covering ASR, diarization, and enhancement.
- Kaldi: The C++ toolkit that defined the hybrid HMM-DNN era; still in use in production and research.
- Vosk: Lightweight offline recognizer with bindings for most languages.
- pyannote.audio: The de facto open diarization library.
Commercial ASR APIs
Consumer assistants sit on top of engines that are also sold directly to developers.
| Provider | Product | Notable For |
|---|---|---|
| Google Cloud | Speech-to-Text / Chirp | Broad language coverage, Workspace and Android integration |
| Microsoft Azure | Azure AI Speech | The engine behind Windows Voice Typing and Copilot; strong enterprise compliance story |
| Amazon AWS | Transcribe | Call analytics, medical variant, AWS-native |
| OpenAI | Whisper API, gpt-4o-transcribe, Realtime API | Ecosystem reach and speech-to-speech |
| Deepgram | Nova | Low-latency streaming, voice agents |
| AssemblyAI | Universal | Speaker turn detection, entity accuracy, audio intelligence features |
| Speechmatics | Ursa | Accent and dialect robustness across diverse voices |
| ElevenLabs | Scribe | Diarization and high-accuracy batch transcription |
| Picovoice | Cheetah / Leopard | Fully on-device, privacy-preserving, embedded targets |
| Rev, Otter.ai, Fireflies, Sonix | Meeting and media transcription | Human-in-the-loop options and workflow integration |
Applications
- Healthcare: Ambient clinical documentation is the highest-value commercial ASR market. Systems listen to the doctor-patient conversation and draft the clinical note, addressing physician burnout from EHR data entry. Requires medical vocabulary, speaker diarization, and HIPAA compliance.
- Contact Centers: Real-time agent assist, automated QA across every call rather than a sampled few, and compliance monitoring.
- Media and Accessibility: Broadcast captioning, subtitle generation, podcast search, and archive indexing.
- Legal: Deposition and court transcription, where verbatim accuracy including disfluencies is a legal requirement, so clean-up editing is a liability rather than a feature.
- Education: Lecture capture, pronunciation assessment, and reading fluency tutoring for early readers.
- Automotive and Field Work: Hands-busy, eyes-busy environments like driving, warehouse picking, surgery, and aviation where voice is the only viable input.
- Security and Intelligence: Keyword spotting and mass transcription of intercepted audio, which raises the surveillance concerns noted below.
Accessibility
Speech recognition is an assistive technology before it is a convenience feature.
- Motor Impairment: Full voice control of a computer, as with Windows Voice Access, can be the only practical input method for users with limited hand mobility or RSI.
- Deaf and Hard of Hearing: Live captioning of in-person conversations, calls, and media.
- Atypical Speech: Mainstream ASR performs poorly on dysarthric, stuttered, accented, and child speech, precisely the users who would benefit most. Google's Project Euphonia and Relate, Voiceitt, and the Speech Accessibility Project collect atypical speech data to close this gap, and fine-tuned models have cut error rates on dysarthric speech by roughly half versus unmodified baselines.
How Speech Recognition Works
YouTube search... ...Google search
Automatic Speech Recognition (ASR) converts an acoustic waveform into text. Every system, old or new, solves the same problem: audio is continuous, variable in speed and pitch, and carries no spaces between words, while text is discrete and segmented.
The Processing Pipeline
- Audio Capture: Raw waveform sampled at 16 kHz (telephony often 8 kHz), 16-bit PCM. Sample rate mismatch is one of the most common causes of poor accuracy in production.
- Preprocessing: Noise suppression, echo cancellation, automatic gain control, and Voice Activity Detection to discard silence.
- Feature Extraction: The waveform is sliced into overlapping frames (typically 25 ms windows every 10 ms) and converted into a compact spectral representation. Historically, this meant MFCCs (Mel-Frequency Cepstral Coefficients), but today it usually means log-mel filterbank energies. Modern self-supervised models can also learn features directly from the raw waveform.
- Acoustic Modeling: Maps audio frames to sound units (phonemes, characters, or subword tokens).
- Language Modeling: Scores which word sequences are plausible, resolving homophones ("recognize speech" vs. "wreck a nice beach").
- Decoding: Beam search finds the highest-scoring text hypothesis given both models.
- Post-processing: Punctuation and capitalization restoration, Inverse Text Normalization (turning "twenty twenty six" into "2026"), profanity filtering, and custom vocabulary boosting.
Traditional (Hybrid) Systems
Before deep learning dominated, ASR was a pipeline of separately trained components:
- DTW (Dynamic Time Warping): 1970s template matching that stretched and compressed audio to align it against stored reference words.
- HMM-GMM: Hidden Markov Models modeled the sequence of phonetic states over time, while Gaussian Mixture Models scored how well each audio frame matched each state. This was the industry standard for roughly 30 years.
- Pronunciation Lexicon: A hand-built dictionary mapping every word to its phoneme sequence. Out-of-vocabulary words simply couldn't be recognized.
- N-gram Language Model: Word-sequence probabilities estimated from large text corpora.
- DNN-HMM: From roughly 2012, deep neural networks replaced the GMM component, delivering the first large accuracy jump of the deep learning era.
The weakness of hybrid systems is that each component is optimized independently, so improvements in one don't necessarily improve the whole.
End-to-End Neural Systems
See also: End-to-End Speech
End-to-end (E2E) models replace the entire pipeline with a single neural network trained to map audio directly to text, eliminating the phoneme lexicon entirely.
- CTC (Connectionist Temporal Classification): Solves the alignment problem by allowing the network to emit a "blank" symbol, so it doesn't need frame-level labels. It is fast and naturally streaming, but assumes output tokens are conditionally independent, which weakens its implicit language modeling.
- RNN-T (RNN Transducer): Adds a prediction network conditioned on previously emitted text. This is the workhorse architecture for on-device streaming dictation, including Google's and Apple's keyboards, because it emits words as you speak with low latency.
- Attention Encoder-Decoder (LAS, "Listen, Attend and Spell"): Uses an Attention mechanism to let the decoder look across the whole utterance. It is more accurate but inherently non-streaming, since it wants the full audio before decoding.
- Transformer and Conformer: The Conformer (2020) interleaves convolution blocks with self-attention, capturing both local acoustic detail and long-range context. It remains the dominant encoder design.
- Self-Supervised Learning (SSL): wav2vec 2.0, HuBERT, and WavLM pretrain on enormous amounts of unlabeled audio by predicting masked segments, then fine-tune on a small labeled set. This is what made high-quality ASR feasible for low-resource languages.
- Weakly Supervised Scale: OpenAI's Whisper took the opposite bet, using 680,000 hours of noisy, web-scraped, labeled audio to trade data purity for robustness and zero-shot generalization.
- Speech-Augmented Language Models (SALM): The current frontier fuses a speech encoder directly onto a LLM decoder (for example NVIDIA's Canary-Qwen, IBM's Granite Speech, Alibaba's Qwen3-ASR). The LLM supplies world knowledge and context, sharply improving rare names, jargon, and punctuation.
Streaming vs. Batch
- Batch (offline): The full recording is available. The model can use future context, so accuracy is highest. Used for podcasts, meeting recordings, and media captioning.
- Streaming (online): Audio arrives in chunks and partial hypotheses are emitted and then revised. Required for live captions, dictation, and voice agents. The key metric isn't just accuracy but word emission latency, which is how long after you say a word it appears.
Measuring Accuracy
Metrics
- Word Error Rate (WER): The standard metric. (Substitutions + Insertions + Deletions) ÷ total reference words. Lower is better. A WER of 5% means roughly one error every twenty words.
- Character Error Rate (CER): Used for languages without clear word boundaries, such as Mandarin and Japanese.
- RTFx (Inverse Real-Time Factor): Throughput. An RTFx of 100 means one hour of audio transcribes in 36 seconds.
- Latency: For streaming, time-to-first-token and word emission latency matter more than throughput. Conversational voice agents generally target end-to-end response latency under a 400 millisecond threshold (the pace of natural human conversation). Modern leading models achieve sub-105ms median Time to First Audio (TTFA).
- Entity Accuracy: Whether the model got the phone number, email address, dollar figure, or drug name right. A transcript can have a low WER and still be useless if it fails here.
- DER (Diarization Error Rate): For speaker attribution.
Why Benchmark Numbers Mislead
Headline WER figures are almost always measured on clean, read speech such as LibriSpeech or TED talks. Real-world accuracy degrades sharply with background noise, accented or dialectal speech, far-field microphones, overlapping speakers, domain-specific jargon, and code-switching. A model that scores 3% WER on audiobooks can easily exceed 20% on a noisy contact-center call. Always evaluate on audio that resembles your actual use case.
Benchmarks and Datasets
- Open ASR Leaderboard (Hugging Face): Standardized text normalization, reports both WER and RTFx across English, multilingual, and long-form tracks.
- LibriSpeech: 1,000 hours of read audiobooks. The classic benchmark, now largely saturated.
- Common Voice (Mozilla): Crowd-sourced, massively multilingual, openly licensed.
- Switchboard / CallHome: Conversational telephone speech. This is the benchmark on which Microsoft claimed human parity in 2016.
- AMI Meeting Corpus: Multi-party meetings with overlapping speech.
- CHiME Challenges: Deliberately adversarial noisy and far-field conditions.
- FLEURS / VoxPopuli / GigaSpeech / People's Speech: Large-scale multilingual and long-form corpora.
- Speech Accessibility Project (University of Illinois): Dysarthric and atypical speech from speakers with Parkinson's, ALS, cerebral palsy, and Down syndrome.
Limitations and Bias
- Demographic Disparity: Published audits have repeatedly found substantially higher error rates for Black speakers than white speakers on identical commercial systems, and elevated error rates for non-native accents, regional dialects, children, and elderly speakers. The cause is training data composition, not the architecture.
- Low-Resource Languages: Thousands of languages have effectively no transcribed audio available. SSL and massively multilingual models help, but coverage remains heavily skewed toward English and major European and East Asian languages.
- Hallucination: Generative ASR models, Whisper in particular, can fabricate fluent text that was never spoken, especially during silence, background noise, or non-speech audio. This is a documented failure mode with real consequences in medical and legal transcription, and it doesn't show up in average WER scores.
- The Cocktail Party Problem: Overlapping simultaneous speakers remain the single hardest unsolved condition.
- Context and Jargon: Proper nouns, acronyms, product names, and specialist vocabulary are the most common real-world errors. Custom vocabulary and keyword boosting only partly mitigate this.
- Punctuation and Formatting: A perfectly accurate word stream with wrong sentence boundaries is still hard to read.
Privacy, Security, and Ethics
- Voice as Biometric Data: A voiceprint is personally identifying. Several jurisdictions regulate it specifically, like Illinois' BIPA and the EU's GDPR treatment of biometric data, with meaningful penalties.
- Cloud vs. On-Device: The central architectural privacy decision. Cloud processing means audio leaves the device; on-device processing trades some accuracy for the guarantee that it doesn't. This is exactly the distinction between Windows Voice Typing and Voice Access described above.
- Always-On Listening: Wake word devices buffer audio continuously by design. Accidental activations have resulted in unintended recordings being stored and, in some cases, reviewed by human annotators.
- Recording Consent: Two-party consent laws in several US states and much of Europe make silently transcribing a call or meeting legally risky regardless of the technology used.
- Spoofing and Anti-Spoofing: Voice authentication is vulnerable to replay attacks and synthetic voice cloning, which now requires only seconds of reference audio. The ASVspoof challenge series benchmarks countermeasures. See Fake.
- Retention and Training Use: Whether a vendor retains audio and whether it uses customer audio to train models are separate questions, and both should be checked before deploying.
History
- 1952: Bell Labs' Audrey recognizes spoken digits from a single speaker.
- 1962: IBM Shoebox demonstrates 16-word recognition at the World's Fair.
- 1971-76: DARPA Speech Understanding Research; CMU's Harpy handles a 1,011-word vocabulary.
- 1980s: Hidden Markov Models become the dominant statistical framework.
- 1990: Dragon Dictate ships the first consumer dictation product; Dragon NaturallySpeaking brings continuous speech to consumers in 1997.
- 2011: Siri launches, making cloud voice assistants mainstream.
- 2012: Deep neural networks replace GMMs in acoustic models, producing the first large error-rate drop in years.
- 2014-16: Baidu's Deep Speech demonstrates end-to-end ASR; Microsoft claims human parity on the Switchboard conversational benchmark in 2016.
- 2020: wav2vec 2.0 shows that self-supervised pretraining drastically cuts the labeled-data requirement; the Conformer architecture is introduced.
- 2022: OpenAI open-sources Whisper, trained on 680,000 hours.
- 2024-26: Speech encoders fuse with LLMs; full-duplex speech-to-speech assistants and sub-100 ms streaming become standard.
Historical Videos