Difference between revisions of "Speech Recognition"
(→Automatic Speech Recognition (ASR)) |
m (→ChatGPT Voice) |
||
| Line 84: | Line 84: | ||
<youtube>Bg0rdpFKCQI</youtube> | <youtube>Bg0rdpFKCQI</youtube> | ||
| − | + | ||
| − | + | = ChatGPT Voice (GPT-Live) = | |
ChatGPT Voice is built for natural, human-like conversation. In July 2026, OpenAI upgraded the default engine to GPT-Live, making the system full-duplex. This means the AI listens and speaks at the same time. You do not have to wait for it to finish a paragraph; you can interrupt it mid-sentence, laugh, or change the subject, and it instantly adapts its response and tone. | ChatGPT Voice is built for natural, human-like conversation. In July 2026, OpenAI upgraded the default engine to GPT-Live, making the system full-duplex. This means the AI listens and speaks at the same time. You do not have to wait for it to finish a paragraph; you can interrupt it mid-sentence, laugh, or change the subject, and it instantly adapts its response and tone. | ||
| Line 101: | Line 101: | ||
<youtube>jKkr8czmG4U</youtube> | <youtube>jKkr8czmG4U</youtube> | ||
| − | |||
= <span id="Automatic Speech Recognition (ASR)"></span>Automatic Speech Recognition (ASR) = | = <span id="Automatic Speech Recognition (ASR)"></span>Automatic Speech Recognition (ASR) = | ||
Revision as of 00:26, 14 September 2026
YouTube ... Quora ...Google search ...Google News ...Bing News
- End-to-End Speech ... Synthesize Speech ... Speech Recognition ... Music
- Video/Image ... Vision ... Enhancement ... Fake ... Reconstruction ... Colorize ... Occlusions ... Predict image ... Image/Video Transfer Learning ... Art ... Photography
- Agents ... Robotic Process Automation ... Assistants ... Personal Companions ... Productivity ... Email ... Negotiation ... LangChain
- Collective Animal Intelligence ... Animal Ecology ... Animal Language ... Bird Identification
- Large Language Model (LLM) ... Natural Language Processing (NLP) ...Generation ... Classification ... Understanding ... Translation ... Tools & Services
- Attention Mechanism ...Transformer ...Generative Pre-trained Transformer (GPT) ... GAN ... BERT
- Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM)
- Artificial Intelligence (AI) ... Generative AI ... Machine Learning (ML) ... Deep Learning ... Neural Network ... Reinforcement ... Learning Techniques
- Conversational AI ... ChatGPT | OpenAI ... Bing/Copilot | Microsoft ... Gemini | Google ... Claude | Anthropic ... Perplexity ... You ... phind ... Ernie | Baidu
- ImageBind | Meta
- Iused
- Speechmatics Introduces Ursa: A Speech-To-Text System That Delivers Unprecedented Performance Across A Diverse Range of Voices | Tanushree Shenwai - MarkTechPost
If you want to control your PC or dictate text directly into an application, Windows 11 Voice Access and Flow Pro are your two primary options for voice dictation. Voice Access is perfect for navigating your computer hands-free. You can say "Open Edge" or "Click File," and the operating system responds immediately. Wispr Flow Pro is better for actual content creation because it cleans up your grammar and removes filler words before the text hits your document.
ChatGPT, Gemini, and Claude serve as standalone advisors rather than system controllers. While ChatGPT offers a slick Windows desktop app that lets you keep a voice conversation running while you work in other windows, it will not press buttons or type text into your local software. You have to copy and paste their text outputs manually.
Contents
Windows 11 Built-in Voice Tools
Windows 11 actually offers two distinct built-in tools for voice: Voice Typing for simple text dictation, and Voice Access for complete hands-free control of your computer. Both are free and fully integrated, but they handle audio and execute tasks very differently.
Voice Typing (Win + H)
This is the quick, everyday dictation tool designed specifically for writing.
- How it Works: You press
Windows Key + H, click into any text box (like an email, Word document, or browser search), and start talking. The tool types whatever you say directly into the active field. - Cloud Processing: Voice Typing relies on Microsoft's Azure servers to process speech. This means you must have an active internet connection to use it, and your audio data is sent to the cloud for transcription.
- Capabilities: It automatically recognizes pauses for punctuation, or you can say commands out loud (e.g., "comma," "new line"). However, it only types text; it cannot click buttons or open applications.
Voice Access
Voice Access is a comprehensive accessibility system built to let you operate your entire PC without touching a mouse or keyboard.
- Total PC Control: Where Voice Typing only writes, Voice Access navigates. You can command your operating system with phrases like "Open Edge," "Click File," "Scroll down," or "Press Enter".
- Offline Processing: Once you download the initial language model during setup, Voice Access runs completely locally on your device. It works offline and never transmits your audio to the cloud, making it excellent for privacy.
- Integrated Dictation: It includes its own robust dictation system. You can use voice commands to navigate to a text field, and then seamlessly start dictating text, allowing for a fully hands-free workflow.
Which Should You Use?
If you just need to quickly draft a document or respond to a chat, press Win + H and use Voice Typing. If you want to lean back from your desk, navigate through applications, or if you need to work entirely offline, turn on Voice Access.
Wispr Flow
Wispr Flow is an AI-powered dictation tool that goes beyond basic speech-to-text. Instead of typing exactly what you say, it acts as a real-time editor. If you stumble, use filler words, or change your mind mid-sentence, Flow cleans up the output before the text hits your screen. It works universally across Windows, macOS, iOS, and Android, typing wherever you place your cursor.
Core Features
- Smart Auto-Editing: Flow understands backtracking. If you say, "Let's meet Tuesday, wait, actually make it Thursday," it simply types "Let's meet Thursday." It also automatically catches punctuation from your natural pauses.
- Personal Dictionary: You can teach it your specific terminology. It learns difficult names, brand terms, and acronyms, so you don't have to manually fix the same typos every day.
- Voice Snippets: Think of these as spoken text expanders. You can say a trigger phrase like "insert meeting link," and Flow pastes your full scheduling URL and standard greeting.
- Context-Aware Tone: It looks at the active window to adjust how it formats your words. It writes casually when you're in a Slack window, but switches to a formal structure when you open an email client.
- Whisper Mode: The engine handles very quiet speech, so you can dictate in an office or coffee shop without disturbing the people around you.
Developer and Coding Support
Flow is heavily used by developers because it handles technical jargon seamlessly.
- Syntax Awareness: It knows the difference between conversational English and code. It automatically formats variables in camelCase or snake_case and preserves correct spacing for command-line instructions.
- Smart IDE Tagging: If you use AI editors like Cursor or Windsurf, Flow recognizes when you say a filename out loud and automatically tags that file in your prompt workspace.
Pricing and Privacy
Because Flow uses advanced AI models to parse your speech, it requires an internet connection and processes your audio in the cloud.
- Free Plan: Allows you to dictate up to 2,000 words per week.
- Pro Plan: Costs $15 per month (or $144 annually) for unlimited dictation and access to the most advanced AI models.
- Security: The platform is SOC 2 Type II certified and offers HIPAA compliance controls to ensure data is handled securely.
ChatGPT Voice (GPT-Live)
ChatGPT Voice is built for natural, human-like conversation. In July 2026, OpenAI upgraded the default engine to GPT-Live, making the system full-duplex. This means the AI listens and speaks at the same time. You do not have to wait for it to finish a paragraph; you can interrupt it mid-sentence, laugh, or change the subject, and it instantly adapts its response and tone.
Windows 11 Desktop Integration
While tools like Voice Access control your computer, ChatGPT Voice operates as a powerful consultant that sits on top of your workflow.
- The Desktop App: You can trigger the voice interface using a keyboard shortcut (like
Alt + Space) to start talking without leaving your current window. - Screen Context: The Windows desktop app includes a screen vision feature. If you are looking at a complex spreadsheet or a block of code, you can ask ChatGPT to "take a look at this." It captures the active window so it understands exactly what you are talking about.
- App Isolation: ChatGPT Voice stays inside its own application ecosystem. It will format code or draft emails for you, but you must manually copy and paste that text into your external programs.
Core Capabilities
- Agent Orchestration: If you use ChatGPT Work or Codex, you can use your voice to direct background agents. You can tell the desktop app to "start a task to run the tests," and it will coordinate the work while you keep talking.
- Live Web Search: The GPT-Live update integrated web searching directly into voice sessions. You can ask for current flight prices or news updates, and it will pull live data without forcing you back to the text interface.
- Real-Time Translation: The system acts as a live interpreter for over 50 languages, maintaining the speaker's pace and natural inflection.
Automatic Speech Recognition (ASR)
Whisper
YouTube search... ...Google search
Whisper is an Automatic Speech Recognition Service (ASR) by OpenAI trained on 680,000 hours of multilingual and multitask supervised data collected from the web. 'We’ve trained and are open-sourcing a neural net called Whisper that approaches human level robustness and accuracy on English speech recognition.'
Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification. The Whisper v2-large model is currently available through our API with the whisper-1 model name. Currently, there is no difference between the open source version of Whisper and the version available through our API. However, through our API, we offer an optimized inference process which makes running Whisper through our API much faster than doing it through other means. For more technical details on Whisper, you can read the paper. - OpenAI