OpenAI has added voice mode to its ChatGPT desktop app, allowing users to control AI agents and perform computer tasks through spoken commands. The update, released Thursday, integrates the company’s new ChatGPT-Live voice models into the macOS and Windows versions of the app.
The feature builds on OpenAI’s earlier voice capabilities, which were limited to mobile devices and lacked action-based functionality. Now, desktop users can dictate multi-step commands, receive real-time responses, and even grant the app screen access to interact with websites and applications.
What ChatGPT Voice Can Do on Desktop
The desktop version of ChatGPT Voice supports both ChatGPT Work and Codex, OpenAI’s coding-focused AI model. Users can issue complex instructions—such as creating code threads, submitting pull requests, or debugging errors—without typing.
On macOS, the app can capture screen content, including alt-text, via a feature called Appshots. This allows ChatGPT to analyze and respond to on-screen information. OpenAI demonstrated the feature in a video where a developer used a single voice command to trigger a sequence of coding tasks.
How It Compares to Mobile and Competitors
Unlike the mobile version, which focused on conversational fluidity, the desktop update prioritizes task execution. Users can interrupt the AI mid-response and provide additional input as needed. OpenAI also confirmed that iOS users can access Codex’s voice mode remotely through the desktop app.
Anthropic, a rival AI firm, recently enhanced its voice mode for Claude, enabling similar app integrations with Gmail, Slack, and Canva. Both companies are expanding voice-driven AI tools, though OpenAI’s update emphasizes desktop productivity.
What’s Next for ChatGPT Voice
OpenAI has not announced further updates, but the addition of voice mode to desktop suggests a push toward more seamless AI integration. Users should watch for potential expansions, such as deeper app compatibility or additional platform support.