It has taken four months since its first public introduction to finally bring OpenAI’s new human-like conversational voice interface for ChatGPT called “ChatGPT Advanced Voice Mode” to users outside the tiny testing group and a long waitlist.
All OpenAI subscribers to ChatGPT Plus and Team plans will have access to the new ChatGPT Advanced Voice Mode, though rollout of access is gradual over the next few days, according to OpenAI. Access starts with the U.S. In the subsequent week, the company will make ChatGPT Advanced Voice Mode available to its subscribers on Edu and Enterprise plans.
The company will also be adding the capacity to store “custom instructions” for the voice assistant and “memory” of the behaviors the user wants it to exhibit, similar to features rolled out earlier this year for the text version of ChatGPT.
And five new, differently styled voices are shipping today too: Arbor, Maple, Sol, Spruce, and Vale-in addition to the first four already available-Breeze, Juniper, Cove, and Ember-between which users could chat using ChatGPT’s older, less advanced voice mode.
It will now mean that the users of ChatGPT, people for Plus, and small business teams for Teams can interact with this chatbot by speaking to it instead of typing out a prompt. Users will know they’ve entered Advanced Voice Assistant via a pop-up when they access voice mode on the app.
We’ve used learnings since the alpha to improve accents in ChatGPT’s most in demand foreign languages, and overall conversational speed and fluidity,” the company said. “You’ll notice, too, a new look for Advanced Voice Mode, with an animated blue orb.”.
Originally, voice mode had four voices – Breeze, Juniper, Cove and Ember, but new update will bring five new voices by the name of Arbor, Maple, Sol, Spruce and Vale. OpenAI did not provide a voice sample of the new voices.
These features are only on the GPT-4o model and not on the preview model newly released, o1. ChatGPT users can also make voice mode personal using custom instructions and memory to engage according to their preference in any conversation they have .
AI voice chat race
Since the development of AI voices assistants, such as Siri by Apple and Alexa by Amazon, developers have been working towards making the experience of generative AI chats more human-like.
The voices ChatGPT has been prepared with pre its voice mode launch through the Read-Aloud function; however, Advanced Voice Mode is alleged to give users a more human-like conversation experience for a feature that other AI developers want to copy.
Former Google Deepminder Alan Cowen founded Hume AI. The company on Monday released the second iteration of its Empathic Voice Interface, an AI voice assistant that has learned to understand emotion based on the cadence of someone’s voice. The humanlike assistant can be accessed by developers via proprietary API.
In July, the French AI firm Kyutai open sourced its AI voice assistant, Moshi.
That aside, Google also added voices to its Gemini chatbot through Gemini Live, as it tries to catch up with OpenAI. Meta is also developing voices sounding like popular actors to add to its Meta AI platform, Reuters says.
OpenAI claims it’s making AI voices widely available to more users across its platforms, bringing the technology into the hands of so many more people than those other firms.
Arrives with delays and controversy
However, the idea of AI voices that speak in real-time and respond with the right emotional intonation has not been taken quite so seriously.
OpenAI’s effort to bring voices to ChatGPT has been controversial from the word go. Recently, during its event in May to announce GPT-4o and the voice mode, people noticed that one of the voices, Sky, sounded like the actress Scarlett Johanssen.
It didn’t help that OpenAI CEO Sam Altman posted the word “her” on social media, a reference to the movie where Johansson voiced an AI assistant. The controversy raised concerns about AI developers using voices of well-known individuals.
The company denied it referenced Johansson and claimed it was not trying to hire actors whose voices sound similar to others.
There, the company said the users are restricted to only the nine voices from OpenAI. The company also added that it evaluated its safety before releasing it.
“We tested the model’s voice capabilities with external red teamers, who collectively speak a total of 45 different languages, and represent 29 different geographies,” the company said in an announcement to reporters.
It, however pushed back the roll out of ChatGPT Advanced Voice Mode from its planned roll out date in late June to “late July or early August,” then further restricted that to a group of OpenAI-selected initial users such as University of Pennsylvania Wharton School of Business professor Ethan Mollick, citing continuing safety testing or “read teaming” the voice mode to avoid its use in potential fraud and wrong doing.
Clearly, the company feels it has done enough to roll out the mode more widely now — and it’s in keeping with OpenAI’s generally more careful approach of late, working hand-in-hand with the U.S. and U.K. governments and allowing them to preview new models such as its o1 series prior to launch.


