OpenAI has released 'GPT-Live-1', an API model that enables developers and companies to build natural voice agents. While existing voice services often experienced latency and disrupted conversation flows due to multiple stages—such as speech recognition, reasoning models, and speech synthesis—the newly unveiled model improves these limitations by processing real-time bidirectional conversations through a single architecture.

Architecture Simplification and Performance Improvement

GPT-Live-1 processes listening and speaking simultaneously, naturally handling user interruptions, background noise, and silence. It can be linked with backend text models to handle complex reasoning or tool-calling tasks and supports telephony networks, making it applicable for actual phone services such as customer support or restaurant reservations. In early evaluations, the learning tool Speak reduced unnecessary conversation breaks that occur during thinking by approximately 80% compared to previous methods.

Pricing and Terms of Use

GPT-Live-1 is available via API at a rate of $0.05 per minute for the frontend voice layer, allowing developers to combine it with backend models and agents that suit their products. To use customized voices, users must verify eligibility and application procedures through the sales team.