Skip to content
Tech News
← Back to articles

GPT‑Live‑1 in the API

read original get Shure MV7+ USB Podcast Microphone → more articles
Why This Matters

OpenAI is bringing GPT-Live-1, its full-duplex voice model already in ChatGPT, to the API — letting developers build voice agents that listen and speak simultaneously instead of stitching together speech-to-text, LLM, and text-to-speech chains. That single-model approach targets the biggest weakness of current voice assistants: awkward latency and premature interruptions, with early user Speak reporting nearly 80% fewer interruptions than turn-based systems. It also opens the door to telephony-based agents for reservations, support, and other business workflows.

Key Takeaways
Worth a Look

Shure MV7+ USB Podcast Microphone — If you're building voice apps with a real-time speech model, clean input audio makes a huge difference, and the MV7+ is a dynamic USB/XLR mic designed to capture close-up speech while rejecting room noise. It plugs straight into a laptop for prototyping voice agents and doubles as a serious podcast and streaming mic.

See Shure MV7+ USB Podcast Microphone on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

We’re launching GPT‑Live‑1 in the API, giving developers a powerful, natural voice model for building voice-enabled apps and business workflows. First introduced in ChatGPT , GPT‑Live‑1 is capable of listening and speaking at the same time, and, as seen with Codex and ChatGPT Work ⁠(opens in a new window), can delegate deeper reasoning and actions to the models and tools it is paired with.

For the API release of GPT‑Live‑1, we’ve focused on new capabilities that let developers steer and customize voice experiences around their users, workflows, and goals. A core GPT‑Live‑1 strength, smooth interruption handling, is already delivering business impact: in early evaluations, Speak found that GPT‑Live‑1 gave learners more time to think before the language tutor responded, cutting interruptions by almost 80% versus previous turn-based systems.

Key strengths of GPT‑Live‑1 in the API:

Interruption handling: Improves interruption handling via a single model that reasons over incoming and outgoing audio together, avoiding the latency and brittle handoffs of chained STT–LLM–TTS architectures.

Reasoning & tool calling delegation: GPT‑Live‑1 can delegate reasoning and tool calls to a backend text model like GPT‑6 Astra or a third-party model.

Tone, pace, and style: Lets developers shape an agent’s tone, pace, and conversational style through the system prompt.

Silent context management & background noise: Better handles background noise and silence without interrupting the conversation or narrating every step out loud.

Long-session reliability: Improves context retention and conversational quality across extended interactions.

Telephony support: Enables deployment of full-duplex voice agents for phone calls, from restaurant reservations to customer support.

Try GPT-Live-1 Start a session and speak naturally. Interrupt, laugh, change your mind - try it at home or in a loud space like a coffee shop or city street. Start session See what it can do Talk over it—naturally. Ask for help, then interrupt mid-response to change the question or add detail.

... continue reading