OpenAI introduced GPT‑Realtime‑2 and two new voice variants available only through the API
OpenAI Expands Voice API Capabilities
The company announced the launch of new features for its API that will allow developers to create more “lively” voice applications: they can talk with users, transcribe speech into text, and translate conversations on the fly.
New Realtime Models
Through the Realtime interface, three models are now available:
Model Purpose Key Features GPT‑Realtime‑2 Real‑time Voice Interaction Analyzes requests, calls tools, handles corrections, and smoothly continues dialogue. Built on GPT‑5 logic, enabling more complex tasks compared to GPT‑Realtime‑1.5. GPT‑Realtime‑Translate Real‑time Translation Supports 70+ input languages and 13 output languages. Preserves meaning even with context changes, regional accents, or specialized terminology. GPT‑Realtime‑Whisper Speech Transcription Low‑latency streaming model that turns audio into text almost instantly.
> “Our new models translate audio in real time from simple dialogue into voice interfaces that can actually work: listen, reason, translate, transcribe, and take action as the conversation unfolds,” the company said.
Pricing
Model Price GPT‑Realtime‑2 $32 per 1 million input audio tokens, $0.40 per 1 million cached tokens, and $64 per 1 million output audio tokens GPT‑Realtime‑Translate $0.034 per minute GPT‑Realtime‑Whisper $0.017 per minute
How to Try It
Developers can test the new models on OpenAI’s online Playground platform and immediately see how they work in real time.
Comments (0)
Share your thoughts — please be polite and stay on topic.
Log in to comment