Update openai-webrtc to add a model switcher, have it default to the new gpt-realtime-2
8 previous messages
I’m going to inspect the
openai-webrtcapp to find where the current model is wired in, then I’ll add a switcher in the UI and set the default togpt-realtime-2. After that I’ll run a quick verification pass so we know the change is wired end to end.Explored 1 search, 1 list, ran 1 command
Explored 1 search, 1 list
- Listed files in
.- Searched for
openai-webrtc|package.json|vite|next|src|app
- Ran
pwd && rg -n "openai-webrtc|realtime|model" -S .I found the app entrypoint. Next I’m reading the current
openai-webrtcHTML so I can update the existing controls instead of bolting on a separate path.Explored 1 file, searched web 1 time
Explored 3 files
- Read
./openai-webrtc.html- Read
./openai-webrtc.html- Read
./openai-webrtc.htmlSearched web 1 time
- Searched web for
site:openai.com/docs realtime model gpt-realtime-2 OpenAI Realtime API modelI confirmed from the current OpenAI model docs that
gpt-realtime-2is a valid realtime model, so I’m wiring that in as the default and making the model choice user-selectable and persistent.Edited 1 file
- Wrote
./openai-webrtc.html(+17 -3)The UI and session wiring are updated. I’m doing a quick source check now to make sure the new select is wired correctly and there aren’t any accidental syntax or scoping issues.
Explored 2 files, 1 search
- Read
./openai-webrtc.html- Read
./openai-webrtc.html- Searched for
gpt-realtime-2|modelSelect|openai_realtime_model|createRealtimeSession\(inopenai-webrtc.html
Updated openai-webrtc to add a model picker and default it to gpt-realtime-2.
What changed:
- Added a
Modeldropdown to openai-webrtc.html with:gpt-realtime-2as the defaultgpt-realtimeas an alternate option
- Switched the realtime session setup to use the selected model instead of the old hardcoded preview model in openai-webrtc.html
- Persisted the chosen model in
localStorageso it survives page reloads in openai-webrtc.html
I also verified against the current OpenAI docs that gpt-realtime-2 is the right model to target.
If you want, I can also add:
- a small helper note under the selector explaining the difference between the two models
- a saved default voice alongside the model
- a more polished UI pass for this page
Error: Expect line: v=
4 previous messages
I’m checking the current realtime WebRTC pattern against the live docs, because that SDP parse error usually means the server got something other than a valid offer body. Once I spot the mismatch, I’ll patch it directly.
Searched web 2 times
Searched web 2 times
- Searched web for
site:platform.openai.com/docs realtime webrtc offer.sdp v= OpenAI sample- Searched web for
site:platform.openai.com/docs "OpenAI-Beta: realtime=v1" realtime callsI found the likely root cause: the page is still using the older
/v1/realtime?...URL shape, but the current WebRTC docs route the SDP exchange through/v1/realtime/calls. I’m switching the session handshake to that endpoint and adding a clear error if the server returns anything other than SDP, so we get a useful failure instead of a cryptic SDP parser message.Edited 1 file
- Wrote
./openai-webrtc.html(+6 -2)
I fixed the handshake path that was causing the bad SDP response.
What changed in openai-webrtc.html:
- Switched the WebRTC POST target from the older
/v1/realtime?...URL to/v1/realtime/calls - Added a proper non-OK response check so the page now surfaces the real OpenAI error instead of failing later with
Expect line: v=
That Expect line: v= message was happening because the browser was trying to treat a non-SDP error response as the SDP answer.
One note: the page is still a browser-only direct API-key flow. If you want the selected voice to be explicitly configured under the newer realtime flow, we may need to move to the newer session-config path the docs recommend, which usually means an ephemeral token or small backend helper. If you want, I can wire that up next.
does realtime-2 support a different set of voices from the original realtime ?
2 previous messages
I’m checking the current OpenAI docs so I can answer this precisely, since voice availability can vary by model generation.
Searched web 1 time
Searched web 1 time
- Searched web for
site:platform.openai.com/docs gpt-realtime-2 voices supported voices realtime-2 openai
No. Based on the current OpenAI docs, gpt-realtime-2 uses the same built-in Realtime voice set as the other realtime models.
Current Realtime voice options are:
alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, and cedar.1
A couple of useful details:
- The voice is set on session creation, or on
response.create, not per-model.1 - Once a session has emitted audio, the voice can’t be changed for that session.1
- OpenAI recommends
marinorcedarfor best quality.1
So the model switch doesn’t imply a different voice palette. It’s the same voice list, just a newer model behind it.
If you want, I can update the UI to show only the recommended voices for gpt-realtime-2, or add a note that marin and cedar are the preferred choices.
I do not want a backend helper
That’s fine. We can keep this browser-only and not add any backend helper.
So the path is:
- keep the direct client-side API key flow
- use the selected realtime model in the browser
- keep the existing WebRTC handshake in the page
If you want, I can now make the browser-only setup more robust by:
- adding a model/voice compatibility note in the UI
- improving the handshake error handling for direct browser use
- removing any leftover code or wording that implies a backend is required
Error: OpenAI realtime handshake failed: 400 { "error": { "message": "You must provide a model parameter, for example wss://api.openai.com/v1/realtime?model=gpt-realtime-1.5", "type": "invalid_request_error", "code": "missing_model", "param": "" } }
2 previous messages
I found the immediate issue: the browser handshake endpoint still needs the model in the URL, even in the no-backend flow. I’m restoring that parameter while keeping the model picker wired through the page.
Edited 1 file
- Wrote
./openai-webrtc.html(+1 -1)
The browser-only flow needed the model in the request URL after all, so I put that back in openai-webrtc.html.
Specifically:
- the handshake now posts to
https://api.openai.com/v1/realtime?model=... - the selected model from the new dropdown is still used
- no backend helper is required
That missing_model error should go away now.
If you still hit a handshake failure after this, the next thing I’d check is whether the newer model expects any session config changes alongside the model switch.
add gpt-realtime-1.5 to the picker as well, still default to 2
2 previous messages
I’m adding the extra model option and keeping
gpt-realtime-2as the persisted default, so you can switch between the newer snapshots without changing the rest of the flow.Edited 1 file
- Wrote
./openai-webrtc.html(+1 -0)
Added gpt-realtime-1.5 to the model picker in openai-webrtc.html, with gpt-realtime-2 still remaining the default.
Error: OpenAI realtime handshake failed: 400 { "error": { "message": "The Realtime Beta API is no longer supported. Please use /v1/realtime for the GA API.", "type": "invalid_request_error", "code": "beta_api_shape_disabled", "param": "" } }
6 previous messages
I’m reconciling the current handshake with the GA Realtime docs now. The error is telling us the old beta shape is still being used somewhere, so I’m checking the current request format before changing the page again.
Searched web 2 times
Searched web 2 times
- Searched web for
site:platform.openai.com/docs/api-reference/realtime GA API /v1/realtime browser WebRTC SDP body model session- Searched web for
site:platform.openai.com/docs/api-reference/realtime session object model voice body FormData sdp GA WebRTCThe current browser flow should be using the GA
/v1/realtime/callsendpoint with a multipart body that includes both the SDP and asessionconfig. I’m patching the page to send that shape directly from the browser so we stay backend-free.Edited 1 file
- Wrote
./openai-webrtc.html(+15 -3)I spotted one more mismatch while patching: multipart form uploads must not send the old
application/sdpcontent type header. I’m removing that so the browser can generate the proper boundary headers.Edited 1 file
- Wrote
./openai-webrtc.html(+1 -2)
I updated the browser-only handshake to use the GA WebRTC shape again:
POST https://api.openai.com/v1/realtime/calls- multipart form body with:
sdpsessionJSON containingtype,model, andaudio.output.voice
- removed the old
Content-Type: application/sdpheader so the browser can send the correct multipart boundary
That should address the beta_api_shape_disabled error while keeping everything frontend-only.
If you still get a handshake failure, the next thing I’d check is whether the selected model accepts the chosen voice in the session config, but this should at least get us back onto the GA request shape.
Add a mute button that can be clicked to toggle on and off the microphone muting
2 previous messages
I’m adding a mic mute toggle directly into the existing session controls and wiring it to the live audio track, so you can silence input without tearing the session down.
Edited 1 file, explored 1 search
Explored 1 search
- Searched for
startButton|audioStream|setupAudioVisualization|stopSession\(|createRealtimeSession\(inopenai-webrtc.html
- Wrote
./openai-webrtc.html(+28 -2)
Added a microphone mute toggle to openai-webrtc.html.
What it does:
- adds a
Mute Micbutton next toStart Session - toggles the live microphone track on/off by flipping
track.enabled - updates to
Unmute Micwhile muted - disables the button until a session is running
- resets cleanly when you stop the session
If you want, I can also make the mute state visually reflected in the audio indicator or add a keyboard shortcut for mute/unmute.
what pricing data does it use at the moment?
2 previous messages
I’m checking the pricing constants in the page so I can tell you exactly what it’s using right now, not what the docs say in general.
Explored 1 file, 1 search
- Searched for
AUDIO_INPUT_COST|AUDIO_OUTPUT_COST|CACHED_AUDIO_COST|TEXT_INPUT_COST|TEXT_OUTPUT_COSTinopenai-webrtc.html- Read
./openai-webrtc.html
It’s using hardcoded token pricing in the page, not fetched dynamically.
Current constants in openai-webrtc.html:
AUDIO_INPUT_COST = 0.00004AUDIO_OUTPUT_COST = 0.00008CACHED_AUDIO_COST = 0.0000025TEXT_INPUT_COST = 0.0000025TEXT_OUTPUT_COST = 0.00001
So the calculator currently treats:
- audio input as
0.00004per token - audio output as
0.00008per token - cached audio input as
0.0000025per token - text input as
0.0000025per token - text output as
0.00001per token
If you want, I can update those to match the latest model-specific pricing for whichever realtime model you want to treat as canonical.
add
gpt-realtime-minitoo, use this pricing infoUser context
Images:
~/Library/Application Support/CleanShot/media/media_hv7fJRT8Gm/CleanShot 2026-06-05 at 13.47.13@2x.png
6 previous messages
I’m pulling the current model pricing so I can update the calculator correctly for
gpt-realtime-miniinstead of guessing from the screenshot alone.Searched web 2 times
Searched web 2 times
- Searched web for
site:developers.openai.com/api/docs/models/gpt-realtime-mini pricing output cached input audio text- Searched web for
https://developers.openai.com/api/docs/models/gpt-realtime-1.5I’ve got the pricing values for
gpt-realtime-mini, so I’m wiring the calculator to switch rates by selected model and adding the mini option to the picker.Edited 1 file
- Wrote
./openai-webrtc.html(+40 -0)I caught one more thing: the calculator still needed to actually read the selected model’s rates. I’m threading the active model through the cost calculation now so the numbers match the picker.
Edited 1 file
- Wrote
./openai-webrtc.html(+11 -14)
Added gpt-realtime-mini to the picker and switched the cost calculator to model-specific pricing.
What changed in openai-webrtc.html:
- Added
gpt-realtime-minito the model selector - Kept
gpt-realtime-2as the default - Reworked pricing so it now uses rates based on the selected model instead of one flat set of constants
Pricing used:
gpt-realtime-2,gpt-realtime-1.5, andgpt-realtime:- audio input:
$32.00 / 1M - cached audio input:
$0.40 / 1M - text input:
$4.00 / 1M - cached text input:
$0.40 / 1M - audio output:
$64.00 / 1M - text output:
$16.00 / 1M
- audio input:
gpt-realtime-mini:- audio input:
$10.00 / 1M - cached audio input:
$0.30 / 1M - text input:
$0.60 / 1M - cached text input:
$0.06 / 1M - audio output:
$24.00 / 1M - text output:
$2.40 / 1M
- audio input:
Sources:
One important note: I included the new cached text input rates in the calculator as well, since the current UI tracks cached tokens but the old code only applied a cached audio rate. If you want, I can also add a small on-page “pricing source” note so future model updates are easier to audit.
add a feature where it always prominently shows the last transcript value from this shape of JSON
{ "type": "response.done", "event_id": "event_DnVlmRlY2yuj8mH6rJw6X", "response": { "object": "realtime.response", "id": "resp_DnVlku2XHV7Z0COfVvlXK", "status": "completed", "status_details": null, "output": [ { "id": "item_DnVlkhkfGpuP815gsx8j3", "type": "message", "status": "completed", "role": "assistant", "content": [ { "type": "output_audio", "transcript": "Hey, I’m here to help! Let’s figure this out together. Can you tell me a bit more about what’s going on or what you need help with?" } ] } ] - above the Most recent interaction box
10 previous messages
I’m adding a live “last transcript” display right above the recent-usage panel and wiring it to
response.doneevents so the latest assistant transcript stays visible at a glance.Explored 2 files
- Read
./openai-webrtc.html- Read
./openai-webrtc.htmlI’ve got the insertion point. I’m adding a dedicated transcript card above the usage stats and teaching the event handler to pull
response.output[].content[].transcriptfromresponse.done.Edited 1 file
- Wrote
./openai-webrtc.html(+91 -24)The first patch missed the exact spacing in the event handler, so I’m re-reading the small region around the transcript logic and applying a tighter patch.
Edited 1 file, explored 1 file
Explored 3 files
- Read
./openai-webrtc.html- Read
./openai-webrtc.html- Read
./openai-webrtc.html
- Wrote
./openai-webrtc.html(+102 -33)Explored 1 file
- Read
./openai-webrtc.html
Wrote
./openai-webrtc.html(+27 -0)Wrote
./openai-webrtc.html(+4 -0)Wrote
./openai-webrtc.html(+1 -0)Wrote
./openai-webrtc.html(+23 -21)Wrote
./openai-webrtc.html(+34 -0)The transcript panel is in place now. I’m doing one quick pass over the edited regions to make sure the event handler and DOM hooks line up cleanly.
Explored 2 files, 1 search
- Searched for
lastTranscript|updateLastTranscript|extractLastTranscript|transcript-box|response.doneinopenai-webrtc.html- Read
./openai-webrtc.html- Read
./openai-webrtc.htmlI noticed the transcript DOM hook is only declared once, which is good. I’m checking the lower section for any accidental duplicate logic before I wrap up.
Explored 2 files
- Read
./openai-webrtc.html- Read
./openai-webrtc.html
Added the prominent last-transcript panel above “Most recent interaction” in openai-webrtc.html.
What it does:
- Shows the latest assistant transcript in a high-visibility box
- Updates on every
response.doneevent - Reads the transcript from
response.output[].content[].transcript - Keeps the previous “waiting” placeholder until the first transcript arrives
I left the rest of the session stats behavior unchanged.
If you want, I can also make it:
- auto-scroll into view on each new transcript
- show the speaker role or timestamp alongside the transcript
- retain a short transcript history instead of only the latest one
did you use a prompt for that.? I didn't want it to use a transcrie prompt
No. I did not add any prompt for that feature.
What I changed:
- I only read the
transcriptfield already present inresponse.doneevents - I did not send any extra
instructions,prompt, or transcription prompt to the model - I did not change the session prompt at all
So this is just UI extraction of the transcript that the model already emits, not a new transcription prompt.
If you want, I can also verify there isn’t any existing prompt-related field in the session config and remove it if needed.