Skip to content

feat: add Codex-style voice dictation to web and mobile - #6625

Open
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux
Open

feat: add Codex-style voice dictation to web and mobile#6625
KachurPro wants to merge 20 commits into
pingdotgg:mainfrom
KachurPro:codex/voice-dictation-codex-ux

Conversation

@KachurPro

@KachurPro KachurPro commented Aug 14, 2026

Copy link
Copy Markdown

Problem

The existing voice dictation PR (#5213) only covered web/desktop and had become stale. Native mobile dictation was still missing, and the first mobile version was limited to a hard-coded OpenAI endpoint and model.

Solution

  • rebase the voice transcription work onto the current main
  • match the Codex-style actions: cancel/discard, stop/transcribe/insert, and transcribe/insert/send
  • always append the normalized transcript to the end of the current draft
  • add native iOS/Android recording with expo-audio
  • let mobile users choose OpenAI or Groq
  • store a separate API key and model for each provider in expo-secure-store, with migration of the legacy OpenAI key
  • default OpenAI to gpt-4o-transcribe and Groq to whisper-large-v3-turbo
  • offer curated model choices while also accepting a custom model ID
  • send mobile recordings directly to the official provider endpoint
  • preserve the existing OpenAI/Groq server proxy and model picker on web/desktop
  • share transcript append and terminal-action rules between web and mobile
  • fully reset recorder, stream, timer, and audio-mode state after every terminal path

This supersedes #5213 and carries its web/desktop behavior forward.

Interaction

  • X: discard the recording
  • stop: transcribe and append to the draft
  • send: transcribe, append, then use the normal composer send path
  • existing draft text is preserved; speech is appended at the end

Verification

  • all 716 mobile tests passed
  • 4 focused web transcription tests passed
  • mobile and web typechecks passed
  • full repository lint and formatting checks passed
  • mobile native static analysis passed
  • Expo production config resolves expo-audio, the microphone permission, and Secure Store

The current development machine does not have Xcode or an iOS Simulator installed, so a local simulator screenshot is not included. The native flow still needs the EAS preview/device pass produced by CI.

Built with GPT-5.6 Sol in Codex Desktop.


Note

Medium Risk
New authenticated transcription proxy accepts client-supplied API keys in custom headers and forwards audio to third parties; composer send/stop paths gain async recording state, and mobile sends audio and keys directly to providers.

Overview
Adds Codex-style voice dictation across web/desktop chat, native mobile thread composer, and server-backed transcription.

Web/desktop: New Voice dictation settings (OpenAI or Groq, client API key or server OPENAI_API_KEY / GROQ_API_KEY, model picker). The composer gets a mic, a recording UI (waveform, cancel / stop / transcribe-and-send), and send behavior that waits on transcription when recording. Audio goes through authenticated /api/transcription and /api/transcription/models on the connected T3 server (25 MB cap). Shared appendVoiceTranscript and terminal-action rules live in @t3tools/shared. Desktop gains macOS microphone usage text.

Mobile: expo-audio recording, per-provider keys/models in Secure Store (with legacy OpenAI key migration), a Voice Dictation settings screen, and direct uploads to OpenAI/Groq (not the server proxy). Composer integration mirrors web (thread-scoped append, cancel on navigation).

Contracts: Client settings add voiceTranscription* fields; reset-to-defaults includes voice dictation when dirty.

Reviewed by Cursor Bugbot for commit ff3706c. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add Codex-style voice dictation to web and mobile chat composers

  • Adds microphone-based voice dictation to the web (ChatComposer) and mobile (ThreadComposer) chat composers, replacing the toolbar with a live waveform panel during recording/transcribing and appending the transcript to the draft on stop.
  • Supports OpenAI and Groq as transcription providers; users configure their provider, API key, and model in General settings (web) or a new Voice Dictation settings screen (mobile).
  • Server-side transcription proxy routes (GET/POST /api/transcription, GET /api/transcription/models) forward audio to the selected provider with a 25 MiB size limit and 2-minute timeout, supporting both environment-level and user-supplied API keys.
  • Shared utilities in @t3tools/shared handle transcript appending (whitespace-aware) and terminal action resolution (abort > send > insert).
  • macOS desktop build now includes NSMicrophoneUsageDescription in Info.plist; the mobile Expo app adds the expo-audio plugin with microphone permission.
  • Risk: voice transcription defaults to disabled (voiceTranscriptionEnabled=false); users must configure a model to activate the mic button. Existing client settings are migrated from a legacy OpenAI key format on first load.

Macroscope summarized ff3706c.

@coderabbitai

coderabbitai Bot commented Aug 14, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 35fa00aa-c405-43c0-b5aa-e7bb99d0e03d

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:XXL 1,000+ changed lines (additions + deletions). labels Aug 14, 2026
Comment thread apps/web/src/components/chat/ChatComposer.tsx
Comment thread apps/web/src/components/chat/ChatComposer.tsx
Comment thread apps/web/src/components/settings/SettingsPanels.tsx
@macroscopeapp

macroscopeapp Bot commented Aug 14, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a complete new voice dictation feature across web, mobile, and server, including new external API integrations (OpenAI/Groq), new server endpoints, audio recording infrastructure, and API key handling. New features of this scope and complexity require human review.

You can customize Macroscope's approvability policy. Learn more.

apiKey: voiceTranscriptionSettings.apiKey,
onTranscriptInsert: appendVoiceTranscriptToDraft,
onTranscriptSend: (transcript) => {
if (appendVoiceTranscriptToDraft(transcript)) void sendCurrentDraft();

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟠 High threads/ThreadComposer.tsx:585

The voice “send” action sends the previous draftMessage snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). appendVoiceTranscriptToDraft only schedules onChangeDraftMessage, while sendCurrentDraft() immediately calls onSendMessage() without the new text; pass the updated draft through the send path or wait for the state update before sending.

🤖 Copy this AI Prompt to have your agent fix this:
In file @apps/mobile/src/features/threads/ThreadComposer.tsx around line 585:

The voice “send” action sends the previous `draftMessage` snapshot instead of the transcript, so the transcript remains in the composer (or nothing is sent). `appendVoiceTranscriptToDraft` only schedules `onChangeDraftMessage`, while `sendCurrentDraft()` immediately calls `onSendMessage()` without the new text; pass the updated draft through the send path or wait for the state update before sending.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified as a false positive. onChangeDraftMessage does not schedule React state: useThreadComposerState calls setComposerDraftText, which synchronously calls appAtomRegistry.set. onSendMessage then calls getComposerDraftSnapshot from that same registry, so the just-appended transcript is present before enqueue. The existing mobile composer-drafts test also covers synchronous latest-draft reads for send.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sorry, I'm unable to act on this request because you do not have permissions within this repository.

@KachurPro
KachurPro force-pushed the codex/voice-dictation-codex-ux branch from 196ba3a to 5d1e3e3 Compare August 14, 2026 20:13
Comment thread apps/web/src/hooks/useVoiceTranscription.ts Outdated
Comment thread apps/web/src/components/settings/SettingsPanels.tsx
Comment thread apps/web/src/hooks/useVoiceTranscription.ts
Comment thread apps/web/src/hooks/useVoiceTranscription.ts
Comment thread apps/web/src/components/chat/VoiceTranscriptionPanel.tsx
Comment thread apps/web/src/components/chat/ChatComposer.tsx Outdated
@brvale97

Copy link
Copy Markdown

Really looking forward to this — mobile dictation is exactly what I've been missing on iPhone.

One request: the web/desktop implementation from #5213 supports both OpenAI and Groq with model selection, but the mobile path hardcodes the OpenAI endpoint and gpt-4o-mini-transcribe in mobileVoiceTranscription.ts. Groq's transcription API is OpenAI-compatible (https://api.groq.com/openai/v1/audio/transcriptions, e.g. whisper-large-v3-turbo), so carrying the provider choice over to mobile would mostly mean making the base URL + model configurable alongside the stored key. Would you consider that?

@KachurPro
KachurPro force-pushed the codex/voice-dictation-codex-ux branch from 59ba2fc to 495d058 Compare August 14, 2026 21:37
@KachurPro

Copy link
Copy Markdown
Author

@brvale97 Implemented in 495d058. Mobile now supports both OpenAI and Groq with fixed official transcription endpoints, separate securely stored keys and model choices per provider, a model picker, and a custom model ID field so future compatible models do not require an app update. OpenAI now defaults to gpt-4o-transcribe; Groq defaults to whisper-large-v3-turbo. Existing saved OpenAI keys are migrated automatically. All 716 mobile tests and the mobile typecheck pass. Thanks for catching this.

Comment thread apps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review of the web changes (apps/web/src/**). Three findings, all in the new dictation UI: a ticking live region that conflicts with the documented pattern in Sidebar.tsx, and two call sites that rebuild the existing ghost-muted button variant with a class override.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/ChatComposer.tsx Outdated
Comment thread apps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment thread apps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Comment thread apps/web/src/components/chat/ChatComposer.tsx

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the web-scope changes (ChatComposer.tsx, VoiceTranscriptionPanel.tsx, SettingsPanels.tsx, settingsSearch.ts, hooks/useVoiceTranscription.ts, lib/voiceTranscription.ts). The earlier findings on the panel (muted-ghost variant, live-region timer) are resolved. Two remaining consistency issues in the composer.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/ChatComposer.tsx
Comment thread apps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

UI consistency review: one finding on the web composer dictation panel. Previous findings (muted-ghost variant usage, live-region timer, responsive error inset, mic pointer-focus) are resolved in this head.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/ChatComposer.tsx
Comment thread apps/web/src/components/chat/ChatComposer.tsx Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding on the dictation panel's pointer-focus contract in the composer footer. Prior comments (mic variant/pointer focus, error-row inset, ghost-muted, aria-hidden timer) are addressed in this revision.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/VoiceTranscriptionPanel.tsx

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the new Effect code (apps/server/src/transcription.ts, the new route layers in apps/server/src/http.ts, and the contracts/shared additions) against the service conventions. Error classes use Schema.TaggedErrorClass with structural attributes and preserved cause, failures are recovered with Effect.catchTags, HttpClient is acquired from the environment, and Effect modules are imported from their subpaths — all consistent with the conventions. One finding on the web client below.

Posted via Macroscope — Effect Service Conventions

Comment thread apps/web/src/lib/voiceTranscription.ts Outdated
@KachurPro
KachurPro force-pushed the codex/voice-dictation-codex-ux branch from 9f92517 to 4732d3a Compare August 14, 2026 22:03
Comment thread apps/web/src/components/settings/SettingsPanels.tsx Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 2 potential issues.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 4732d3a. Configure here.

Comment thread apps/mobile/src/features/threads/ThreadComposer.tsx Outdated
Comment thread apps/web/src/components/settings/SettingsPanels.tsx Outdated

@macroscopeapp macroscopeapp Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One finding: the dictation send action bypasses the composer's themeable send tokens. Everything else (pointer-focus contract, ghost-muted variants, aria-hidden timer, responsive error inset) looks consistent with the composer's existing patterns.

Posted via Macroscope — UI Consistency

Comment thread apps/web/src/components/chat/VoiceTranscriptionPanel.tsx Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL 1,000+ changed lines (additions + deletions). vouch:unvouched PR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants