AddThisFeature

Voice Dictation

Put a microphone in text fields so people can speak their input instead of typing it.

moderate Forms & Input

What it adds

A dictation control on the app's longer text inputs that transcribes speech into the field as the user talks.

What your agent is told to do

5
  1. 1

    Add the control to the fields where typing is genuinely tedious: long descriptions, notes, comments, and message bodies. A microphone on a postcode field is clutter.

  2. 2

    Show interim results in the field as they arrive, visually distinct from committed text, so the speaker can see it is working and correct course before they finish.

  3. 3

    Reuse the app's existing Voice Transcription for the server-side path so one recognition configuration and one set of language settings serve both dictation and recorded audio.

  4. 4

    Make the recording state unmistakable: a change of colour alone is not enough. Show an active indicator and a stop control that is always reachable.

  5. 5

    Do not clear or replace the field when dictation starts. Speaking is usually an addition to something already written, and wiping existing text is unrecoverable if undo is not wired up.

Edge cases it handles

8
  • Insert transcribed text at the current caret position and respect an active selection, rather than appending to the end or overwriting the whole field.
  • Where the browser has no speech input support, hide the control entirely and leave a fully working text field. A microphone button that errors on press is worse than no button.
  • Stop listening automatically after a stretch of silence, and tell the user why it stopped so they do not keep talking to a field that is no longer listening.
  • Handle spoken punctuation and line breaks consistently, and document the phrases that work. Unpredictable punctuation makes every dictated sentence need a manual pass.
  • Release the microphone when the field loses focus, the dialog closes, or the user navigates away. A stubbornly active recording indicator is alarming and looks like a bug even when it is one.
  • Undo must treat a dictated block as a single step, so one keystroke removes it rather than fifty.
  • Dictation into a field that also autosaves must not save every interim result; commit on a pause or on stop.
  • Background noise and multiple voices produce garbage. Give the user an obvious way to discard the last dictated block without hunting for where it began.

Definition of done

9
  • Long text inputs across the app offer a dictation control.
  • Transcribed text is inserted at the caret and existing content is preserved.
  • Interim results are visually distinguished from committed text.
  • Unsupported browsers show no control and a fully functional text field.
  • Listening stops after sustained silence and on focus loss, navigation, or dialog close, with the microphone released.
  • The recording state is indicated by more than colour and can always be stopped.
  • A dictated block can be undone in a single step.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.