AddThisFeature

Speaker Labels

Split a transcript by who was speaking and let people rename each speaker once.

involved Data & Content

What it adds

Speaker attribution on transcript segments, with editable names, merging and splitting of detected speakers, and labels that survive regeneration.

What your agent is told to do

5
  1. 1

    Store the speaker as an attribute of each transcript segment on the existing transcript record. Do not fork the transcript into a second speaker-aware copy that then drifts from the original.

  2. 2

    Give each detected speaker a neutral placeholder name and a consistent colour, and let a user rename any speaker from the transcript itself rather than from a separate settings screen.

  3. 3

    Make renaming a single action that applies to every segment attributed to that speaker, with an undo, and make it obvious how many segments a rename will touch.

  4. 4

    Provide merge and split for the cases detection gets wrong: merging two placeholder speakers into one person, and reassigning a run of segments to a different speaker.

  5. 5

    Do not guess a speaker's real identity from participant lists or account names and present it as fact. Suggest a match if you have one, and require a person to confirm it before it becomes the label.

Edge cases it handles

8
  • Detection routinely over-segments, splitting one person into three, and under-segments, folding two quiet speakers into one. Both need a repair path: merge several detected speakers into one, and split a speaker's segments apart by reassigning them.
  • Renaming a speaker must update every occurrence in one action, including segments not currently on screen, and must be undoable as a single step rather than segment by segment.
  • Crosstalk, where two people speak at once, will produce overlapping or interleaved segments. Attribute what is confident, mark the rest as uncertain, and never silently drop the overlapping audio's text.
  • Regenerating a transcript must not wipe the names a person has already entered. Key the labels to something durable and re-apply them to the new segmentation, telling the user where re-attachment was not possible.
  • When diarisation confidence is poor, present the transcript as a single unlabelled speaker. A transcript wrongly split across five imaginary people is worse than one with no labels at all.
  • A speaker name is typed once and then follows the transcript into every export, share link, and quoted excerpt. Cap its length so it does not break the transcript layout, and check it before it travels, because renaming it later does not recall the copies already sent.
  • Colours assigned to speakers must stay distinguishable in the app's light and dark themes and must not be the only way a speaker is identified.
  • Two people editing the same transcript must not overwrite each other's renames. Apply the app's existing concurrency handling for shared records.

Definition of done

8
  • Every transcript segment carries a speaker attribute stored on the existing transcript record.
  • Renaming a speaker updates all of that speaker's segments in one undoable action.
  • Detected speakers can be merged, and runs of segments can be reassigned to another speaker.
  • Regenerating a transcript preserves previously entered speaker names where the segments still match.
  • Low-confidence diarisation falls back to a single unlabelled speaker rather than inventing speakers.
  • A speaker name is length-capped, checked before it travels, and escaped in every export and share view.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.