AddThisFeature

AI Duplicate Record Matching

Surface records that probably describe the same thing, with the evidence for each match.

involved AI Safety & Trust

What it adds

A duplicate detection pass combining deterministic keys with semantic similarity, producing reviewable candidate pairs.

What your agent is told to do

5
  1. 1

    Run deterministic matching first on the identifiers the domain already trusts — email, tax number, external system ID, normalized phone — and settle those cases before any model is involved. Semantic similarity is for what deterministic rules miss, not a replacement for them.

  2. 2

    Use similarity to generate candidates, then score each pair on a field-by-field basis and store which fields agreed, which conflicted, and by how much.

  3. 3

    Present matches as a review queue showing the two records side by side with the contributing fields highlighted, and make dismissal permanent so the same pair does not resurface every run.

  4. 4

    Scope every comparison to a single tenant and to records the reviewer may see. A duplicate finder that reaches across a workspace boundary is a data leak with a helpful interface.

  5. 5

    Make merges reversible for a defined window: keep the pre-merge state, record which record absorbed which, and preserve references from both so links elsewhere in the app do not break.

Edge cases it handles

8
  • Neither signal is sufficient alone. Deterministic identifiers catch the obvious cases, similarity catches the messy ones, and a pair should surface only when the combined evidence holds up field by field.
  • Common names, shared support addresses, and generic descriptions produce large false-positive clusters. Require corroboration on a second independent field before proposing a merge on a name alone.
  • Every candidate must show why it was proposed. A confidence percentage with no visible evidence gives a reviewer nothing to reason about and trains them to approve everything.
  • Comparison must never cross tenant, workspace, or permission boundaries, and the review queue must exclude pairs where the reviewer cannot see both sides.
  • Merges must be confirmed by a person or governed by rules narrow enough to be safe, and must be reversible. An automatic merge on a wrong pair destroys two records and the history that explains them.
  • Field-level conflicts need an explicit resolution step. When two records disagree on an address or a phone number, the reviewer chooses rather than the newer record silently winning.
  • Duplicate detection over a large table is expensive. Block candidates on cheap keys before scoring, cap the work per run, and process through the app's existing background-job system rather than during a request.
  • Dismissals and merges must be recorded in an audit trail with who acted and when, since a merge is one of the least recoverable operations in the app.

Definition of done

9
  • Deterministic identifier matches are resolved before semantic scoring runs.
  • Every candidate pair records the fields that agreed and the fields that conflicted.
  • The review interface shows both records with the matching evidence highlighted.
  • Comparisons and the review queue respect tenant and permission boundaries.
  • Merges require confirmation, resolve field conflicts explicitly, and are reversible within a defined window.
  • Dismissed pairs do not reappear in later runs.
  • Detection runs in the background with a bounded cost per run.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.