AddThisFeature

Image Text Extractor

Turn a photo or scan of a document into text the user can copy and edit.

involved Media & Video

What it adds

Text extraction from uploaded images and scanned documents, returned as editable, copyable text alongside the original.

What your agent is told to do

5
  1. 1

    Design around the wait. Extraction is not instant, so accept the upload, show a clear processing state, and let the user leave the page and come back to a finished result rather than holding a spinner.

  2. 2

    Run extraction as a background job on the app's existing queue, with a bounded retry and a visible failed state. A request that blocks until a multi-page scan finishes will time out.

  3. 3

    Reuse the app's existing File Upload path and, where one exists, its Multimodal Image Analysis pipeline, rather than standing up a second image ingestion route with its own limits and storage.

  4. 4

    Present the extracted text next to the original image so the user can check it, and make it editable and copyable. Extraction is never perfect and a result nobody can correct is a result nobody trusts.

  5. 5

    Do not silently present low-confidence output as fact. Surface the confidence, flag the passages the extractor was unsure about, and prompt the user to review them.

Edge cases it handles

8
  • Phone photos arrive rotated, skewed, shadowed, and noisy. Correct orientation, straighten, and clean the image before reading it, or accuracy collapses on exactly the inputs users supply most.
  • Multi-column layouts, tables, and sidebars read as nonsense when scanned line by line across the page. Preserve the reading order of each column rather than interleaving them.
  • Return a confidence figure with the result and mark the low-confidence regions, so poor scans can be flagged for review instead of quietly entering the system as correct text.
  • A long multi-page file must be processed page by page in the background, with progress shown and partial results kept, rather than one request that runs until it times out.
  • Uploaded documents are often sensitive. Delete the source image on the retention schedule the customer was told about, and make that schedule visible in the interface.
  • Enforce page count, file size, and type limits on the server, and reject clearly over-limit files before any processing cost is incurred.
  • An image containing no readable text must return an explicit empty result with an explanation, not an error and not a blank editor.
  • Extracted text may contain personal data. Apply the same access rules to the result as to the source file, and remove both together when the record is deleted.

Definition of done

9
  • A user can upload an image or scan and receive editable, copyable text.
  • Extraction runs in the background with a visible processing, failed, and complete state.
  • Skewed, rotated, and noisy photos are corrected before reading.
  • Multi-column documents preserve their reading order.
  • Results carry a confidence indication and low-confidence passages are flagged.
  • Multi-page files process page by page without timing out and keep partial results.
  • Source images are deleted on the stated retention schedule and share the result's access rules.
  • The feature matches the existing design system.
  • No existing functionality is broken.

Related features

How it works

  1. 1

    Copy the link

    Grab the Markdown instruction URL for this feature.

  2. 2

    Give it to your AI

    Paste it into Claude Code, Cursor, v0, Lovable — whatever you build with.

  3. 3

    It inspects, then implements

    Your agent reads your existing app first, then adds the feature to fit it.

Works with your stack

These instructions are written to adapt. They tell the agent to detect your framework, match your existing design system, and reuse what you already have — rather than assuming a particular stack.

Need it tighter than that? Customize the feature and tell it exactly what you're running.