TL;DR: In a Node.js and Express catalogue importer, dedupe supplier images by using normalized metadata as a quick filter and a SHA-256 hash of the original bytes as the final decision. Store one immutable source object per digest, cache OCR output by that digest plus an extractor version, and record why each upload was accepted or reused. This keeps repeated storage and processing out of the import without pretending that filenames or dimensions prove identity.
The short mental model is a funnel. Before: every row creates an image object and an OCR job. After: cheap metadata narrows the search, a byte hash establishes identity, and only a new digest crosses the processing boundary. Fast first, certain second.
How should Node.js dedupe supplier images with metadata and a hash?
Supplier feeds reuse names such as front.jpg. They also rename the same file, and two distinct product photos can share a MIME type, byte length, width, and height. Metadata is useful as an index hin
Discussion
Be the first to comment
Add your perspective to get the discussion started.