Canonical item
The shared identity of a title: platform, region, edition and other attributes that define what the product actually is.
01 · Flagship case study · Field tested
A production-minded pipeline for turning Japanese second-hand store photos into traceable buying decisions. The difficult part is not recognizing a game. It is deciding whether this exact physical copy is worth buying without silently mixing incompatible evidence.
The problem
I started with a practical question while looking through retro-game stores in Japan: which copies are actually worth buying for resale?
A title can look profitable until the details change. The disc may be missing. The copy may be a budget reissue. The market result may be for a different platform. A nearby price label may belong to another game. A high retail asking price may be mistaken for a sold value.
The project grew from photo reading into a system for preserving evidence, resolving identity, filtering incompatible market data and producing decisions that can explain themselves.
Architecture
The pipeline deliberately keeps recognition, physical-copy state, market evidence and economics separate. Each stage can refuse to continue when the evidence is not good enough.
No downstream stage is allowed to manufacture certainty that the evidence layer does not have.
Data model
The shared identity of a title: platform, region, edition and other attributes that define what the product actually is.
One exact copy at one store, on one visit, at one visible price, with its own condition and completeness state.
One sold result, active ask, retail ask, buyback quote or guide observation. Evidence types stay semantically distinct.
The source image or document supporting a claim, linked through hashes and image regions instead of being discarded after OCR.
Hard rules
Reject the comp. A Saturn market result cannot price a PlayStation copy.
BLOCKDo not silently borrow standard-edition pricing for a budget or collector release.
BLOCKDisc-only and complete-copy evidence remain separate unless an explicit rule allows otherwise.
BLOCKStore NULL. Research time is not substituted for an unreadable source value.
UNKNOWNReturn RESEARCH rather than forcing a confident-looking purchase decision.
DEFERFailure → redesign
Reading a dense shelf as a flat block of text lost the spatial relationship between titles and nearby labels.
Keep the original image immutable, derive crops and grid regions, and preserve explicit spatial lineage back to the source.
A complete-copy comp can turn a damaged or incomplete physical copy into a fake bargain.
Physical-copy state became a separate entity, with compatibility gates before any market evidence reaches economics.
A sourcing tool that cannot safely re-run becomes difficult to trust during iterative research.
Hash-based evidence identity, idempotent writes, integrity checks and checkpoint validation became part of the workflow.
Field validation
A recent validation batch illustrates how the system behaves when the evidence is imperfect.
Engineering surface
Typed domain models, modular decision logic, CLI workflows, hashing, validation and testable business rules.
Normalized entities, foreign keys, append-only evidence, idempotent writes and integrity checks.
Image preservation, OCR candidate extraction, crop/grid derivation and evidence-region tracking.
Deduplication, canonical resolution, provenance, evidence typing, compatibility and freshness handling.
Deterministic fees, profit, ROI, confidence gates, maximum-buy thresholds and explicit failure states.
Regression tests, checkpoint validation, repeat-import checks and CI-safe publication boundaries.
Public / private boundary
The public repository is deliberately runnable. A reviewer can inspect the architecture, execute the synthetic demo, exercise the SQLite layer and run the tests. Real store photos, production market datasets, private adapters and operational thresholds remain private.
Typed Python models
Production database and sourcing history
Synthetic end-to-end demo
Real store photographs and market evidence
Compatibility tests + CI
Operational thresholds and source adapters
Repository
Architecture notes, data-model documentation, ADRs, regression tests and the synthetic demo are all available in the repository.