The problem

A price tag is not a valuation.

I started with a practical question while looking through retro-game stores in Japan: which copies are actually worth buying for resale?

A title can look profitable until the details change. The disc may be missing. The copy may be a budget reissue. The market result may be for a different platform. A nearby price label may belong to another game. A high retail asking price may be mistaken for a sold value.

The project grew from photo reading into a system for preserving evidence, resolving identity, filtering incompatible market data and producing decisions that can explain themselves.

Architecture

Evidence first. Decision last.

The pipeline deliberately keeps recognition, physical-copy state, market evidence and economics separate. Each stage can refuse to continue when the evidence is not good enough.

01 / INPUTStore photoimmutable source evidence
→
02 / EVIDENCEHash + regiondedupe and spatial lineage
→
03 / VISIONOCR + candidatestitles, labels, positions
→
04 / IDENTITYCanonical itemplatform, region, edition
→
05 / MARKETCompatible evidencesold, ask, buyback, guide
→
06 / MODELEconomicsfees, ROI, max buy
→
07 / OUTPUTDecisionBUY / MAYBE / RESEARCH / SKIP
RULE

No downstream stage is allowed to manufacture certainty that the evidence layer does not have.

Data model

The separation that makes the system trustworthy.

A

Canonical item

The shared identity of a title: platform, region, edition and other attributes that define what the product actually is.

B

Physical observation

One exact copy at one store, on one visit, at one visible price, with its own condition and completeness state.

C

Market observation

One sold result, active ask, retail ask, buyback quote or guide observation. Evidence types stay semantically distinct.

D

Evidence asset

The source image or document supporting a claim, linked through hashes and image regions instead of being discarded after OCR.

Hard rules

The system is designed to say “I don’t know.”

01Different platform?

Reject the comp. A Saturn market result cannot price a PlayStation copy.

BLOCK
02Different edition?

Do not silently borrow standard-edition pricing for a budget or collector release.

BLOCK
03Different completeness?

Disc-only and complete-copy evidence remain separate unless an explicit rule allows otherwise.

BLOCK
04Obscured price?

Store NULL. Research time is not substituted for an unreadable source value.

UNKNOWN
05Weak market evidence?

Return RESEARCH rather than forcing a confident-looking purchase decision.

DEFER

Failure → redesign

The useful mistakes changed the architecture.

EARLY FAILURE

Whole-shelf OCR attached the wrong price to the wrong game.

Reading a dense shelf as a flat block of text lost the spatial relationship between titles and nearby labels.

REDESIGN

Keep the original image immutable, derive crops and grid regions, and preserve explicit spatial lineage back to the source.

EARLY FAILURE

Headline market value looked useful until copy state disagreed.

A complete-copy comp can turn a damaged or incomplete physical copy into a fake bargain.

REDESIGN

Physical-copy state became a separate entity, with compatibility gates before any market evidence reaches economics.

EARLY FAILURE

Repeat imports risked duplicating evidence and changing history.

A sourcing tool that cannot safely re-run becomes difficult to trust during iterative research.

REDESIGN

Hash-based evidence identity, idempotent writes, integrity checks and checkpoint validation became part of the workflow.

Field validation

Built against messy store data, not a clean demo CSV.

A recent validation batch illustrates how the system behaves when the evidence is imperfect.

uploaded photos10
unique evidence assets9
distinct physical copies49
directly readable prices48
obscured prices guessed0
compatible market observations added55
resulting screen1 BUY · 7 MAYBE · 41 RESEARCH
FK CHECK0INTEGRITYOKREPEAT IMPORTNO-OP

Engineering surface

What the project demonstrates.

Python

Typed domain models, modular decision logic, CLI workflows, hashing, validation and testable business rules.

SQLite / data modeling

Normalized entities, foreign keys, append-only evidence, idempotent writes and integrity checks.

Computer vision workflow

Image preservation, OCR candidate extraction, crop/grid derivation and evidence-region tracking.

Data engineering

Deduplication, canonical resolution, provenance, evidence typing, compatibility and freshness handling.

Decision systems

Deterministic fees, profit, ROI, confidence gates, maximum-buy thresholds and explicit failure states.

Reliability

Regression tests, checkpoint validation, repeat-import checks and CI-safe publication boundaries.

Public / private boundary

Inspectable without publishing the operational sourcing system.

The public repository is deliberately runnable. A reviewer can inspect the architecture, execute the synthetic demo, exercise the SQLite layer and run the tests. Real store photos, production market datasets, private adapters and operational thresholds remain private.

PUBLIC

Typed Python models

PRIVATE

Production database and sourcing history

PUBLIC

Synthetic end-to-end demo

PRIVATE

Real store photographs and market evidence

PUBLIC

Compatibility tests + CI

PRIVATE

Operational thresholds and source adapters

Repository

The public implementation is built to be read, run and challenged.

Architecture notes, data-model documentation, ADRs, regression tests and the synthetic demo are all available in the repository.