A reusable data pipeline for building retailer-level reference datasets for Open Food Facts Canada.
Reference DB transforms raw retailer product data into standardized, validated, classified, grouped, and enriched product datasets.
The pipeline is designed to support multiple retailers while keeping retailer-specific extraction logic isolated from the generic processing pipeline.
Retailer Source
↓
Retailer Adapter
↓
Ingestion
↓
Phase 1 — Validation
↓
Phase 2 — Identity
↓
Phase 3 — Classification + Taxonomy + Grouping
↓
Phase 4 — Variants
↓
Phase 5 — Nutrition
↓
Phase 6 — Scores & Enrichment
↓
Retailer Reference Dataset
reference-db/
│
├── contracts/
│ ├── ingestion/
│ ├── phase_1/
│ ├── phase_2/
│ ├── phase_3/
│ ├── phase_4/
│ ├── phase_5/
│ └── phase_6/
│
├── reference_data/
│ └── taxonomy/
│
├── src/
│ └── reference_db/
│ ├── adapters/
│ ├── ingestion/
│ ├── phase_1/
│ ├── phase_2/
│ ├── phase_3/
│ ├── phase_4/
│ ├── phase_5/
│ └── phase_6/
│
├── tests/
├── data/
├── README.md
├── pyproject.toml
└── .gitignore
The current implementation includes:
- Compliments
The adapter architecture is designed to support additional retailers such as Walmart, Costco, Metro, and Voilà / Sobeys. New retailer-specific adapters can be added without changing the generic pipeline phases.
- Python
- dlt — data ingestion
- DuckDB — analytical storage
- pytest — testing
- Keep retailer-specific logic inside adapters.
- Keep processing phases retailer-agnostic.
- Preserve source identifiers and provenance.
- Never invent missing source values.
- Keep identity separate from variants.
- Keep classification, taxonomy, and product grouping as separate concerns.
- Preserve traceability from processed records back to source data.
- Validate data at phase boundaries using executable data contracts.
- Prefer simple implementations before introducing additional infrastructure.
Validates the ingested product dataset against the ingestion contract and checks required fields, data types, identifiers, and duplicate records.
Builds a standardized product identity representation and separates identity attributes from variant attributes.
Classifies products, assigns the Reference DB taxonomy, and creates product groups based on product identity.
Models variants within existing product groups.
Standardizes and validates nutrition data and links it to products.
Calculates available product scores and assigns enrichment references such as Nutri-Score and Agribalyse references where possible.
Each pipeline stage has a versioned contract defining:
- Expected input schema
- Expected output schema
- Required and optional fields
- Processing responsibilities
- Quality expectations
- Primary identifiers
- Lineage
Contracts are stored under contracts/.
Raw and generated datasets are not committed to this repository by default.
The data/ directory is intended for local or generated pipeline data.
Create and activate the virtual environment:
python -m venv .venv
source .venv/bin/activateInstall the project in editable mode:
pip install -e .Run tests:
pytestThe repository currently contains the project architecture, contracts, and implementation skeleton.
Pipeline phases are being implemented incrementally and tested independently before being integrated into the complete workflow.
Each phase produces a quality result that determines whether its output can continue.
Possible statuses are:
PASS— continue normally.WARNING— continue and report the issue.AMBIGUOUS— quarantine the record for review.FAIL— stop processing the affected pipeline path.
Ambiguity is detected by the phase responsible for the decision.
For example:
- Phase 2 → ambiguous identity or normalization.
- Phase 3 → ambiguous classification, taxonomy, or grouping.
- Phase 4 → ambiguous variant.
Review is not a separate pipeline phase. After a decision is recorded, the affected record can be reprocessed from the phase where the ambiguity occurred.