# Retail analytics reproduction package

This package contains the historical UCI retail analytical source, accepted aggregate outputs, a portable rerun path, a synthetic smoke fixture, and browser-ready revenue, retention and product aggregates. It is a portfolio case study, not a current company operating system.

## Run without the full dataset

With Python 3.12 or later, run `python scripts/verify_package.py` from this directory. This uses only the standard library, validates shipped file hashes, checks browser aggregate identities, and executes the synthetic fixture through an independent calculation. It does not validate the full pipeline, SQL or Power BI.

## Reproduce the full Python analysis

1. Install uv. The original runtime used uv 0.11.32 and Python 3.12.13. `uv.lock` and `pyproject.toml` are retained byte-for-byte.
2. Run `./scripts/bootstrap.ps1`, or `uv sync --locked --python 3.12.13` on another platform.
3. Run `uv run --locked python scripts/acquire_source.py --data-root data`. Alternatively place the checksum-matched workbook at `DATA_ROOT/uci-online-retail/Online Retail.xlsx`.
4. On PowerShell run `./scripts/run.ps1 -DataRoot data`. Outputs are confined to this extracted package. On another platform set `RETAIL_DATA_ROOT` to an absolute data directory and `RETAIL_OUTPUT_ROOT` to this package's `02_analysis/outputs`; run `build_analytics_data.py`, `build_decision_analysis.py`, and `validate_outputs.py` from `02_analysis/code` with `uv run --locked python`, then `scripts/export_explorer.py --project . --out browser-data` and `scripts/verify_package.py --full`.

The full pipeline generates 684,129 sales lines, including 534,129 public-core lines and 150,000 simulated physical-store lines, 300,000 simulated sessions and 1,122 simulated spend rows. Allow disk space for large generated CSVs. The workbook SHA256 is `43465a06f2ccf7c8b5bd2892bc7defb52f97487934fe93b16ae4c3936424676d`.

## Data and metric rules

`browser-data/retail-explorer.json` separates `public_core` from `combined`. Combined includes the public core and simulated physical-store records, so never sum across these scopes. Revenue is signed quantity times unit price after discount, including returns. Orders count distinct non-return, positive-revenue order IDs. AOV uses net revenue divided by that count. Monthly retention is the intersection of current and previous monthly identified purchaser sets divided by the previous set size. First-month retention is null, not zero. Do not sum customer counts across months. Repeat rate counts identified customers with at least two distinct completed positive orders across the entire source window. Do not recalculate a time-filtered repeat KPI from the all-period number.

December 2011 ends on the ninth and is partial. `complete_month=false` is present on that month's records. Use complete months by default. A selected interval retention value should be the arithmetic mean of available monthly rates, matching the accepted definition, with the monthly denominator still based on the immediately preceding source month. A filter must not rewrite that denominator.

Product rows contain the top 20 products by full-period revenue in each scope plus an `ALL_OTHER` aggregate. Ranking is fixed to that full-period scope, not recalculated over every filter. `ALL_OTHER` preserves revenue totals. Product name comes from the retained product dimension. Category classifications and all costs are simulated attributes and are excluded from the public-core product view.

Accepted category, store and customer-segment tables describe the combined simulated business case. All margin and CLV figures use simulated costs. The combined-model 100% repeat rate is a simulation artifact and is not a real customer-performance KPI. Attribution does not establish causal marketing lift.

## Provenance and reuse

See `ATTRIBUTION.md`, `evidence/handoff.md`, `evidence/baseline-file-identities.json` and `verification-report.json`. The stale `headline_metrics.csv` is deliberately not shipped: its Champions monetary-share precision disagrees with the accepted runtime contract and recomputed segment output. No accepted figures are rewritten. Original analysis files, raw data, final PBIX and live registers remain unchanged.

The package does not include raw customer-level data, customer identifiers, credentials, machine-specific paths, the PBIX or a native release claim. SQL and DAX are included for technical inspection and optional reproduction. Historical Power BI acceptance and the six reconciliation scenarios do not validate new browser interactions or arbitrary DAX filters. Source distribution terms are documented separately from any downstream publication decision.
