# Loan outcome evaluation

Private educational portfolio edition, `loan-portfolio-20261004-v1`. Rebuilds an offline model comparison from a checksum-identified historical file. It supplies aggregate model metrics and threshold tables for an evaluation explorer. It contains no applicant scoring service, approval rule, geographic fairness score or fitted model.

## Reproduce

Verified runtime: Python 3.12.13. Install the exact runtime dependencies in `requirements.txt` in an isolated environment.

```text
python -m venv .venv
python -m pip install -r requirements.txt
python -m unittest -v test_evaluation
python evaluate.py --data /your/authorized/lending_club_loan_two.csv --output results --private-audit private_audit
python verify_results.py --results results --private-audit private_audit
```

Use the environment's Python interpreter after creating it. On Windows this is `.venv\Scripts\python.exe`; on Unix-like systems it is `.venv/bin/python`. Paths in the command above are placeholders supplied by the reproducer. The dataset is not bundled or fetched automatically. The recorded duration is `elapsed_seconds` in `results/run_manifest.json`; other machines may differ.

The required CSV is 99,957,365 bytes, SHA-256 `04b6ff3660ba9b7125453c3b08ef04c418ee03d74e90947daa2b2f9b51939621`. The runner rejects other bytes. The recorded source attribution is [Kaggle, Lending Club loan Data](https://www.kaggle.com/datasets/sahilnbajaj/lending-club-loan-data). That attribution does not establish that a current download matches this retained file. Acquisition date, extraction procedure and redistribution rights are unverified. Obtain authorized matching bytes separately. Synthetic tests run without the dataset.

`private_audit/predictions.npz` contains individual labels and prediction scores only when the optional flag is used. Keep this directory private. Do not add it to the website or downloadable package. It is required for the independent verifier, which recomputes ROC-AUC, average precision, Brier error, every confusion matrix and validation threshold selection using separate NumPy calculations.

## Design and outputs

- Positive outcome: `Charged Off = 1`; negative outcome: `Fully Paid = 0`. Unexpected labels cause a failure.
- Training: 60,000 of 102,860 loans issued in 2014, sampled uniformly without replacement with seed 20261004. The training cap applies only to this year. There is no outcome-based training sampling.
- Validation: all 94,264 loans issued in 2015. Test: all 28,088 loans issued in 2016. All 170,818 records outside 2014–2016 are excluded. Unused 2014 training rows total 42,860.
- Numeric imputation, missing indicators, scaling and categorical encoding fit on training rows only. Each model is a complete pipeline. Credit-history length uses the individual loan's issue month.
- Predictors: loan amount, term, purpose, home ownership, log annual income, debt-to-income ratio, revolving utilization, open accounts, public records, mortgage accounts, total accounts, bankruptcies, credit-history months and employment length. Their point-in-time availability is not independently verified.
- Excluded: outcome labels, address/postal fields, employer/title free text, grade, subgrade, interest rate, installment, verification status and application/listing status. Grade, rate and payment fields are lender-generated or contract terms whose timing is insufficiently established for the intended application-information comparison.
- Models: training-prevalence baseline, L2 logistic regression, L1 logistic regression and random forest. No class reweighting, resampling or oversampling. Report minority-outcome metrics alongside prevalence and accuracy.
- L1 chooses `C` from `[0.001, 0.01, 0.1, 1.0]` using validation average precision. Thresholds choose maximum validation F1 on the fixed 0.05–0.95 grid. Test outcomes do not select parameters or thresholds. All four model families are reported; no final operational model is selected.
- `results/evaluation.json`: model and threshold data for the browser. `data_profile.json`: source and split reconciliation. `run_manifest.json`: source, code, software and output identities. `independent_verification.json`: separate calculation checks. `test_results.json`: synthetic test receipt.

The result JSON contains threshold rows at 0.00–1.00 in 0.01 increments. The rule is `score >= threshold` predicts the recorded Charged Off label. “Positive” always means that retrospective outcome, never credit denial or borrower eligibility. A threshold control must preserve this wording and display its cohort denominator.

## Limits beside every performance display

This file contains only Fully Paid and Charged Off outcomes. Outcome selection, the unknown acquisition date and unknown follow-up horizon make the 2014–2016 cohorts unsuitable for a claim of equal-maturity or future-default validation. The 2016 test rate is 13.13%, versus 24.90% in validation; the reasons cannot be separated from this file. Borrower identity is unavailable, so repeat borrowers across years cannot be ruled out.

The model comparison describes this retained dataset only. It does not establish fairness, causal effects, lending profitability, generalization to rejected or current applicants, regulatory compliance or production readiness. L1 coefficient selection is not evidence of reduced discrimination. See `MODEL_CARD.md` and `CORRECTIONS.md` for the exact scope and changes.

Raw records, addresses, personal text, private paths, row-level predictions and trained models are excluded from the downloadable archive. The rights check remains unresolved even though the retained input's bytes are verified. This package is for the authorized owner-private portfolio; it does not grant data redistribution rights.
