Compare three fitted models and a simple baseline. Change the threshold to see which recorded outcomes each model identifies or misses.
Pythonscikit-learnHistorical outcomes
The model also flags many fully paid loans.
Question What errors accompany the validation-selected score threshold on held-out recorded outcomes?
Fixed comparison: L2 logistic regression at its 2015-validation-selected threshold of 0.23; 28,088 selected 2016 test loans. Controls below do not change this summary.
Supported finding
At threshold 0.23, the model flags 2,297 of 3,688 recorded Charged Off outcomes, or 62.28%, and also flags 9,159 Fully Paid loans. Only 20.05% of flagged records are Charged Off. These are retrospective labels, not credit decisions.
Decision implication
Before interpreting model ranking as a useful decision rule, an evaluation reviewer should inspect recall, precision and both error counts together. This example quantifies the trade-off; lending costs and an operational threshold are not established.
Limitation
The extract includes only Fully Paid and Charged Off loans, with unknown follow-up and unresolved borrower independence. Scores are not validated future-default probabilities. No fairness, lending profitability or production-readiness claim is made.
One threshold. Two kinds of error.
Positive means the recorded Charged Off outcome. A lower threshold generally identifies more of these outcomes while also flagging more Fully Paid loans.
Retrospective evaluation: the extract includes only Fully Paid and Charged Off loans, with an unknown follow-up horizon. The 2016 test Charged Off rate is 13.13%, versus 24.90% in validation. Scores are not validated future-default probabilities or credit decisions.
Start here Move the threshold to compare flagged and missed outcomes. Changing model restores its validation threshold.
Move the slider in 0.01 steps or enter an exact threshold. The validation-selected threshold is shown after the data loads.
Loading held-out evaluation…
Recorded outcomes at this threshold
Read the result
Precision = Charged Off among flagged records. Recall = flagged among actual Charged Off records. A false positive is a flagged Fully Paid record; a false negative is a missed Charged Off record. Precision is unavailable when nothing is flagged.
Ranking is only part of evaluation.
Every model uses the same held-out test cohort. These metrics do not change when the threshold moves.
ROC-AUC measures ranking discrimination. Average precision summarizes precision across recall levels; its no-skill reference is the test prevalence, 0.1313. Brier error is mean squared score error; lower is better. Small differences do not establish significant model superiority.
A corrected, separate experiment.
The new pipeline narrows leakage risks and exposes evaluation limits. Historical notebook, report and fairness claims are not reused.
Time split
Uniform sample of 60,000 loans issued in 2014 for training; all retained 2015 loans for validation; all retained 2016 loans for testing. Source coverage ends in 2016 despite the earlier brief’s 2018 endpoint. Unequal follow-up and repeat borrowers cannot be resolved.
Features and fitting
Fourteen application-related inputs. Imputation, scaling and encoding fit on training only; credit history uses each loan’s issue month. Address, postal, lender grade, rate, installment and free-text identity fields are excluded. Point-in-time field availability is not independently certified.
Selection
L1 strength uses validation average precision. Default thresholds maximize validation F1 over 0.05–0.95 in 0.01 steps; the highest threshold breaks ties. Test outcomes do not select parameters. Browser exploration changes only the displayed threshold.
Interpretation
Offline comparison of recorded outcomes only. No applicant scoring, lending profitability, protected-class fairness, reduced discrimination, regulatory compliance or production-readiness claim.
Source
Retained LendingClub CSV attributed to the Kaggle dataset named in the package. Exact bytes are verified. Acquisition date, upstream extraction and redistribution rights remain unverified. Raw borrower records and individual predictions are excluded.
Recreate the evaluation
Inspect aggregate results immediately. Full retraining requires authorized access to the exact historical source file identified in the package.
Included
Python code and tests included. Full retraining needs the original loan CSV.
Software
Python with the pinned packages. Synthetic tests run without borrower data.
Full replay
Authorized exact CSV required for full retraining; borrower records are excluded.
Reproduction package
Python pipeline, dependencies, model card, corrections, synthetic tests, results and independent metric checks.