ECE 57000 — September 25, 2026
Today: use held-out evidence to understand how model flexibility, data, and representation affect prediction on unseen cases.
We do not have outcomes for the genuinely unseen customers we ultimately care about. How can we estimate their error?
\[ \underbrace{\mathcal D_{\mathrm{train}}}_{\text{construct }f} \qquad\Big|\qquad \underbrace{\mathcal D_{\mathrm{holdout}}}_{\text{pretend unseen; score once}} \]
Reserve some observed pairs before fitting. Because their outcomes did not help construct \(f\), predictions on them mimic the information boundary of an unseen case—under the same sampling assumptions.
Split one observed dataset into folds. Each fold takes one turn as untouched evaluation data.
| Fold 1 | Fold 2 | Fold 3 | Fold 4 | Fold 5 | |
|---|---|---|---|---|---|
| Fit 1 | score \(e_1\) | train | train | train | train |
| Fit 2 | train | score \(e_2\) | train | train | train |
| Fit 3 | train | train | score \(e_3\) | train | train |
| Fit 4 | train | train | train | score \(e_4\) | train |
| Fit 5 | train | train | train | train | score \(e_5\) |
\(K\) fits produce \(K\) errors, then we average: \(\widehat R_{\mathrm{CV}}=\frac1K\sum_{j=1}^{K}e_j\)
Leave-one-out cross-validation is the small-data extreme: \(K=n\).
\(k=1\) can overfit individual outcomes; using every example erases useful local structure. Held-out error can reveal both failures.
At one fixed test point \(\boldsymbol{x}\), imagine refitting the method on many possible training samples. Their average prediction is where the method tends to land:
\[ \mu(\boldsymbol{x}) =\mathbb E_{\mathcal D}\!\left[\widehat\eta_{\mathcal D}(\boldsymbol{x})\right], \qquad \operatorname{Bias}(\boldsymbol{x}) =\mu(\boldsymbol{x})-\eta(\boldsymbol{x}). \]
Real data lets us observe how estimates vary across samples. It does not reveal the true \(\eta(\boldsymbol{x})\), so formal bias remains unknown.
At age 40 (the only feature), each blue dot uses a different disjoint training set \(\mathcal D\) of \(n=1{,}647\) or \(1{,}648\) records. Each row reuses the same 25 sets.
\[ \operatorname{Bias}(\boldsymbol{x}) ={\color{#3f7f6b}\mu(\boldsymbol{x})} -{\color{#b8872f}\eta(\boldsymbol{x})} \]
\[ \operatorname{Var}_{\mathcal D}\!\left({\color{#2f6b8a}\widehat\eta_{\mathcal D}(\boldsymbol{x})}\right) =\mathbb E_{\mathcal D}\!\left[ \left({\color{#2f6b8a}\widehat\eta_{\mathcal D}(\boldsymbol{x})} -{\color{#3f7f6b}\mu(\boldsymbol{x})}\right)^2\right] \]
True \(\eta\) is unknown: the diamond-to-ring gap is a bias proxy. The reference may itself be biased.
Draw a new input \(\mathbf{x}\) and binary outcome \(\textnormal{y}\) from the target population, independently of the training set \(\mathcal D\). How far is the predicted probability from the actual outcome, averaged over new cases and refits?
\[ \begin{gathered} \underbrace{\mathbb E_{\mathcal D,\mathbf{x},\textnormal{y}} \left[(\widehat\eta_{\mathcal D}(\mathbf{x})-\textnormal{y})^2\right]} _{\text{expected squared prediction error}} \\[0.6em] =\mathbb E_{\mathbf{x}}\!\left[ \underbrace{\operatorname{Bias}(\mathbf{x})^2}_{\text{systematic mismatch}} + \underbrace{\operatorname{Var}_{\mathcal D}(\widehat\eta_{\mathcal D}(\mathbf{x}))}_{\text{sample sensitivity}} + \underbrace{\eta(\mathbf{x})(1-\eta(\mathbf{x}))}_{\text{unresolved outcome variation}} \right]. \end{gathered} \]
Squared error loss gives this exact sum. Modeling choices can trade bias against variance; the represented inputs determine the remaining outcome noise.
Fix \(\boldsymbol{x}\). Write \(\widehat\eta_{\mathcal D}=\widehat\eta_{\mathcal D}(\boldsymbol{x})\) and \(\eta=\mathbb E[\textnormal{y}\mid\boldsymbol{x}]\). The new outcome is independent of the training set given this input.
\(\displaystyle \mathbb E_{\mathcal D,\textnormal{y}}[(\widehat\eta_{\mathcal D}-\textnormal{y})^2\mid\boldsymbol{x}]\)
\(=\)
\(\displaystyle \mathbb E_{\mathcal D,\textnormal{y}}[((\widehat\eta_{\mathcal D}-\eta)+(\eta-\textnormal{y}))^2\mid\boldsymbol{x}]\)
(Add and subtract \(\eta\))
\(=\)
\(\displaystyle \begin{aligned}[t]&\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]+\mathbb E_{\textnormal{y}}[(\eta-\textnormal{y})^2\mid\boldsymbol{x}]\\&\quad+2\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}-\eta]\,\underbrace{\mathbb E_{\textnormal{y}}[\eta-\textnormal{y}\mid\boldsymbol{x}]}_{\eta-\mathbb E[\textnormal{y}\mid\boldsymbol{x}]=\eta-\eta=0}\end{aligned}\)
(Expand; independence factors the cross term, and centering makes it zero)
\(=\)
\(\displaystyle \mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]+\underbrace{\mathbb E[\textnormal{y}^2\mid\boldsymbol{x}]-\eta^2}_{\operatorname{Var}(\textnormal{y}\mid\boldsymbol{x})}\)
(Expand the remaining outcome square)
\(=\)
\(\displaystyle \underbrace{\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]}_{\text{probability-estimation error}}+\underbrace{\eta(1-\eta)}_{\text{outcome noise}}\)
(Binary \(\textnormal{y}^2=\textnormal{y}\), so \(\eta-\eta^2=\eta(1-\eta)\))
At the same fixed \(\boldsymbol{x}\), let \(\mu=\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}]\). Continue from the preceding slide, keeping its outcome-noise term.
\(\displaystyle \mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]+\eta(1-\eta)\)
\(=\)
\(\displaystyle \mathbb E_{\mathcal D}[((\widehat\eta_{\mathcal D}-\mu)+(\mu-\eta))^2]+\eta(1-\eta)\)
(Add and subtract \(\mu\))
\(=\)
\(\displaystyle \begin{aligned}[t]&\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\mu)^2]+(\mu-\eta)^2+\eta(1-\eta)\\&\quad+2(\mu-\eta)\,\underbrace{\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}-\mu]}_{\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}]-\mu=\mu-\mu=0}\end{aligned}\)
(Expand; \(\mu-\eta\) is constant across training sets, and the centered fit has mean zero)
\(=\)
\(\displaystyle \underbrace{(\mu-\eta)^2}_{\text{squared bias}}+\underbrace{\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\mu)^2]}_{\text{variance}}+\underbrace{\eta(1-\eta)}_{\text{outcome noise}}\)
(The cross term is zero; reorder the remaining terms)
This exact sum uses squared error loss. Other losses need different analyses.
Now draw a new input \(\mathbf{x}\) from the target population and its outcome \(\textnormal{y}\). Average the pointwise identity over those inputs:
\[ \begin{gathered} \underbrace{\mathbb E_{\mathcal D,\mathbf{x},\textnormal{y}} [(\widehat\eta_{\mathcal D}(\mathbf{x})-\textnormal{y})^2]}_{\text{average error on new cases, across refits}} \\[0.6em] =\underbrace{\mathbb E_{\mathbf{x}}[\operatorname{Bias}(\mathbf{x})^2]}_{\text{average squared bias}} +\underbrace{\mathbb E_{\mathbf{x}}[\operatorname{Var}_{\mathcal D}(\widehat\eta_{\mathcal D}(\mathbf{x}))]}_{\text{average variance}} +\underbrace{\mathbb E_{\mathbf{x}}[\eta(\mathbf{x})(1-\eta(\mathbf{x}))]}_{\text{average outcome noise}}. \end{gathered} \]
Why it matters: smoothing can reduce sample sensitivity while increasing systematic mismatch. Held-out error measures their combined effect, plus outcome noise.
| Smaller \(k\) | Larger \(k\) |
|---|---|
| uses closer evidence | reaches farther from the query |
| follows more sample-specific detail | averages more observed outcomes |
| often lower local bias, higher variance | often higher local bias, lower variance |
Neither extreme is automatically best. Held-out error measures the combined effect for the population and representation we care about.
| Change | What happens to a KNN neighborhood? |
|---|---|
| More representative training data | genuinely similar examples can be closer |
| Better scaling or features | distance favors differences relevant to the outcome |
| Larger \(k\) | more outcomes are averaged, but evidence comes from farther away |
All three change the evidence inside \(\widehat\eta_k(\boldsymbol{x})\) and therefore its bias and variance.
Choosing a predictor and measuring its performance are two different uses of data.
We fit KNN with \(k=1,5,25\) and choose the lowest training error. Inputs are distinct, and each training point counts as its own neighbor.
Is this a reliable way to choose the best predictor for new cases?
A. Yes: evaluating every candidate on the same examples makes the comparison fair.
B. No: training error rewards reproducing outcomes the predictor already used.
C. No: we should always choose the largest \(k\) to get the most stable predictions.
B. Here, \(k=1\) gets zero training error by retrieving each point’s own label. That does not establish how well it predicts a new outcome.
Fit each candidate using \(\mathcal D_{\mathrm{train}}\). Evaluate its predictions on the same separate \(\mathcal D_{\mathrm{validation}}\), then choose the lowest validation error.
\[ \mathcal D_{\mathrm{train}} \xrightarrow{\ \text{fit}\ } f_1,f_5,f_{25} \xrightarrow{\ \text{score on validation}\ } \widehat k=\arg\min_{k\in\{1,5,25\}}\widehat R_{\mathrm{validation}}(f_k). \]
The validation outcomes are unavailable when fitting each candidate. They give us evidence about prediction on cases outside that candidate’s training set.
We choose \(k\) with the lowest validation error and report its 4.5% as the selected model’s estimated population error.
Across repetitions of this procedure, is the reported error lower on average, unbiased, or higher on average?
A. Lower than the selected model’s true population error on average.
B. Unbiased for the selected model’s true population error.
C. Higher than the selected model’s true population error on average.
A: optimistic selection bias. Choosing the lowest observed error also favors favorable validation noise. Validation appropriately helps choose \(k\), but the winning score is no longer an unbiased final evaluation.
| What are we evaluating? | Did validation outcomes help construct it? |
|---|---|
| One candidate fitted at a fixed \(k\) | No: only the training data determined that fit. |
| The selected pipeline, including the choice of \(k\) | Yes: validation errors determined which candidate won. |
Selection is part of learning. Information from validation enters through our choice of \(k\), even though the KNN fitting step never reads those labels.
The same reuse occurs when validation guides features, preprocessing, or model families. It leaks into a claimed final evaluation through those choices.
Reserve all three roles before development:
| Data | Role | What may it influence? |
|---|---|---|
| Training | Fit each candidate. | Its learned parameters or stored examples. |
| Validation | Compare candidates and choose \(k\). | The selected pipeline. |
| Test | Score the settled pipeline once. | The reported performance estimate. |
Freeze the pipeline before using the test outcomes. An independent, representative test set gives an unbiased error estimate for that fixed pipeline, with finite-sample uncertainty.