Why Fitting Examples Can Mislead Us

ECE 57000 — September 25, 2026

David I. Inouye

Wednesday connected prediction goals, local evidence, and evaluation

  • We defined the prediction contract using only information available when the prediction is made.
  • Under an explicit i.i.d. working assumption, KNN uses nearby historical outcomes to estimate a local outcome probability.
  • Training error scores evidence the predictor already used; a holdout creates untouched evidence that can stand in for unseen cases.

Today: use held-out evidence to understand how model flexibility, data, and representation affect prediction on unseen cases.

A holdout lets observed data play the role of unseen data

We do not have outcomes for the genuinely unseen customers we ultimately care about. How can we estimate their error?

\[ \underbrace{\mathcal D_{\mathrm{train}}}_{\text{construct }f} \qquad\Big|\qquad \underbrace{\mathcal D_{\mathrm{holdout}}}_{\text{pretend unseen; score once}} \]

Reserve some observed pairs before fitting. Because their outcomes did not help construct \(f\), predictions on them mimic the information boundary of an unseen case—under the same sampling assumptions.

Cross-validation rotates which fold is held out

Split one observed dataset into folds. Each fold takes one turn as untouched evaluation data.

Fold 1 Fold 2 Fold 3 Fold 4 Fold 5
Fit 1 score \(e_1\) train train train train
Fit 2 train score \(e_2\) train train train
Fit 3 train train score \(e_3\) train train
Fit 4 train train train score \(e_4\) train
Fit 5 train train train train score \(e_5\)

\(K\) fits produce \(K\) errors, then we average: \(\widehat R_{\mathrm{CV}}=\frac1K\sum_{j=1}^{K}e_j\)

Leave-one-out cross-validation is the small-data extreme: \(K=n\).

The neighbor count controls how much local detail survives

A controlled bank-response illustration using age and completed prior contacts. With k equal to 1 the estimated probability surface follows individual labels, with k equal to 25 it tracks a smoother local pattern, and with k equal to all 500 examples it collapses to one global rate.

\(k=1\) can overfit individual outcomes; using every example erases useful local structure. Held-out error can reveal both failures.

Bias compares the average fitted estimate with the target relationship

At one fixed test point \(\boldsymbol{x}\), imagine refitting the method on many possible training samples. Their average prediction is where the method tends to land:

\[ \mu(\boldsymbol{x}) =\mathbb E_{\mathcal D}\!\left[\widehat\eta_{\mathcal D}(\boldsymbol{x})\right], \qquad \operatorname{Bias}(\boldsymbol{x}) =\mu(\boldsymbol{x})-\eta(\boldsymbol{x}). \]

Real data lets us observe how estimates vary across samples. It does not reveal the true \(\eta(\boldsymbol{x})\), so formal bias remains unknown.

Bias measures offset; variance measures spread

At age 40 (the only feature), each blue dot uses a different disjoint training set \(\mathcal D\) of \(n=1{,}647\) or \(1{,}648\) records. Each row reuses the same 25 sets.

At age 40, blue dots are training-set probability estimates, green diamonds are their mean approximating mu, and gold rings are full-data reference estimates used only as proxies for unknown eta. Spread about the diamond illustrates variance; the diamond-to-ring offset is only a proxy for bias.

\[ \operatorname{Bias}(\boldsymbol{x}) ={\color{#3f7f6b}\mu(\boldsymbol{x})} -{\color{#b8872f}\eta(\boldsymbol{x})} \]

\[ \operatorname{Var}_{\mathcal D}\!\left({\color{#2f6b8a}\widehat\eta_{\mathcal D}(\boldsymbol{x})}\right) =\mathbb E_{\mathcal D}\!\left[ \left({\color{#2f6b8a}\widehat\eta_{\mathcal D}(\boldsymbol{x})} -{\color{#3f7f6b}\mu(\boldsymbol{x})}\right)^2\right] \]

True \(\eta\) is unknown: the diamond-to-ring gap is a bias proxy. The reference may itself be biased.

Expected squared error combines bias, variance, and label noise

Draw a new input \(\mathbf{x}\) and binary outcome \(\textnormal{y}\) from the target population, independently of the training set \(\mathcal D\). How far is the predicted probability from the actual outcome, averaged over new cases and refits?

\[ \begin{gathered} \underbrace{\mathbb E_{\mathcal D,\mathbf{x},\textnormal{y}} \left[(\widehat\eta_{\mathcal D}(\mathbf{x})-\textnormal{y})^2\right]} _{\text{expected squared prediction error}} \\[0.6em] =\mathbb E_{\mathbf{x}}\!\left[ \underbrace{\operatorname{Bias}(\mathbf{x})^2}_{\text{systematic mismatch}} + \underbrace{\operatorname{Var}_{\mathcal D}(\widehat\eta_{\mathcal D}(\mathbf{x}))}_{\text{sample sensitivity}} + \underbrace{\eta(\mathbf{x})(1-\eta(\mathbf{x}))}_{\text{unresolved outcome variation}} \right]. \end{gathered} \]

Squared error loss gives this exact sum. Modeling choices can trade bias against variance; the represented inputs determine the remaining outcome noise.

Centering the outcome at \(\eta\) separates estimation error and noise

Fix \(\boldsymbol{x}\). Write \(\widehat\eta_{\mathcal D}=\widehat\eta_{\mathcal D}(\boldsymbol{x})\) and \(\eta=\mathbb E[\textnormal{y}\mid\boldsymbol{x}]\). The new outcome is independent of the training set given this input.

\(\displaystyle \mathbb E_{\mathcal D,\textnormal{y}}[(\widehat\eta_{\mathcal D}-\textnormal{y})^2\mid\boldsymbol{x}]\)

\(=\)

\(\displaystyle \mathbb E_{\mathcal D,\textnormal{y}}[((\widehat\eta_{\mathcal D}-\eta)+(\eta-\textnormal{y}))^2\mid\boldsymbol{x}]\)

(Add and subtract \(\eta\))

\(=\)

\(\displaystyle \begin{aligned}[t]&\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]+\mathbb E_{\textnormal{y}}[(\eta-\textnormal{y})^2\mid\boldsymbol{x}]\\&\quad+2\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}-\eta]\,\underbrace{\mathbb E_{\textnormal{y}}[\eta-\textnormal{y}\mid\boldsymbol{x}]}_{\eta-\mathbb E[\textnormal{y}\mid\boldsymbol{x}]=\eta-\eta=0}\end{aligned}\)

(Expand; independence factors the cross term, and centering makes it zero)

\(=\)

\(\displaystyle \mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]+\underbrace{\mathbb E[\textnormal{y}^2\mid\boldsymbol{x}]-\eta^2}_{\operatorname{Var}(\textnormal{y}\mid\boldsymbol{x})}\)

(Expand the remaining outcome square)

\(=\)

\(\displaystyle \underbrace{\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]}_{\text{probability-estimation error}}+\underbrace{\eta(1-\eta)}_{\text{outcome noise}}\)

(Binary \(\textnormal{y}^2=\textnormal{y}\), so \(\eta-\eta^2=\eta(1-\eta)\))

Centering the fit at \(\mu\) separates squared bias and variance

At the same fixed \(\boldsymbol{x}\), let \(\mu=\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}]\). Continue from the preceding slide, keeping its outcome-noise term.

\(\displaystyle \mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\eta)^2]+\eta(1-\eta)\)

\(=\)

\(\displaystyle \mathbb E_{\mathcal D}[((\widehat\eta_{\mathcal D}-\mu)+(\mu-\eta))^2]+\eta(1-\eta)\)

(Add and subtract \(\mu\))

\(=\)

\(\displaystyle \begin{aligned}[t]&\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\mu)^2]+(\mu-\eta)^2+\eta(1-\eta)\\&\quad+2(\mu-\eta)\,\underbrace{\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}-\mu]}_{\mathbb E_{\mathcal D}[\widehat\eta_{\mathcal D}]-\mu=\mu-\mu=0}\end{aligned}\)

(Expand; \(\mu-\eta\) is constant across training sets, and the centered fit has mean zero)

\(=\)

\(\displaystyle \underbrace{(\mu-\eta)^2}_{\text{squared bias}}+\underbrace{\mathbb E_{\mathcal D}[(\widehat\eta_{\mathcal D}-\mu)^2]}_{\text{variance}}+\underbrace{\eta(1-\eta)}_{\text{outcome noise}}\)

(The cross term is zero; reorder the remaining terms)

This exact sum uses squared error loss. Other losses need different analyses.

Averaging over inputs gives population squared prediction error

Now draw a new input \(\mathbf{x}\) from the target population and its outcome \(\textnormal{y}\). Average the pointwise identity over those inputs:

\[ \begin{gathered} \underbrace{\mathbb E_{\mathcal D,\mathbf{x},\textnormal{y}} [(\widehat\eta_{\mathcal D}(\mathbf{x})-\textnormal{y})^2]}_{\text{average error on new cases, across refits}} \\[0.6em] =\underbrace{\mathbb E_{\mathbf{x}}[\operatorname{Bias}(\mathbf{x})^2]}_{\text{average squared bias}} +\underbrace{\mathbb E_{\mathbf{x}}[\operatorname{Var}_{\mathcal D}(\widehat\eta_{\mathcal D}(\mathbf{x}))]}_{\text{average variance}} +\underbrace{\mathbb E_{\mathbf{x}}[\eta(\mathbf{x})(1-\eta(\mathbf{x}))]}_{\text{average outcome noise}}. \end{gathered} \]

  • Inputs count according to how often they occur in the target population. The decomposition stays exact, but its terms now summarize all inputs.
  • Average squared bias: positive and negative biases at different inputs cannot cancel each other.

Why it matters: smoothing can reduce sample sensitivity while increasing systematic mismatch. Held-out error measures their combined effect, plus outcome noise.

Increasing \(k\) exchanges local detail for stability

Smaller \(k\) Larger \(k\)
uses closer evidence reaches farther from the query
follows more sample-specific detail averages more observed outcomes
often lower local bias, higher variance often higher local bias, lower variance

Neither extreme is automatically best. Held-out error measures the combined effect for the population and representation we care about.

Data and representation improve the same local-evidence mechanism

Change What happens to a KNN neighborhood?
More representative training data genuinely similar examples can be closer
Better scaling or features distance favors differences relevant to the outcome
Larger \(k\) more outcomes are averaged, but evidence comes from farther away

All three change the evidence inside \(\widehat\eta_k(\boldsymbol{x})\) and therefore its bias and variance.

4. How can we choose a model without fooling ourselves?

Which neighbor count should we choose?

We fit KNN with \(k=1,5,25\) and choose the lowest training error. Inputs are distinct, and each training point counts as its own neighbor.

Is this a reliable way to choose the best predictor for new cases?

A. Yes: evaluating every candidate on the same examples makes the comparison fair.

B. No: training error rewards reproducing outcomes the predictor already used.

C. No: we should always choose the largest \(k\) to get the most stable predictions.

B. Here, \(k=1\) gets zero training error by retrieving each point’s own label. That does not establish how well it predicts a new outcome.

A held-out validation set lets us compare candidate values of \(k\)

Fit each candidate using \(\mathcal D_{\mathrm{train}}\). Evaluate its predictions on the same separate \(\mathcal D_{\mathrm{validation}}\), then choose the lowest validation error.

\[ \mathcal D_{\mathrm{train}} \xrightarrow{\ \text{fit}\ } f_1,f_5,f_{25} \xrightarrow{\ \text{score on validation}\ } \widehat k=\arg\min_{k\in\{1,5,25\}}\widehat R_{\mathrm{validation}}(f_k). \]

The validation outcomes are unavailable when fitting each candidate. They give us evidence about prediction on cases outside that candidate’s training set.

What does the selected model’s 4.5% error tell us?

We choose \(k\) with the lowest validation error and report its 4.5% as the selected model’s estimated population error.

Across repetitions of this procedure, is the reported error lower on average, unbiased, or higher on average?

A. Lower than the selected model’s true population error on average.

B. Unbiased for the selected model’s true population error.

C. Higher than the selected model’s true population error on average.

A: optimistic selection bias. Choosing the lowest observed error also favors favorable validation noise. Validation appropriately helps choose \(k\), but the winning score is no longer an unbiased final evaluation.

Validation is held out of each fit, but participates in selection

What are we evaluating? Did validation outcomes help construct it?
One candidate fitted at a fixed \(k\) No: only the training data determined that fit.
The selected pipeline, including the choice of \(k\) Yes: validation errors determined which candidate won.

Selection is part of learning. Information from validation enters through our choice of \(k\), even though the KNN fitting step never reads those labels.

The same reuse occurs when validation guides features, preprocessing, or model families. It leaks into a claimed final evaluation through those choices.

Separate validation for choosing from test data for reporting

Reserve all three roles before development:

Data Role What may it influence?
Training Fit each candidate. Its learned parameters or stored examples.
Validation Compare candidates and choose \(k\). The selected pipeline.
Test Score the settled pipeline once. The reported performance estimate.

Freeze the pipeline before using the test outcomes. An independent, representative test set gives an unbiased error estimate for that fixed pipeline, with finite-sample uncertainty.