ECE 57000 — September 23, 2026
We will go through the revised Arc 3 opening again, but more quickly this time.
The slides are unchanged from the new arc sequence. Because the ideas should now be review, we will focus on their connections and move faster through details.
David I. Inouye
Last updated: September 22, 2026
What would justify trusting a prediction for someone we have never seen?
| Setting | Prediction question with two possible answers |
|---|---|
| Bank phone campaign | Will this customer open a term-deposit account? |
| Wine screening | Will this wine receive a high sensory rating? |
| Activity recognition | Is this person walking? |
Each uses information available now to predict an outcome we do not yet know.
Our running case: the bank phone campaign. A term-deposit account holds money for a fixed period to earn interest.
| Practical question | Bank-case answer | Mathematical counterpart |
|---|---|---|
| What information is available? | New customer information known before the call | ? |
| What answer do we want? | Whether the customer opens the offered account | ? |
| What makes a predictor good? | Few mistakes across unseen customers | ? |
We will fill in one mathematical counterpart at a time.
Question: What information is available?
Bank-case answer: New customer information known before the call.
Formalization:
\[\boldsymbol{x}\in\mathcal X\]
\(\boldsymbol{x}\) is one customer’s available information; \(\mathcal X\) is the set of valid input representations.
Question: What answer do we want?
Bank-case answer: Whether the customer opens the offered account.
Formalization:
\[ \boldsymbol{x} \quad\xrightarrow{\quad f\quad}\quad \widehat y\in\{0,1\} \]
\(y\in\{0,1\}\) is the actual outcome. The prediction is correct when \(\widehat y=y\).
A distribution \(p\) describes how common different unseen customer–outcome pairs are. A random pair \((\mathbf{x},\textnormal{y})\) drawn from \(p\) contains the available information and what actually happens.
The population error is the chance that our prediction is wrong:
\[ R_p(f)=\Pr_{(\mathbf{x},\textnormal{y})\sim p} \bigl(f(\mathbf{x})\ne\textnormal{y}\bigr). \]
Goal: choose \(f^*\in\operatorname*{arg\,min}_{f:\mathcal X\to\{0,1\}} R_p(f)\).
This completes the goal—but it does not tell us how to find \(f\).
| Practical question | Bank-case answer | Mathematical counterpart |
|---|---|---|
| What information is available? | New customer information known before the call | Input \(\boldsymbol{x}\in\mathcal X\) |
| What answer do we want? | Whether the customer opens the offered account | Outcome \(y\in\{0,1\}\); prediction \(\widehat y=f(\boldsymbol{x})\) |
| What makes a predictor good? | Few mistakes across unseen customers | Small population error \(R_p(f)\) |
The unresolved question: How do we find the right function \(f\)?
In ordinary computation, the rule is given. For example, a circuit simulator is given equations and parameters, then computes an output from an input.
\[f\ \text{given},\qquad \boldsymbol{x}\longmapsto f(\boldsymbol{x})\]
\[\mathcal D=\{(\boldsymbol{x}_i,y_i)\}_{i=1}^{n} \quad\longmapsto\quad \text{infer } f\]
The historical input–outcome pairs \(\mathcal D\) are the training data. They give examples of the mapping, but not the rule itself.
The real prediction context defines what each \(\boldsymbol{x}_i\) may contain.
Prediction moment: before calling this customer.
| Historical column | Status at prediction time |
|---|---|
| Recorded age and job | Available before the call |
| Duration of this call | Unavailable until after the call |
| Whether the customer opens this account | The outcome \(y_i\), not an input |
The dataset records each call’s duration after that call ends.
Can call duration be used to predict whether this customer will open the account before the call begins?
A. Yes, because it appears in the historical training data.
B. Yes, if the historical dataset is sufficiently large.
C. No, because it is unavailable at the prediction moment.
C. No. Training inputs must represent what would actually be known when the prediction is made; more records do not repair that mismatch.
Larger principle: available columns describe a dataset; allowed inputs are determined by the real decision we are trying to support.
Training pairs are the evidence. Performance on unseen pairs is the claim. An explicit assumption must connect them:
\[ \underbrace{\text{training pairs}}_{\text{observed evidence}} \quad\xrightarrow{\quad\text{assumption}\quad}\quad \underbrace{\text{unseen pairs}}_{\text{desired claim}} \]
The most common simple starting point is i.i.d. sampling: model the training pairs and unseen pairs as independent draws from the same distribution \(p\).
I.i.d. is deliberately naive. When it reasonably approximates the problem, it provides a useful foundation for analysis. We will use it for now.
The first “I” in i.i.d. is independent.
Knowing one sampled pair \((\mathbf{x}_i,\textnormal{y}_i)\) should not change the probability of another draw once the population distribution is fixed. The unseen case should also be independent of the sample used to construct \(f\).
Possible violation: if the same customer appears after several campaign contacts, those rows and outcomes may be linked rather than independent examples.
The second “I” in i.i.d. is identically distributed.
Training pairs and the unseen pairs we care about are modeled as draws from the same joint distribution \(p\)—both the inputs and how they relate to the outcome.
Possible violation: if the bank changes who it contacts or changes the offered account, deployment cases may no longer follow the training distribution.
Same distribution does not mean the same customers or repeated feature values.
The contract asks for low error on unseen customers using only information available before the call. Training pairs become evidence under our working assumption that they and the unseen cases are connected.
First idea: find past customers similar to the new customer and consult their outcomes.
Next question: What should “similar” mean, and why should their outcomes help?
KNN makes learning from local evidence visible without making the algorithm the point.
For a new customer \(\boldsymbol{x}\):
\[ \widehat\eta_k(\boldsymbol{x}) =\frac{1}{k}\sum_{i\in\mathcal N_k(\boldsymbol{x})}y_i, \qquad \widehat y=\mathbb 1\!\left[\widehat\eta_k(\boldsymbol{x})\geq\tfrac12\right]. \]
The method constructs a prediction directly from the stored examples.
Distance measures differences in the coordinates we supply—not an intrinsic notion of customer similarity.
For each feature \(j\), estimate its training mean and standard deviation, then use
\[ x'_{ij}=\frac{x_{ij}-\widehat\mu_j}{\widehat s_j}. \]
Interpretation: every retained coordinate has empirical variance one before distance is computed.
The population quantity is
\[ \eta(\boldsymbol{x}) =\Pr(\textnormal{y}=1\mid\mathbf{x}=\boldsymbol{x}). \]
KNN substitutes nearby observed outcomes for repeated outcomes at exactly the same input:
\[ \widehat\eta_k(\boldsymbol{x}) =\frac{\text{nearby customers who opened the account}}{k}. \]
Similar customers need not have identical outcomes; their probabilities must be locally related.
Here \(k=\lfloor\sqrt n\rfloor\), so \(k\to\infty\) while \(k/n\to0\): the same query averages more labels inside a smaller neighborhood.
Local averaging recovers slowly varying probabilities sooner. Fine-scale or choppy structure can remain blurred until the sample is extraordinarily dense.
Fit
Predict one input
Small reference sets can be cheap; very large ones motivate models that compress training evidence into learned parameters.
Observed success may reflect memorization, sample luck, or the wrong flexibility.
Assume every recorded input is distinct. We fit each KNN model, then compute its error on those same examples.
Which value of \(k\) should we select using only this error estimate?
A. \(k=1\)
B. \(k=15\)
C. \(k=151\)
A. \(k=1\). Every training example is its own nearest neighbor, so this procedure reports zero error.
That makes \(k=1\) look perfect—but our goal concerns unseen customers.
| Quantity | Examples being scored | Question answered |
|---|---|---|
| Training error | Used to construct \(f\) | How well does \(f\) reproduce evidence it already saw? |
| Population error \(R_p(f)\) | Unseen draws from \(p\) | How often will \(f\) fail on the cases we care about? |
The same observations cannot simultaneously serve as unrestricted construction material and untouched evidence of generalization.
We do not have outcomes for the genuinely unseen customers we ultimately care about. How can we estimate their error?
\[ \underbrace{\mathcal D_{\mathrm{train}}}_{\text{construct }f} \qquad\Big|\qquad \underbrace{\mathcal D_{\mathrm{holdout}}}_{\text{pretend unseen; score once}} \]
Reserve some observed pairs before fitting. Because their outcomes did not help construct \(f\), predictions on them mimic the information boundary of an unseen case—under the same sampling assumptions.