Regression: a numerical price, \(\widehat y=\boldsymbol{w}^{\mathsf T}\boldsymbol{x}+b\).
Classification: a positive/negative review, with an estimated class probability.
Return to our real Cornell case: 18,746 word features, 1,200 training reviews.
1. Model — now: turn the score into a probability.
2. Objective — next: measure how well those probabilities predict the labels.
3. Solution — afterward: learn the coefficients by numerical optimization.
For one review, let \(\boldsymbol{x}\in\mathbb R^d\) contain its word features.
\[ z=\boldsymbol{w}^{\mathsf T}\boldsymbol{x}+b =\sum_{j=1}^d w_jx_j+b. \]
With \(d=18{,}746\), the model learns 18,747 parameters, not one parameter per review.
Illustrative two-feature review, with invented weights:
| Feature | Measurement \(x_j\) | Weight \(w_j\) | Contribution \(w_jx_j\) |
|---|---|---|---|
| “engaging” | 2 | \(+0.8\) | \(+1.6\) |
| “dull” | 1 | \(-1.1\) | \(-1.1\) |
| Offset | — | \(b=-0.2\) | \(-0.2\) |
\[z=1.6-1.1-0.2=0.3\quad\Longrightarrow\quad\widehat y=1.\]
The model combines evidence; a word’s contribution depends on both its value and its weight.
\[ \widehat y=\mathbf1\{\boldsymbol{w}^{\mathsf T}\boldsymbol{x}+b\ge0\}, \qquad \text{boundary: }\boldsymbol{w}^{\mathsf T}\boldsymbol{x}+b=0. \]
Learning chooses which direction matters for the labels, rather than which examples are nearby in every coordinate.
\[ \widehat\eta(\boldsymbol{x})=\sigma(z)=\frac{1}{1+e^{-z}}, \qquad z=\boldsymbol{w}^{\mathsf T}\boldsymbol{x}+b. \]
| Score \(z\) | Probability of label 1 |
|---|---|
| \(-4\) | 1.8% |
| \(-2\) | 11.9% |
| \(0\) | 50% |
| \(2\) | 88.1% |
| \(4\) | 98.2% |
Logistic regression uses this probability model.
\[ \widehat\eta(\boldsymbol{x})\ge\tfrac12 \quad\Longleftrightarrow\quad \boldsymbol{w}^{\mathsf T}\boldsymbol{x}+b\ge0. \]
Model: a linear score followed by a sigmoid.
Now: what loss should tell us whether its probabilities fit the labels?
Next: solve for the coefficients that minimize that loss.
For one positive review, compare scores \(z=-3\) and \(z=-0.1\).
| Score | Predicted probability | Predicted class | Classification error |
|---|---|---|---|
| \(-3\) | 4.7% | 0 | 1 |
| \(-0.1\) | 47.5% | 0 | 1 |
The probability moved toward the observed label, but accuracy sees no improvement yet.
A smooth loss can provide information about such changes before the class flips.
Let \(a=\widehat\eta(\boldsymbol{x})\) and \(y\in\{0,1\}\).
\[ \ell(y,a)=-y\log a-(1-y)\log(1-a) =\begin{cases}-\log a,&y=1,\\-\log(1-a),&y=0.\end{cases} \]
Give the observed outcome high probability → small loss.
Give the observed outcome nearly zero probability → large loss.
For a positive outcome, \(y=1\):
| Predicted probability \(a\) | Class at threshold 0.5 | Log loss \(-\log a\) |
|---|---|---|
| 0.99 | Correct | 0.010 |
| 0.60 | Correct | 0.511 |
| 0.40 | Wrong | 0.916 |
| 0.01 | Wrong | 4.605 |
Unlike classification error, log loss distinguishes how confidently a prediction was made.
Two observed outcomes are both positive: \(y_1=y_2=1\).
| Model | Predicted probabilities | Accuracy | Average log loss |
|---|---|---|---|
| A | \((0.60,\;0.60)\) | 100% | 0.511 |
| B | \((0.99,\;0.49)\) | 50% | 0.362 |
What does this comparison establish?
A. Accuracy and log loss always rank models identically.
B. Better log loss can coexist with worse thresholded accuracy.
C. Model B must generalize better to every new population.
B. The criteria assess different outputs. Neither two-example score establishes population performance.
| Numerical regression | Logistic probability prediction |
|---|---|
| Output: \(\widehat y=z\) | Output: \(a=\sigma(z)\) |
| Loss: \(\tfrac12(z-y)^2\) | Loss: \(-y\log a-(1-y)\log(1-a)\) |
| Penalize numerical residual size. | Penalize low probability of the observed label. |
Squared error on probabilities is also possible, but composing it with a sigmoid generally loses convexity in the coefficients.
With \(z_i=\boldsymbol{w}^{\mathsf T}\boldsymbol{x}_i+b\) and \(a_i=\sigma(z_i)=e^{z_i}/(1+e^{z_i})\):
\[\log a_i=z_i-\log(1+e^{z_i}),\qquad \log(1-a_i)=-\log(1+e^{z_i}).\]
\(\displaystyle \ell(y_i,a_i)\)
\(=\)
\(\displaystyle -y_i\log a_i-(1-y_i)\log(1-a_i)\)
(Usual binary log loss)
\(\displaystyle \,\)
\(=\)
\(\displaystyle -y_i z_i+\big[y_i+(1-y_i)\big]\log(1+e^{z_i})\)
(Substitute the two log identities)
\(\displaystyle \,\)
\(=\)
\(\displaystyle \log(1+e^{z_i})-y_i z_i\)
(Combine the coefficients)
Average the same loss over observations:
\[J(\boldsymbol{w},b)=\frac1n\sum_{i=1}^n\ell(y_i,a_i) =\frac1n\sum_{i=1}^n\left[\log(1+e^{z_i})-y_i z_i\right].\]
OLS gave linear stationary equations; logistic loss gives nonlinear equations in the coefficients.
Next: what does the objective’s shape guarantee?