Why the Shape of the Objective Matters

David I. Inouye

Which loss would be easiest to optimize?

Which shape gives us the clearest route to a global minimum?

A. The jumpy loss.

B. The smooth loss with several minima.

C. The convex bowl.

C. A local minimum is already global; there are no worse local minima to escape.

A useful objective lets us reason about the solution

  • Jumpy: a small change may reveal little—or abruptly change the loss.
  • Smooth, several minima: downhill progress can end at a worse local minimum.
  • Convex: a local minimum is global, making optimality easier to analyze.

Convexity is a useful class of well-behaved objectives.

It gives a global guarantee; speed still depends on the problem and algorithm.

Objective shape answers three linked questions

  1. Define: what does convexity mean geometrically?
  2. Check: are our OLS and logistic objectives convex?
  3. Interpret: what does that guarantee—and what else do we need?

Convexity is a statement about the entire objective shape

A convex objective may have a single bottom—or a flat set of equally good solutions.

Definition: a convex function lies below every chord

\(J\) is convex on a convex domain if, for every \(\boldsymbol{u},\boldsymbol{v}\) in the domain and \(0\le t\le1\),

\[{\color{#27847f}{J\big(}}{\color{#b8872f}{(1-t)}}{\color{#2f6b8a}{\boldsymbol{u}}}+{\color{#b8872f}{t}}{\color{#756486}{\boldsymbol{v}}}{\color{#27847f}{\big)}} \le {\color{#b8872f}{(1-t)}}{\color{#2f6b8a}{J(\boldsymbol{u})}}+{\color{#b8872f}{t}}{\color{#756486}{J(\boldsymbol{v})}}.\]

Loss at the averaged input ≤ average of the endpoint losses.

Convexity makes every local minimum global

Suppose \(\boldsymbol{u}\) is a local minimum, but some \(\boldsymbol{v}\) has lower loss.

\[J((1-t)\boldsymbol{u}+t\boldsymbol{v})\le(1-t)J(\boldsymbol{u})+tJ(\boldsymbol{v}).\]

Since \(J(\boldsymbol{v})<J(\boldsymbol{u})\), the right side is less than \(J(\boldsymbol{u})\) for every \(t>0\).

Choosing an arbitrarily small \(t\) gives a nearby point with lower loss—a contradiction.

Therefore, a local minimum must be global.

For a twice differentiable function, convexity means nonnegative curvature

Let \(J\) be twice continuously differentiable on an open convex domain.

Parameters Equivalent criterion for convexity
One scalar \(\theta\) \(J''(\theta)\ge0\) everywhere
A vector \(\boldsymbol{\theta}\) \(H(\boldsymbol{\theta})=\nabla^2J(\boldsymbol{\theta})\succeq0\) everywhere

The Hessian is the matrix of second partial derivatives: \(H_{jk}=\frac{\partial^2J}{\partial\theta_j\partial\theta_k}\).

Positive semidefinite (PSD) means \(\boldsymbol{v}^{\mathsf T}H\boldsymbol{v}\ge0\) for every direction \(\boldsymbol{v}\). Equivalently, all eigenvalues of \(H\) are \(\ge0\): no eigenvector direction curves downward; zero allows a flat direction at that point.

Derivation hints: in one dimension, \(J''\ge0\) makes slopes nondecreasing, giving the chord inequality. In several dimensions, restrict to a line:

\[g(t)=J(\boldsymbol{\theta}+t\boldsymbol{v}),\qquad g''(t)=\boldsymbol{v}^{\mathsf T}H(\boldsymbol{\theta}+t\boldsymbol{v})\boldsymbol{v}.\]

OLS and logistic log loss have nonnegative directional curvature

Append the intercept to \(\widetilde X\) and the coefficients to \(\boldsymbol{\theta}\).

Objective Hessian \(H=\nabla^2J\)
OLS \(\frac1n\widetilde X^{\mathsf T}\widetilde X\)
Logistic log loss \(\frac1n\widetilde X^{\mathsf T}W\widetilde X\), with \(W_{ii}=a_i(1-a_i)\ge0\)

For logistic loss, \(\ell(z,y)=\log(1+e^z)-yz\) gives \(\ell'=\sigma(z)-y\) and \(\ell''=a(1-a)\).

\[\boldsymbol{v}^{\mathsf T}H\boldsymbol{v} =\frac1n\sum_i a_i(1-a_i)(\widetilde{\boldsymbol{x}}_i^{\mathsf T}\boldsymbol{v})^2\ge0.\]

For OLS, replace the weights by 1. Both objectives are convex in the coefficients.

Convexity makes a local optimum globally meaningful

  • Local minimum → global minimum. Zero gradient also certifies a global minimum for a differentiable convex objective.
  • Uniqueness needs more: strict convexity excludes multiple minimizers, if a minimizer exists.
  • Speed needs more: smoothness, learning rate, and conditioning govern convergence.

For an \(L\)-smooth convex objective with a minimizer, gradient descent with step \(1/L\) has an \(O(1/t)\) objective-gap bound. Strong convexity can give geometric convergence.

Next: a useful shape does not itself tell us how to compute the solution.