For centered data, minimizing reconstruction error is equivalent to maximizing \(\|XW\|_F^2\).
Every orthonormal \(W\) retains at most \(\sum_{j=1}^k\sigma_j^2\); choosing \(W=V_k\) attains that bound.
The discarded squared singular values give reconstruction error: \(\|X-X_k\|_F^2=\sum_{j>k}\sigma_j^2\).
Each PCA coordinate has sample variance \(\sigma_j^2/(n-1)\); choosing \(k\) still depends on the intended use.
Next: revisit the variance interpretation, finish PCA’s geometric interpretation, then implement it in NumPy.
Squared singular values give the variance of each PCA coordinate
The \(j\)th coordinate across observations is \(Z_{:j}=X\boldsymbol{v}_j=\sigma_j\boldsymbol{u}_j\). Since \(X\) is centered, this coordinate also has mean zero.
A curved structure can be low-dimensional without being linear
Would one linear feature reconstruct every point exactly? Could a nonlinear coordinate do so? What does “one-dimensional” mean in each claim?
PCA solves the stated problem; usefulness still needs evidence
Chosen objective: squared error in the selected measurements and units.
Mathematical solution: leading right singular vectors of centered data.
Remaining questions: which features and preprocessing, how many components, and whether the retained variation matters for the intended use.
Next: translate this mathematical contract into arrays and reliable numerical computation—then ask what the observed data justify claiming about new cases.
NumPy and Numerical Computation for PCA
David I. Inouye
Dynamic Content · Last Updated: September 17, 2026
The PCA equations give us a computation to implement
Let \(\widetilde X\in\mathbb R^{n\times d}\) contain the original observations as rows.
P = np.array([[1., 2.], [3., 4.]])Q = np.array([[2., 0.], [1., 2.]])print("P * Q (entrywise):\n", P * Q)print("P @ Q (matrix product):\n", P @ Q)
P * Q (entrywise):
[[2. 0.]
[3. 8.]]
P @ Q (matrix product):
[[ 4. 4.]
[10. 8.]]
For PCA, Z = X @ W maps (n, d) @ (d, k) to (n, k).
The upper-left entries are \(1\cdot2=2\) and \(1\cdot2+2\cdot1=4\). Both results have shape (2, 2). Shape alone cannot tell us which operation we intended.
Centering averages across observations for each feature
For each feature, \(\mu_j=\frac1n\sum_{i=1}^n\widetilde X_{ij}\).
A = np.array([[2., 10.], [4., 14.], [6., 18.]])mu = A.mean(axis=0)X = A - muprint("mu (feature means):", mu)print("X (centered data):\n", X)
mu (feature means): [ 4. 14.]
X (centered data):
[[-2. -4.]
[ 0. 0.]
[ 2. 4.]]
Axis rule: count positions in the shape tuple from zero. Aggregate over the chosen dimension and remove it by default. The resulting shapes are:
T = np.zeros((2, 3, 4))print("T.shape:", T.shape)print("axis=0:", T.mean(axis=0).shape)print("axis=1:", T.mean(axis=1).shape)print("axis=2:", T.mean(axis=2).shape)
keepdims preserves a reduced axis; None inserts a new axis where specified.
(2, 1, 4) broadcasts back against (2, 3, 4).
(2, 4) would align as (1, 2, 4) and fail at the middle position.
Division broadcasts one standard deviation per feature
For positive feature standard deviations \(s_j\), \(H_{ij}=(\widetilde X_{ij}-\mu_j)/s_j\). In matrix notation, \(H=X\operatorname{diag}(1/s_1,\ldots,1/s_d)\).
(3, 2) / (2,) reuses each feature’s scale across observations. The same shape rules apply to elementwise +, -, *, and /. Scaling changes the relative weight of features in PCA’s reconstruction error.
Code can run successfully while centering the wrong objects
A = np.array([[2., 10.], [4., 14.], [6., 18.]])B = A - A.mean(axis=1, keepdims=True)print("B:\n", B)print("B.mean(axis=1):", B.mean(axis=1))print("B.mean(axis=0):", B.mean(axis=0))
What does this code subtract? Which averages become zero? Would this prepare the data for the PCA problem we defined? Repair the expression and explain why.
NumPy returns the right singular vectors as rows of Vt
U, s, Vt = np.linalg.svd( X, full_matrices=False)W = Vt[:k, :].T
full_matrices=False omits extra directions multiplied by zero rows or columns of rectangular \(\Sigma\). This reduced SVD keeps \(m=\min(n,d)\) singular values; selecting the first \(k\) is a separate truncation.
Returned array
Shape
Mathematical role
U
(n, m)
Left singular vectors
s
(m,)
Singular values, largest first
Vt
(m, d)
Rows are right singular vectors transposed
W
(d, k)
First \(k\) right singular vectors as columns
Encoding and reconstruction follow the same equations as before
Xraw stores \(\widetilde X\), the original observations.