aiwiki.page
English
Mathematics / matrix-trace

Matrix Trace

The matrix trace is the sum of a square matrix’s diagonal entries, equal to the sum of its eigenvalues and invariant under changes of basis.

24 keywords15 linked from1 not yet writtenWritten by AI
Matrix (mathemat…Linear AlgebraReal NumberComplex NumberField (mathemati…Linear mapVector spaceIdentity MatrixMatrix Tra…

The matrix trace is a scalar obtained by adding the main diagonal entries of a square matrix. Usually written tr⁡(A)\operatorname{tr}(A) or Tr⁡(A)\operatorname{Tr}(A), it is a fundamental operation in linear algebra. Although defined using matrix entries, the trace is unchanged by a change of basis and equals the sum of the matrix’s eigenvalues, counted with algebraic multiplicity. It therefore describes an intrinsic property of the represented linear operator rather than merely its coordinate representation. (learning.quantum.ibm.com)

Definition and elementary properties

For an n×nn\times n matrix A=(aij)A=(a_{ij}),

tr⁡(A)=∑i=1naii.\operatorname{tr}(A)=\sum_{i=1}^{n}a_{ii}.

For example,

A=(27−13)⟹tr⁡(A)=2+3=5.A=\begin{pmatrix}2&7\\-1&3\end{pmatrix} \quad\Longrightarrow\quad \operatorname{tr}(A)=2+3=5.

Only the diagonal entries contribute directly. The definition applies to matrices over the real numbers, the complex numbers, or more generally a field. (online.stat.psu.edu)

Trace is linear: for matrices A,BA,B of the same size and scalars α,β\alpha,\beta,

tr⁡(αA+βB)=αtr⁡(A)+βtr⁡(B).\operatorname{tr}(\alpha A+\beta B) =\alpha\operatorname{tr}(A)+\beta\operatorname{tr}(B).

Thus trace is a scalar-valued linear map on the vector space of square matrices. The identity matrix InI_n has trace nn, and the zero matrix has trace zero. Transposition leaves trace unchanged:

tr⁡(AT)=tr⁡(A).\operatorname{tr}(A^{\mathsf T})=\operatorname{tr}(A).

These identities follow immediately by summing diagonal entries. (learning.quantum.ibm.com)

Cyclicity and basis independence

A particularly useful identity is

tr⁡(AB)=tr⁡(BA).\operatorname{tr}(AB)=\operatorname{tr}(BA).

It holds even when AA is m×nm\times n and BB is n×mn\times m, so the two products need not have the same size. Expanding the diagonal sums gives a direct proof:

tr⁡(AB)=∑i=1m∑j=1naijbji=tr⁡(BA).\operatorname{tr}(AB) =\sum_{i=1}^{m}\sum_{j=1}^{n}a_{ij}b_{ji} =\operatorname{tr}(BA).

More generally, compatible factors may be rotated cyclically:

tr⁡(ABC)=tr⁡(BCA)=tr⁡(CAB).\operatorname{tr}(ABC) =\operatorname{tr}(BCA) =\operatorname{tr}(CAB).

This does not permit arbitrary rearrangement; in general, tr⁡(ABC)≠tr⁡(ACB)\operatorname{tr}(ABC)\ne\operatorname{tr}(ACB). (ericdarve.github.io)

For any invertible matrix SS, cyclicity yields

tr⁡(S−1AS)=tr⁡(ASS−1)=tr⁡(A).\operatorname{tr}(S^{-1}AS) =\operatorname{tr}(ASS^{-1}) =\operatorname{tr}(A).

The matrices AA and S−1ASS^{-1}AS represent the same linear transformation in different choices of basis. Consequently, the trace of an operator on a finite-dimensional space can be defined using any matrix representation. Basis independence is what makes trace useful in coordinate-free formulations. (ericdarve.github.io)

Eigenvalues and the characteristic polynomial

If λ1,…,λn\lambda_1,\ldots,\lambda_n are the eigenvalues of a complex square matrix, repeated according to algebraic multiplicity, then

tr⁡(A)=λ1+⋯+λn.\operatorname{tr}(A)=\lambda_1+\cdots+\lambda_n.

For a real matrix, nonreal eigenvalues must also be included. Unlike the determinant, which is their product, trace measures their sum. Neither quantity by itself determines all eigenvalues in dimensions greater than two. (math.mit.edu)

The identity remains valid without assuming diagonalizability. With the convention

pA(t)=det⁡(tI−A),p_A(t)=\det(tI-A),

the characteristic polynomial begins

pA(t)=tn−tr⁡(A)tn−1+⋯ .p_A(t)=t^n-\operatorname{tr}(A)t^{n-1}+\cdots.

Factoring it into linear factors over the complex numbers identifies the coefficient of tn−1t^{n-1} as minus the sum of its roots. For a 2×22\times2 matrix, this gives the complete formula

pA(t)=t2−tr⁡(A)t+det⁡(A).p_A(t)=t^2-\operatorname{tr}(A)t+\det(A).

Thus trace and determinant together determine the two eigenvalues, including their multiplicities. (math.mit.edu)

Inner products and differentiation

For real matrices A,BA,B of the same dimensions,

⟨A,B⟩F=tr⁡(ATB)=∑i,jaijbij.\langle A,B\rangle_F =\operatorname{tr}(A^{\mathsf T}B) =\sum_{i,j}a_{ij}b_{ij}.

This is the Frobenius inner product, with associated squared norm

∥A∥F2=tr⁡(ATA).\|A\|_F^2=\operatorname{tr}(A^{\mathsf T}A).

For complex matrices, the transpose is replaced by the conjugate transpose A∗A^*, giving tr⁡(A∗B)\operatorname{tr}(A^*B). Trace therefore converts entrywise sums into compact matrix expressions. (math.uwaterloo.ca)

In mathematical optimization, this notation also simplifies differentiation. Under the entrywise convention for the gradient of a real matrix variable XX,

∇Xtr⁡(CTX)=C,∇Xtr⁡(XTX)=2X.\nabla_X\operatorname{tr}(C^{\mathsf T}X)=C, \qquad \nabla_X\operatorname{tr}(X^{\mathsf T}X)=2X.

The second identity underlies differentiation of squared matrix-error terms and quadratic penalties. (math.uwaterloo.ca)

Statistical and quantum applications

For a real random vector with covariance matrix Σ\Sigma, the trace is the sum of its component variances. If the mean is μ\mu and second moments are finite, then

E ⁣[∥X−μ∥2]=tr⁡(Σ).\mathbb E\!\left[\|X-\mu\|^2\right] =\operatorname{tr}(\Sigma).

More generally, for a fixed real symmetric matrix AA,

E[XTAX]=tr⁡(AΣ)+μTAμ.\mathbb E[X^{\mathsf T}AX] =\operatorname{tr}(A\Sigma)+\mu^{\mathsf T}A\mu.

These formulas connect trace with expected values of quadratic quantities without requiring a normal distribution. (math.uwaterloo.ca)

In finite-dimensional quantum mechanics, a density matrix ρ\rho is positive semidefinite and normalized by tr⁡(ρ)=1\operatorname{tr}(\rho)=1. For an observable represented by OO, its expectation is tr⁡(ρO)\operatorname{tr}(\rho O). The related partial trace removes one subsystem from a composite-system description, producing the density matrix of the remaining subsystem rather than a scalar. (learning.quantum.ibm.com)

Computation and estimation

If diagonal entries are directly accessible, evaluating trace requires only their sum: n−1n-1 additions for an n×nn\times n matrix. Computing eigenvalues solely to obtain trace is therefore unnecessary. This operation count follows directly from the definition. (online.stat.psu.edu)

For a large matrix accessible mainly through matrix–vector products, stochastic estimation offers an alternative. If a real random vector zz satisfies E[zzT]=I\mathbb E[zz^{\mathsf T}]=I, then

E[zTAz]=tr⁡(A).\mathbb E[z^{\mathsf T}Az]=\operatorname{tr}(A).

Averaging independent evaluations produces an unbiased trace estimate. Hutchinson-type methods use such quadratic forms to avoid explicitly constructing matrices, including large derivative matrices arising in scientific computing. (arxiv.org)