2 Matrices and Matrix Operations

A progressive guide to matrix notation, core operations, transposes, inverses, row reduction, and reliable computational methods for solving linear systems.

Reading and representing matrices

A is a rectangular array organized by rows and columns. If it has mm rows and nn columns, its size is m×nm\times n, and an entry is written aija_{ij}, where ii is the row number and jj is the column number.

For example,

A=[2−14035]A=\begin{bmatrix}2&-1&4\\0&3&5\end{bmatrix}

has size 2×32\times3, and its entry a23a_{23} is 55. A square has the same number of rows and columns. A row has one row, while a column has one column.

The of size n×nn\times n is written InI_n. Its main diagonal contains ones and all other entries are zero. The zero , written 00, contains zeros everywhere.

Matrices provide a compact form for linear systems. The equations

2x−y=4,x+3y=72x-y=4,\qquad x+3y=7

can be written as

[2−113][xy]=[47].\begin{bmatrix}2&-1\\1&3\end{bmatrix} \begin{bmatrix}x\\y\end{bmatrix} = \begin{bmatrix}4\\7\end{bmatrix}.

The associated augmented is

[2−14137].\left[\begin{array}{cc|c}2&-1&4\\1&3&7\end{array}\right].

Takeaway: Read indices as row first, column second, and use dimensions to determine which operations are possible.

Addition, equality, and scalar multiplication

Addition and subtraction are defined only for matrices with the same dimensions. The operations are performed entry by entry:

(A+B)ij=aij+bij.(A+B)_{ij}=a_{ij}+b_{ij}.

For example,

[12−34]+[5−120]=[61−14].\begin{bmatrix}1&2\\-3&4\end{bmatrix} + \begin{bmatrix}5&-1\\2&0\end{bmatrix} = \begin{bmatrix}6&1\\-1&4\end{bmatrix}.

A scalar multiple multiplies every entry by the same scalar. Thus,

3[1−204]=[3−6012].3\begin{bmatrix}1&-2\\0&4\end{bmatrix} = \begin{bmatrix}3&-6\\0&12\end{bmatrix}.

The zero is the additive identity because A+0=AA+0=A. equality also depends on both dimensions and corresponding entries: two matrices are equal exactly when they have the same dimensions and every corresponding entry is equal.

These operations should not be confused with . Entrywise addition is possible only for matching dimensions, whereas multiplication uses a row-by-column rule and has a different dimension condition.

Takeaway: Before adding or subtracting, verify equal dimensions; for scalar multiplication, apply the scalar to every entry.

Computing products

For matrices AA of size m×nm\times n and BB of size n×pn\times p, the product ABAB is defined and has size m×pm\times p. The inner dimensions must match.

Each entry is obtained by taking the dot product of a row of AA with a column of BB:

(AB)ij=∑k=1naikbkj.(AB)_{ij}=\sum_{k=1}^{n}a_{ik}b_{kj}.

For

A=[1234],B=[5012],A=\begin{bmatrix}1&2\\3&4\end{bmatrix},\qquad B=\begin{bmatrix}5&0\\1&2\end{bmatrix},

we get

AB=[1(5)+2(1)1(0)+2(2)3(5)+4(1)3(0)+4(2)]=[74198].AB= \begin{bmatrix} 1(5)+2(1)&1(0)+2(2)\\ 3(5)+4(1)&3(0)+4(2) \end{bmatrix} = \begin{bmatrix}7&4\\19&8\end{bmatrix}.

is generally not commutative: AB≠BAAB\ne BA, and BABA may not even be defined. It is associative and distributive when all products involved are defined:

A(BC)=(AB)C,A(B+C)=AB+AC,(A+B)C=AC+BC.A(BC)=(AB)C, \qquad A(B+C)=AB+AC, \qquad (A+B)C=AC+BC.

The acts as a multiplicative identity whenever the dimensions are compatible:

AIn=A,ImA=A.AI_n=A,\qquad I_mA=A.

Multiplying a by a column vector can also be viewed as forming a linear combination of the 's columns. If the columns are a1,…,an\mathbf{a}_1,\ldots,\mathbf{a}_n, then

A[x1x2⋮xn]=x1a1+x2a2+⋯+xnan.A\begin{bmatrix}x_1\\x_2\\\vdots\\x_n\end{bmatrix} =x_1\mathbf{a}_1+x_2\mathbf{a}_2+\cdots+x_n\mathbf{a}_n.

Takeaway: Check compatibility first, then compute each result entry using one row of the first and one column of the second.

Transposes and structural identities

The ATA^T exchanges rows and columns. An m×nm\times n becomes an n×mn\times m , with entries satisfying

(AT)ij=aji.(A^T)_{ij}=a_{ji}.

For example,

[123456]T=[142536].\begin{bmatrix}1&2&3\\4&5&6\end{bmatrix}^T = \begin{bmatrix}1&4\\2&5\\3&6\end{bmatrix}.

Important identities are

(AT)T=A,(A+B)T=AT+BT,(cA)T=cAT.(A^T)^T=A, \qquad (A+B)^T=A^T+B^T, \qquad (cA)^T=cA^T.

For a product, the order reverses:

(AB)T=BTAT.(AB)^T=B^TA^T.

The reversal is required to preserve dimension compatibility. A square satisfying AT=AA^T=A is symmetric. For complex matrices, the analogous conjugate both transposes the and takes the complex conjugate of each entry.

Takeaway: by reflecting entries across the main diagonal, and reverse factor order when transposing a product.

Inverses and solving with matrices

An is a square with an inverse. The inverse A−1A^{-1} satisfies

AA−1=A−1A=I.AA^{-1}=A^{-1}A=I.

If Ax=bA\mathbf{x}=\mathbf{b} and AA is invertible, multiplying by the inverse gives

x=A−1b.\mathbf{x}=A^{-1}\mathbf{b}.

For

A=[abcd],A=\begin{bmatrix}a&b\\c&d\end{bmatrix},

an inverse exists when ad−bc≠0ad-bc\ne0, and then

A−1=1ad−bc[d−b−ca].A^{-1}=\frac{1}{ad-bc} \begin{bmatrix}d&-b\\-c&a\end{bmatrix}.

The quantity ad−bcad-bc is the determinant for this 2×22\times2 . If it is zero, the is singular and has no inverse.

To find an inverse by Gauss–Jordan elimination, augment the with an and row-reduce:

[A∣In]⟶[In∣A−1].[A\mid I_n]\longrightarrow[I_n\mid A^{-1}].

For example,

[12103401]⟶[10−210132−12].\left[\begin{array}{cc|cc}1&2&1&0\\3&4&0&1\end{array}\right] \longrightarrow \left[\begin{array}{cc|cc}1&0&-2&1\\0&1&\frac{3}{2}&-\frac{1}{2}\end{array}\right].

Therefore,

[1234]−1=[−2132−12].\begin{bmatrix}1&2\\3&4\end{bmatrix}^{-1} = \begin{bmatrix}-2&1\\\frac{3}{2}&-\frac{1}{2}\end{bmatrix}.

Useful identities include

(AB)−1=B−1A−1,(AT)−1=(A−1)T,(A−1)−1=A.(AB)^{-1}=B^{-1}A^{-1}, \qquad (A^T)^{-1}=(A^{-1})^T, \qquad (A^{-1})^{-1}=A.

The order reversal in (AB)−1(AB)^{-1} is essential. Although the inverse gives a theoretical formula for solving systems, numerical computation usually favors a direct solve rather than explicitly forming the inverse.

Takeaway: An inverse exists only for an invertible square ; use row reduction or an appropriate factorization to compute or apply it.

Row operations and elimination

An results from applying one elementary row operation to an . The three operations are:

  1. Interchange two rows: Ri↔RjR_i\leftrightarrow R_j.

  2. Multiply a row by a nonzero scalar: Ri→cRiR_i\to cR_i, where c≠0c\ne0.

  3. Add a multiple of one row to another: Ri→Ri+cRjR_i\to R_i+cR_j, where i≠ji\ne j.

If EE represents one of these operations, then left multiplication applies it to another :

EA=the matrix obtained from A by that row operation.EA=\text{the matrix obtained from }A\text{ by that row operation}.

For example, adding three times row one to row two in a 3×33\times3 uses

E=[100310001].E=\begin{bmatrix}1&0&0\\3&1&0\\0&0&1\end{bmatrix}.

Every is invertible, and its inverse reverses the associated row operation. Thus, a sequence of row operations can be represented as

Ek⋯E2E1A.E_k\cdots E_2E_1A.

This explains why row reduction preserves the solution set of a linear system: each elementary row operation produces an equivalent system.

transforms an augmented [A∣b][A\mid\mathbf{b}] into row-echelon form, after which back-substitution finds the solution. Gauss–Jordan elimination continues to reduced row-echelon form. For a unique solution, the final augmented has the form

[I∣x].[I\mid\mathbf{x}].

For example,

[1253411]→R2→R2−3R1[1250−2−4].\left[\begin{array}{cc|c}1&2&5\\3&4&11\end{array}\right] \xrightarrow{R_2\to R_2-3R_1} \left[\begin{array}{cc|c}1&2&5\\0&-2&-4\end{array}\right].

The second row gives y=2y=2, and substituting into the first gives x=1x=1.

Takeaway: Elementary matrices encode row operations, while uses those operations to solve systems without changing their solutions.

Reliable numerical computation

Exact arithmetic permits any nonzero pivot, but floating-point arithmetic can magnify rounding errors. Partial pivoting reduces this risk by interchanging the current row with a lower row whose entry in the pivot column has the largest absolute value. This also helps avoid division by a very small pivot.

A can be mathematically invertible but numerically ill-conditioned. In that situation, small errors in the entries or right-hand side can produce relatively large changes in the computed solution. The condition number measures one aspect of this sensitivity.

When several systems share the same coefficient , is efficient. Write

PA=LU,PA=LU,

where PP is a permutation , LL is lower triangular, and UU is upper triangular. To solve Ax=bA\mathbf{x}=\mathbf{b}, solve two triangular systems:

Ly=Pb,Ux=y.L\mathbf{y}=P\mathbf{b}, \qquad U\mathbf{x}=\mathbf{y}.

The factorization requires an initial setup, but each additional right-hand side can then be handled efficiently.

For numerical work, solve a system directly rather than computing A−1bA^{-1}\mathbf{b} explicitly whenever possible. Direct solvers generally avoid unnecessary work and reduce numerical error. In NumPy, numpy.linalg.solve(A, b) is the standard choice for a square system, .T represents a , and @ performs .

Common checks:

  • Do not add or subtract matrices with different dimensions.

  • Do not multiply entries position by position when is required.

  • Do not assume AB=BAAB=BA.

  • Do not replace (AB)T(AB)^T with ATBTA^TB^T.

  • Do not invert a nonsquare or singular .

  • Do not divide by a zero or extremely small pivot without considering a row interchange.

Takeaway: Choose algorithms based on numerical reliability: pivot carefully, reuse LU factors for repeated systems, and prefer direct solves to explicit inverses.