04 Gradients and Optimization

A guided introduction to directional derivatives, gradients, tangent and normal geometry, unconstrained extrema, and Lagrange multipliers for constrained optimization.

Directional Change in a Chosen Direction

A scalar-valued function assigns a number to each point. Its partial derivatives describe change in coordinate directions, while a describes change along any chosen direction.

If u\mathbf{u} is a , the of ff at a\mathbf{a} is

Duf(a)=lim⁡h→0f(a+hu)−f(a)h.D_{\mathbf{u}}f(\mathbf{a})=\lim_{h\to 0}\frac{f(\mathbf{a}+h\mathbf{u})-f(\mathbf{a})}{h}.

For a differentiable function, compute it using

Duf=∇f⋅u.D_{\mathbf{u}}f=\nabla f\cdot\mathbf{u}.

The direction must have length 11. If the problem gives a nonzero vector v\mathbf{v}, normalize it first:

u=v∥v∥.\mathbf{u}=\frac{\mathbf{v}}{\|\mathbf{v}\|}.

For example, let f(x,y)=x2+xyf(x,y)=x^2+xy, evaluate the rate of change at (1,2)(1,2), and use the direction of v=⟨3,4⟩\mathbf{v}=\langle 3,4\rangle. The is ∇f(x,y)=⟨2x+y,x⟩\nabla f(x,y)=\langle 2x+y,x\rangle, so ∇f(1,2)=⟨4,1⟩\nabla f(1,2)=\langle 4,1\rangle. Since ∥v∥=5\|\mathbf{v}\|=5, the unit direction is u=⟨35,45⟩\mathbf{u}=\left\langle\frac{3}{5},\frac{4}{5}\right\rangle. Therefore,

Duf(1,2)=⟨4,1⟩⋅⟨35,45⟩=165.D_{\mathbf{u}}f(1,2)=\langle 4,1\rangle\cdot\left\langle\frac{3}{5},\frac{4}{5}\right\rangle=\frac{16}{5}.

The function increases at a rate of 165\frac{16}{5} units per unit distance in that direction.

Takeaway: Find the , normalize the given direction, and take their dot product.

The as a Rate and Direction

For a differentiable function of three variables,

∇f=⟨∂f∂x,∂f∂y,∂f∂z⟩.\nabla f=\left\langle\frac{\partial f}{\partial x},\frac{\partial f}{\partial y},\frac{\partial f}{\partial z}\right\rangle.

For two variables, ∇f=⟨fx,fy⟩\nabla f=\langle f_x,f_y\rangle. The dot-product formula connects the to every :

Duf=∇f⋅u.D_{\mathbf{u}}f=\nabla f\cdot\mathbf{u}.

Because u\mathbf{u} has length 11, the Cauchy–Schwarz inequality gives

Duf≤∥∇f∥.D_{\mathbf{u}}f\leq\|\nabla f\|.

Thus:

  • The maximum rate of increase is ∥∇f∥\|\nabla f\|, occurring in the direction of ∇f\nabla f.

  • The maximum rate of decrease is −∥∇f∥-\|\nabla f\|, occurring in the direction of −∇f-\nabla f.

  • A zero occurs in directions perpendicular to ∇f\nabla f.

If ∇f(a)=0\nabla f(\mathbf{a})=\mathbf{0}, the first-order information does not select a preferred direction. The point could still be a local maximum, local minimum, or saddle point, so further analysis is required.

Takeaway: The simultaneously encodes all first-order directional rates and identifies the direction of steepest ascent.

Level Sets, Normals, and Tangent Planes

A level curve of a function of two variables is defined by f(x,y)=cf(x,y)=c. If r(t)=⟨x(t),y(t)⟩\mathbf{r}(t)=\langle x(t),y(t)\rangle traces that curve, then f(r(t))=cf(\mathbf{r}(t))=c. Differentiating gives

∇f(r(t))⋅r′(t)=0.\nabla f(\mathbf{r}(t))\cdot\mathbf{r}'(t)=0.

Therefore, the tangent vector r′(t)\mathbf{r}'(t) is perpendicular to the . A tangent direction t\mathbf{t} satisfies ∇f⋅t=0\nabla f\cdot\mathbf{t}=0.

For a surface defined implicitly by F(x,y,z)=cF(x,y,z)=c, the at P=(x0,y0,z0)P=(x_0,y_0,z_0) is a normal vector, provided ∇F(P)≠0\nabla F(P)\neq\mathbf{0}. The tangent plane is

∇F(P)⋅⟨x−x0,y−y0,z−z0⟩=0.\nabla F(P)\cdot\langle x-x_0,y-y_0,z-z_0\rangle=0.

In coordinate form,

Fx(P)(x−x0)+Fy(P)(y−y0)+Fz(P)(z−z0)=0.F_x(P)(x-x_0)+F_y(P)(y-y_0)+F_z(P)(z-z_0)=0.

For a graph z=f(x,y)z=f(x,y), write it as F(x,y,z)=f(x,y)−z=0F(x,y,z)=f(x,y)-z=0. A normal vector is ⟨fx,fy,−1⟩\langle f_x,f_y,-1\rangle, and the tangent-plane equation at (x0,y0,f(x0,y0))(x_0,y_0,f(x_0,y_0)) is

z−f(x0,y0)=fx(x0,y0)(x−x0)+fy(x0,y0)(y−y0).z-f(x_0,y_0)=f_x(x_0,y_0)(x-x_0)+f_y(x_0,y_0)(y-y_0).

This equation is also the first-order linear approximation to the function.

Takeaway: Tangent directions lie along a level set, while the supplies a perpendicular normal direction.

Finding and Classifying Unconstrained Extrema

To locate possible local extrema of a differentiable function of two variables, first solve ∇f=0\nabla f=\mathbf{0}. Also include interior points where a relevant partial derivative does not exist. These candidates are critical points.

For each , compute

D=fxxfyy−(fxy)2.D=f_{xx}f_{yy}-(f_{xy})^2.

The second-derivative test gives:

  • If D>0D>0 and fxx>0f_{xx}>0, the point is a local minimum.

  • If D>0D>0 and fxx<0f_{xx}<0, the point is a local maximum.

  • If D<0D<0, the point is a saddle point.

  • If D=0D=0, the test is inconclusive.

For a restricted domain, local analysis is not enough to find global extrema. Check interior critical points, boundary candidates, and corners when the boundary has corners. Then compare all objective values.

A zero alone does not prove that a point is an extremum. Classification requires the second-derivative test or another appropriate argument.

Takeaway: Critical points provide candidates; classification and boundary analysis determine which candidates are actual extrema.

One-Constraint Optimization with Lagrange Multipliers

A constrained optimization problem asks for the largest or smallest value of an objective function while one or more equations restrict the allowed points. With one constraint, the problem has the form

optimize f(x,y,z)subject tog(x,y,z)=c.\text{optimize }f(x,y,z)\quad\text{subject to}\quad g(x,y,z)=c.

At a constrained extremum, movement is limited to tangent directions on the constraint surface. The derivative of the objective must be zero in every allowed tangent direction. Since ∇g\nabla g is normal to the constraint surface, the objective must be parallel to it. This gives

∇f=λ∇g,g(x,y,z)=c.\nabla f=\lambda\nabla g, \qquad g(x,y,z)=c.

Use the following procedure:

  1. Identify the objective function and the constraint.

  2. Compute ∇f\nabla f and ∇g\nabla g.

  3. Solve ∇f=λ∇g\nabla f=\lambda\nabla g together with g=cg=c.

  4. Evaluate the objective at every candidate point.

  5. Compare the values to identify the maximum and minimum.

For example, optimize f(x,y)=xyf(x,y)=xy subject to x2+y2=1x^2+y^2=1. With g(x,y)=x2+y2g(x,y)=x^2+y^2, the equations are

⟨y,x⟩=λ⟨2x,2y⟩,x2+y2=1.\langle y,x\rangle=\lambda\langle 2x,2y\rangle, \qquad x^2+y^2=1.

They imply x2=y2x^2=y^2. If y=xy=x, the product is 12\frac{1}{2}; if y=−xy=-x, the product is −12-\frac{1}{2}. Therefore,

max⁡f=12\max f=\frac{1}{2}

at (12,12)\left(\frac{1}{\sqrt{2}},\frac{1}{\sqrt{2}}\right) and (−12,−12)\left(-\frac{1}{\sqrt{2}},-\frac{1}{\sqrt{2}}\right), while

min⁡f=−12\min f=-\frac{1}{2}

at (12,−12)\left(\frac{1}{\sqrt{2}},-\frac{1}{\sqrt{2}}\right) and (−12,12)\left(-\frac{1}{\sqrt{2}},\frac{1}{\sqrt{2}}\right).

Takeaway: Lagrange equations locate constrained candidates, but evaluating and comparing the objective values is essential.

Multiple Constraints and Sensitivity

When two constraint equations restrict a function of three variables,

g(x,y,z)=c,h(x,y,z)=d,g(x,y,z)=c, \qquad h(x,y,z)=d,

the allowed set is generally the intersection of two surfaces. A tangent direction is perpendicular to both constraint normals, so the objective must lie in their span:

∇f=λ∇g+μ∇h.\nabla f=\lambda\nabla g+\mu\nabla h.

Solve the complete system

{∇f=λ∇g+μ∇h,g=c,h=d.\begin{cases} \nabla f=\lambda\nabla g+\mu\nabla h,\\ g=c,\\ h=d. \end{cases}

The multipliers λ\lambda and μ\mu are unknowns alongside the coordinates. After finding all candidates, evaluate the objective and compare the resulting values. As with one constraint, the usual method assumes suitable smoothness and that the relevant constraint gradients do not fail to provide the required normal directions.

A multiplier can also describe sensitivity. If the constrained optimum has value M(c)M(c) when the constraint level is cc, then under appropriate differentiability conditions,

dMdc=λ.\frac{dM}{dc}=\lambda.

Thus, a small increase in the constraint level changes the optimized value by approximately λ\lambda times that increase. This interpretation is called the of the constraint.

Takeaway: With multiple constraints, combine the constraint gradients and interpret multipliers as approximate marginal changes in the optimal value.

A Reliable Problem-Solving Checklist

Several checks prevent common errors in and optimization problems:

  • Normalize a direction before using it in a .

  • Include the constraint equation when solving Lagrange equations.

  • Find and evaluate every candidate rather than reporting the first one found.

  • Examine boundaries and corners when the domain is restricted.

  • Do not assume every Lagrange candidate is a maximum; compare objective values or use additional analysis.

  • Check the regularity condition, such as ∇g≠0\nabla g\neq\mathbf{0}, when applying the standard one-constraint theorem.

  • Distinguish tangent vectors, which lie along a constraint, from constraint gradients, which are normal to it.

The central geometric connection is consistent throughout: gradients describe first-order change, gradients are normal to level sets, tangent directions are orthogonal to those normals, and constrained extrema occur when the objective is compatible with the available constraint normals.