10. Modeling Functions and Data

A progressive guide to selecting, constructing, interpreting, and evaluating function models for real-world data.

1. Begin with the Modeling Process

A is an equation that approximates how an output changes with an input. The input may represent time, distance, temperature, or another explanatory quantity; the output is the quantity being measured or predicted.

A useful modeling process is:

  1. Identify the variables and their units.

  2. Graph the observations in a .

  3. Look for shape, symmetry, repetition, growth, decay, or breaks.

  4. Select a function family that matches both the pattern and the context.

  5. Estimate the parameters.

  6. Compare predictions with observations.

  7. Interpret the parameters, domain, and limitations.

A model is an approximation rather than a claim that real data follow an equation perfectly. The best model depends on the context, the interval being studied, and the purpose of the prediction.

Takeaway: Begin with variables, units, a graph, and context before choosing an equation.

2. Choose a Function Family

The graph suggests an initial function family, but appearance alone is not enough. Check whether the model has meaningful units, a reasonable domain, and plausible predictions.

  • A nearly constant rate of change suggests a linear model, such as f(x)=mx+bf(x)=mx+b.

  • One or more bends or turning points may suggest a .

  • Constant percentage growth or decay suggests an , such as f(x)=abxf(x)=ab^x.

  • Rapid change that gradually levels off may suggest a .

  • Repeated cycles suggest a .

  • A reciprocal relationship or a denominator involving the input may suggest a .

  • Different rules on different intervals suggest a .

The choice should reflect the mechanism as well as the visual pattern. For example, a polynomial may fit a short interval of data, while an may better represent sustained percentage growth.

Takeaway: Match the function family to the data's shape, mechanism, domain, and units.

3. Interpret Polynomial Behavior

A has the form

f(x)=anxn+an−1xn−1+⋯+a1x+a0.f(x)=a_nx^n+a_{n-1}x^{n-1}+\cdots+a_1x+a_0.

The degree is the highest power of xx. The leading coefficient helps determine end behavior, zeros identify inputs where the output is zero, and turning points can represent local maxima or minima. The constant term a0a_0 is the predicted output at x=0x=0, provided 00 is in the domain.

A quadratic model is useful when data rise and then fall, or fall and then rise:

f(x)=ax2+bx+c.f(x)=ax^2+bx+c.

For a thrown ball, the model

h(t)=−16t2+48t+5h(t)=-16t^2+48t+5

uses time tt in seconds and height h(t)h(t) in feet. The negative leading coefficient indicates that the graph opens downward. The vertex estimates the maximum height, and solving h(t)=0h(t)=0 estimates when the ball reaches the ground.

Higher-degree polynomials can represent more complicated patterns, but excessive degree may create unwanted oscillations and poor predictions outside the observed interval.

Takeaway: Use polynomial features to interpret turning points and zeros, but avoid unnecessary complexity.

4. Model Growth, Decay, and Leveling Off

An can be written as

f(x)=abx.f(x)=ab^x.

Here, aa is the initial value and bb is the multiplicative factor for each one-unit increase in xx. If b>1b>1, the model represents growth; if 0<b<10<b<1, it represents decay. The percentage change per unit is (b−1)×100%(b-1)\times100\%.

For example,

P(t)=1200(1.04)tP(t)=1200(1.04)^t

starts at P(0)=1200P(0)=1200 and represents approximately 4%4\% growth per time unit.

A can be written as

f(x)=a+bln⁡(x−h),f(x)=a+b\ln(x-h),

with domain x>hx>h. It describes rapid initial change followed by slower change. In

L(x)=60+12ln⁡(x),L(x)=60+12\ln(x),

the input must satisfy x>0x>0, and multiplying xx by the same factor adds a constant amount to the output.

Exponential behavior can be diagnosed by transforming y=abxy=ab^x into

ln⁡(y)=ln⁡(a)+xln⁡(b).\ln(y)=\ln(a)+x\ln(b).

A plot of ln⁡(y)\ln(y) against xx should be approximately linear when an exponential relationship is plausible. Predictions should then be converted back to the original units.

Takeaway: Exponential models track multiplicative change; logarithmic models track rapid change that gradually levels off.

5. Represent Repeating Phenomena

A can be written as

f(t)=Asin⁡(Bt−C)+Df(t)=A\sin(Bt-C)+D

or with cosine. Its parameters describe the cycle:

  • Amplitude: ∣A∣\lvert A\rvert, the distance from the midline to a maximum or minimum.

  • Period: T=2π∣B∣T=\frac{2\pi}{\lvert B\rvert} when angles are measured in radians.

  • Phase shift: CB\frac{C}{B}, with direction determined by the expression inside the function.

  • Midline: y=Dy=D, the average of the maximum and minimum values.

  • Maximum: D+∣A∣D+\lvert A\rvert.

  • Minimum: D−∣A∣D-\lvert A\rvert.

Suppose a temperature ranges from 10∘C10^\circ\text{C} to 26∘C26^\circ\text{C} over a 12-month cycle and reaches its maximum in month 7. Then

D=10+262=18,∣A∣=26−102=8,D=\frac{10+26}{2}=18,\qquad \lvert A\rvert=\frac{26-10}{2}=8,

and

B=2π12=π6.B=\frac{2\pi}{12}=\frac{\pi}{6}.

One possible model is

T(t)=18+8cos⁡(π6(t−7)).T(t)=18+8\cos\left(\frac{\pi}{6}(t-7)\right).

It predicts a maximum of 26∘C26^\circ\text{C} at t=7t=7, a minimum of 10∘C10^\circ\text{C} six months later, and a repeating period of 12 months.

Takeaway: Estimate the maximum, minimum, timing, and period before selecting the trigonometric form.

6. Handle Restrictions and Changing Rules

A is written as

f(x)=p(x)q(x),q(x)≠0.f(x)=\frac{p(x)}{q(x)},\qquad q(x)\ne0.

The denominator identifies excluded input values. These values may produce vertical asymptotes or holes, but a mathematical singularity should not automatically be interpreted as a physical outcome; real systems may change behavior before it is reached.

A reciprocal model is

f(x)=kx.f(x)=\frac{k}{x}.

For a fixed distance dd, travel time may vary inversely with speed vv:

t(v)=dv.t(v)=\frac{d}{v}.

A uses different formulas on different intervals. For example, a delivery cost can be modeled by

C(w)={8,0<w≤2,8+2(w−2),w>2,C(w)= \begin{cases} 8, & 0<w\le2,\\ 8+2(w-2), & w>2, \end{cases}

where ww is package weight in pounds and C(w)C(w) is cost in dollars. The first two pounds cost a flat $8\$8, and each additional pound costs $2\$2.

When constructing a piecewise model, verify that every relevant input is covered, intervals do not overlap unintentionally, boundary values are assigned correctly, and any jump or corner has a contextual explanation.

Takeaway: Restrictions and boundaries are part of the model, not details to check afterward.

7. Estimate Parameters and Quantify Error

Parameters can come from known features, selected points, or regression. For example, two points determine a linear model, while maximum, minimum, period, and timing can determine a .

For regression, the for observation ii is

ei=yi−y^i,e_i=y_i-\hat y_i,

where yiy_i is observed and y^i\hat y_i is predicted. Least-squares fitting minimizes the sum of squared residuals:

SSE=∑i=1n(yi−y^i)2.\text{SSE}=\sum_{i=1}^{n}(y_i-\hat y_i)^2.

Squaring prevents positive and negative errors from canceling and gives greater weight to large errors. Other useful measures are

RMSE=1n∑i=1n(yi−y^i)2\text{RMSE}=\sqrt{\frac{1}{n}\sum_{i=1}^{n}(y_i-\hat y_i)^2}

and

MAE=1n∑i=1n∣yi−y^i∣.\text{MAE}=\frac{1}{n}\sum_{i=1}^{n}\lvert y_i-\hat y_i\rvert.

RMSE is expressed in the response units, while MAE is the average absolute error in those same units. Software can calculate parameters, but it cannot decide whether the function family is appropriate.

Takeaway: Regression estimates parameters; interpretation and validation determine whether the resulting model is defensible.

8. Validate Fit and Limit Predictions

Graph the observations and model on the same axes, then inspect the residuals. A useful plot generally has points scattered around zero without a clear curve, trend, funnel shape, or isolated extreme value.

Common patterns provide diagnostic information:

  • A curved pattern suggests that a linear model may be missing nonlinear structure.

  • A funnel shape suggests that the amount of error changes with the size of the prediction.

  • A trend over time may indicate dependence or a missing time-related variable.

  • A large isolated may indicate an outlier, measurement problem, or unusual event.

The R2R^2 describes the proportion of variation accounted for by a model when that measure is appropriate. However, a high R2R^2 does not guarantee a suitable function family, constant variability, patternless residuals, or reliable predictions outside the observed range.

stays within the range of observed inputs and is usually safer. goes beyond that range and can be unreliable, especially for polynomials, exponentials, and logarithms. A model should state the interval and conditions in which it is intended to be used.

Takeaway: Evaluate visual fit, residuals, error measures, domain, and intended prediction range together.

9. Select the Most Defensible Model

When several models seem plausible, compare them systematically:

  1. Context: Does the model reflect the situation or mechanism?

  2. Domain and range: Are all predicted inputs and outputs meaningful?

  3. Shape: Does the graph match the observed behavior?

  4. Parameters: Can the coefficients be interpreted with the correct units?

  5. Residuals: Are the errors small and reasonably patternless?

  6. Complexity: Is the model no more complicated than necessary?

  7. Prediction interval: Is the intended use or ?

The model that passes closest to existing points is not automatically the best model. A is preferable for a genuinely repeating cycle even if a high-degree polynomial can fit the same recorded points. Likewise, an may better represent long-term percentage growth than a polynomial that fits a short interval.

A defensible model combines mathematical fit with contextual meaning, appropriate restrictions, interpretable parameters, and honest limits on prediction.

Final takeaway: Choose the simplest model that explains the important pattern and supports the intended prediction without violating the context.