11 Mathematical Modeling
A structured guide to selecting, fitting, interpreting, validating, and responsibly applying mathematical models to real-world data.
The
A represents a real situation for a particular purpose; it is not the situation itself. It can be used to estimate a quantity, compare alternatives, describe a trend, or support a decision.
The provides a disciplined workflow:
Define the question: decide what must be estimated, predicted, compared, or optimized.
Identify the variables: specify the input or explanatory variables and the output or response variable.
Collect and examine data: organize observations, graph them, and look for patterns, unusual values, and changes in spread.
Select a model: choose a function whose shape and assumptions fit the context.
Fit the model: estimate parameter values so that predictions agree reasonably with observations.
Interpret and validate: explain parameters in context and inspect errors or residuals.
Use the model appropriately: restrict predictions to a reasonable domain and state limitations.
A useful model must do more than match existing points. Its assumptions should make sense, its variables should have clear meanings and units, and its predictions should address the original question.
Takeaway: Modeling combines mathematical fit with contextual judgment; a close fit alone is not enough.
Selecting a Function Model
Model selection should use both numerical patterns and the mechanism suggested by the context. A function's domain must contain only meaningful inputs, and its range should reflect possible outputs.
Common patterns include:
A constant model, , for an output that remains approximately unchanged.
A , , for approximately constant additive change.
A quadratic model, , for curved change with one possible turning point.
A polynomial model, , for more complicated smooth variation.
An , or , for approximately constant multiplicative or percentage change.
A logarithmic model, , for rapid initial change that slows as the input increases.
An inverse-variation model, , when the product of the variables is approximately constant.
A sinusoidal model, , for repeating behavior.
For values near , the successive differences are approximately constant, supporting a . For values near , the successive ratios are approximately constant:
An with a growth factor near is more appropriate in the second case. A model can be reasonable over a short interval while becoming inappropriate for long-term prediction.
Takeaway: Differences suggest additive change, ratios suggest multiplicative change, and context determines whether the pattern is credible.
Fitting Functions to Data
Fitting means choosing parameter values so that model outputs are close to observed data. A fitted prediction can be written as
where denotes estimated parameters.
For a ,
is the estimated rate of change, measured in output units per input unit. The intercept is the predicted output at , but only when that input is meaningful and relevant to the data range.
chooses the parameters that minimize
For example, suppose electricity use , measured in kilowatt-hours, is modeled from average temperature , measured in degrees Celsius, by
The slope predicts an increase of approximately kilowatt-hours for each one-degree increase in average temperature over the observed range. At ,
The predicted use is approximately kilowatt-hours.
Nonlinear parameters can be estimated numerically. For an exponential relationship , taking natural logarithms gives
If a fitted line for against is , then and . This transformation requires positive response values and can change the weighting of errors, so the result should be checked in the original units.
Takeaway: Fitting produces parameter estimates, but interpreting those estimates requires units, a meaningful domain, and awareness of how the fitting method weights errors.
Residuals and Model Validation
A compares an observed value with the value predicted by the model:
Positive residuals indicate underprediction; negative residuals indicate overprediction. Residuals represent the part of the response not explained by the fitted model.
A plot should generally show points scattered around zero without a systematic pattern. Important checks include:
Randomness: a curve or trend suggests that the model misses structure.
Constant spread: a funnel shape suggests that variability changes across the input range.
Outliers: an unusually large may indicate an unusual observation, measurement problem, or missing feature.
Dependence: runs, trends, or cycles in time-ordered residuals suggest that the model has not captured a time effect or another relationship.
Two common error summaries are the mean absolute error and root mean square error:
and
Both are expressed in the original output units. Mean absolute error gives the average absolute error, while root mean square error penalizes large errors more strongly.
The , denoted by , can summarize explained variation for a particular data set and definition. However, a high does not prove that the model is appropriate, causal, or reliable outside the observed range.
Takeaway: patterns often reveal model problems that a single goodness-of-fit number can hide.
Interpreting Parameters and Assumptions
Parameters must be interpreted with their units, associated variables, and model assumptions. For the
is the predicted value when , and is the multiplicative factor for each one-unit increase in . If , the model represents growth; if , it represents decay. The percentage change per time unit is
For
the initial value is , and the predicted increase is per time unit. After four time units,
The prediction is approximately units, subject to the model's assumptions.
examines how much predictions change when parameters or assumptions change. This matters especially for long-term predictions, because a small parameter change can produce a large difference after repeated growth or decay.
Association between variables does not by itself establish causation. A fitted relationship can support prediction without proving that changing one variable causes a change in another.
Takeaway: A parameter's numerical value is meaningful only when its role, units, interpretation, and uncertainty are made clear.
Using Models Responsibly
estimates a value inside the observed input range and is usually more reliable because the model is used where it was fitted. predicts outside that range and is riskier because the relationship, variability, or mechanism may change.
Before relying on a prediction, ask:
Is the data representative of the situation?
Are the measurements accurate and sufficiently numerous?
Were important variables omitted?
Are assumptions about independence, variability, or growth reasonable?
Is the input domain restricted by physical, practical, or mathematical constraints?
Could an outlier or influential observation control the fitted parameters?
Do residuals show curvature, changing spread, dependence, or cycles?
Does the conclusion claim causation without an appropriate study design?
Consider the fitted advertising-sales relationship
where is advertising spending in hundreds of dollars and is weekly sales in hundreds of dollars. The negative quadratic coefficient indicates diminishing increases and a turning point. For , the vertex input is
The model predicts its maximum at approximately , corresponding to dollars in advertising. The predicted sales are
or approximately dollars in weekly sales. This does not prove that spending exactly dollars causes maximum sales. The conclusion is justified only over the data-supported range and after checking residuals, error measures, influential observations, and the quality of the data.
A responsible conclusion states the prediction, the domain, the assumptions, and the limitations. A model can be useful without being exact when its purpose and conditions of use are explicit.
Takeaway: Prefer when appropriate, treat cautiously, and report both what the model predicts and when that prediction is credible.