6 Data Fitting and Empirical Models
A practical guide to building, evaluating, comparing, and responsibly interpreting empirical models from observed data.
From measurements to models
turns measurements into an estimated relationship that can summarize patterns and generate predictions. A general model can be written as
where is an explanatory variable, is a response variable, is a chosen functional form, contains unknown parameters, and represents measurement error, random variation, and effects not included in the model.
An is based primarily on observed data. It may be useful for prediction without providing a complete explanation of the mechanism behind the relationship. After estimating the parameters, the fitted model produces values such as . For example,
predicts distance from time . The estimated intercept is , and the estimated slope is .
Takeaway: A fitted model is an approximation that combines a functional form with parameter estimates and a stated source of unexplained variation.
Inspecting the pattern
Before fitting a formula, inspect a of the paired observations . Look for the overall direction of association, approximate linearity, curvature, clusters, outliers, changing spread, gaps, and a restricted range of values.
A shows an observed relationship, not proof that changing causes to change. The pattern could arise from causation, variables, selection effects, or coincidence.
Begin with the simplest function that appears capable of describing the main structure. Excessively complicated functions can treat random fluctuations as real patterns, which reduces interpretability and may worsen predictions for new data.
Takeaway: Visual inspection guides model choice and can expose structure or limitations before numerical fitting begins.
Fitting a straight-line relationship
represents the fitted relationship with a line:
The intercept is the predicted value of when , provided that is meaningful and relevant to the data. The slope is the predicted change in for a one-unit increase in .
For example, if
then an increase of one unit in is associated with an estimated increase of units in the response. This interpretation describes the fitted association; it does not, by itself, establish causation.
The least-squares method chooses the line that minimizes
For , the coefficients can be computed with
Because deviations are squared, large errors receive more weight than small errors. An unusual observation can therefore substantially change the fitted line.
Takeaway: The slope describes the fitted rate of change, while selects coefficients by minimizing squared discrepancies.
Using residuals as evidence
For observation , the is
A positive means that the observed value is above the fitted value, so the model underpredicts. A negative means that the model overpredicts.
A should generally look like a roughly random cloud centered around zero when the model is reasonable. Specific patterns provide diagnostic clues:
A curved pattern suggests that a straight-line form may be inadequate.
A funnel shape suggests changing variability, also called nonconstant variance or heteroscedasticity.
Runs or trends may indicate dependence, time-related change, or a changing process.
An isolated large may indicate an outlier, a recording problem, or an unusual condition.
Clusters may indicate multiple populations or an omitted explanatory variable.
analysis is often more informative than relying on one numerical statistic because a graph can reveal several forms of model failure at once.
Takeaway: A good fit requires residuals without substantial systematic structure, not merely a visually plausible fitted curve.
Measuring fit responsibly
The sum of squares is
A standard deviation summarizes the typical size of unexplained errors in the same units as the response, which makes it useful for judging practical prediction accuracy.
The is commonly written as
In a standard regression setting, represents the fraction of observed variation in the response accounted for by the fitted model. However, a high does not guarantee a suitable model. Curved patterns, biased predictions, influential observations, or poor behavior outside the observed range can remain.
Takeaway: Evaluate fit using error size and structure together; never treat as sufficient validation by itself.
Beyond a straight line
A linear form is not appropriate for every relationship. A polynomial model can describe curvature over a limited range:
An exponential model can describe growth or decay:
Taking logarithms gives
which can sometimes make linear regression applicable to transformed data. A power model has the form
and transforming both variables gives
Transformations may make a relationship easier to fit, but they also change the interpretation of errors and predictions. Choose a form using the context, , behavior, and purpose of the analysis rather than numerical fit alone.
Takeaway: A transformed or nonlinear model is useful only when its structure and interpretation are appropriate for the problem.
Comparing candidate models
When several candidate models fit the same data, compare them using both statistical evidence and practical usefulness. Important questions include:
Do residuals show less curvature or changing variance for one model?
Are typical prediction errors small enough for the application?
Does the model explain the data without unnecessary terms?
Can its parameters and predictions be interpreted clearly?
Does it predict new observations well?
Does its behavior make sense according to subject-matter constraints?
Adding parameters often lowers training error, but it can produce : the model follows random fluctuations in the available data and performs poorly on new data. A more complex model is justified only when the added structure is supported and improves meaningful predictions.
When possible, use a training set for fitting and a validation or test set for evaluation. For small data sets, cross-validation or another resampling method can provide a more reliable comparison than in-sample fit alone.
Takeaway: Prefer the simplest model that captures meaningful structure and performs well on data not used to fit it.
Knowing when predictions are limited
A fitted model is an approximation whose reliability depends on the conditions represented by the data. Important limitations include:
Poor measurement quality or inconsistent definitions can produce misleading estimates.
A nonrepresentative sample may not describe the broader population.
can make an apparent association difficult to interpret.
beyond the observed range can be highly unreliable.
A relationship may change across time, locations, or other conditions.
Outliers and influential observations can strongly affect parameter estimates.
Regression assumptions, including approximately independent residuals centered around zero with reasonably stable variance, should be investigated rather than assumed.
Association alone does not establish causation.
A responsible interpretation states the population, units, data range, uncertainty, and intended use of the model. Predictions should be treated as supported only to the extent that the data and assumptions justify them.
Takeaway: Good modeling includes explicit limits on what the fitted relationship can support.