6 Data Fitting and Empirical Models

A practical guide to building, evaluating, comparing, and responsibly interpreting empirical models from observed data.

From measurements to models

turns measurements into an estimated relationship that can summarize patterns and generate predictions. A general model can be written as

y=f(x;β)+ε,y=f(x;\boldsymbol{\beta})+\varepsilon,

where xx is an explanatory variable, yy is a response variable, ff is a chosen functional form, β\boldsymbol{\beta} contains unknown parameters, and ε\varepsilon represents measurement error, random variation, and effects not included in the model.

An is based primarily on observed data. It may be useful for prediction without providing a complete explanation of the mechanism behind the relationship. After estimating the parameters, the fitted model produces values such as y^\hat{y}. For example,

d^=12.4+4.8t\hat{d}=12.4+4.8t

predicts distance d^\hat{d} from time tt. The estimated intercept is 12.412.4, and the estimated slope is 4.84.8.

Takeaway: A fitted model is an approximation that combines a functional form with parameter estimates and a stated source of unexplained variation.

Inspecting the pattern

Before fitting a formula, inspect a of the paired observations (xi,yi)(x_i,y_i). Look for the overall direction of association, approximate linearity, curvature, clusters, outliers, changing spread, gaps, and a restricted range of values.

A shows an observed relationship, not proof that changing xx causes yy to change. The pattern could arise from causation, variables, selection effects, or coincidence.

Begin with the simplest function that appears capable of describing the main structure. Excessively complicated functions can treat random fluctuations as real patterns, which reduces interpretability and may worsen predictions for new data.

Takeaway: Visual inspection guides model choice and can expose structure or limitations before numerical fitting begins.

Fitting a straight-line relationship

represents the fitted relationship with a line:

y^=b0+b1x.\hat{y}=b_0+b_1x.

The intercept b0b_0 is the predicted value of yy when x=0x=0, provided that x=0x=0 is meaningful and relevant to the data. The slope b1b_1 is the predicted change in yy for a one-unit increase in xx.

For example, if

y^=18.2+2.7x,\hat{y}=18.2+2.7x,

then an increase of one unit in xx is associated with an estimated increase of 2.72.7 units in the response. This interpretation describes the fitted association; it does not, by itself, establish causation.

The least-squares method chooses the line that minimizes

SSE⁡=∑i=1n(yi−y^i)2.\operatorname{SSE}=\sum_{i=1}^{n}(y_i-\hat{y}_i)^2.

For , the coefficients can be computed with

b1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2,b0=yˉ−b1xˉ.b_1=\frac{\sum (x_i-\bar{x})(y_i-\bar{y})}{\sum (x_i-\bar{x})^2}, \qquad b_0=\bar{y}-b_1\bar{x}.

Because deviations are squared, large errors receive more weight than small errors. An unusual observation can therefore substantially change the fitted line.

Takeaway: The slope describes the fitted rate of change, while selects coefficients by minimizing squared discrepancies.

Using residuals as evidence

For observation ii, the is

ei=yi−y^i.e_i=y_i-\hat{y}_i.

A positive means that the observed value is above the fitted value, so the model underpredicts. A negative means that the model overpredicts.

A should generally look like a roughly random cloud centered around zero when the model is reasonable. Specific patterns provide diagnostic clues:

  • A curved pattern suggests that a straight-line form may be inadequate.

  • A funnel shape suggests changing variability, also called nonconstant variance or heteroscedasticity.

  • Runs or trends may indicate dependence, time-related change, or a changing process.

  • An isolated large may indicate an outlier, a recording problem, or an unusual condition.

  • Clusters may indicate multiple populations or an omitted explanatory variable.

analysis is often more informative than relying on one numerical statistic because a graph can reveal several forms of model failure at once.

Takeaway: A good fit requires residuals without substantial systematic structure, not merely a visually plausible fitted curve.

Measuring fit responsibly

The sum of squares is

SSE⁡=∑ei2.\operatorname{SSE}=\sum e_i^2.

A standard deviation summarizes the typical size of unexplained errors in the same units as the response, which makes it useful for judging practical prediction accuracy.

The is commonly written as

R2=1−SSE⁡SST⁡,SST⁡=∑(yi−yˉ)2.R^2=1-\frac{\operatorname{SSE}}{\operatorname{SST}}, \qquad \operatorname{SST}=\sum (y_i-\bar{y})^2.

In a standard regression setting, R2R^2 represents the fraction of observed variation in the response accounted for by the fitted model. However, a high R2R^2 does not guarantee a suitable model. Curved patterns, biased predictions, influential observations, or poor behavior outside the observed range can remain.

Takeaway: Evaluate fit using error size and structure together; never treat R2R^2 as sufficient validation by itself.

Beyond a straight line

A linear form is not appropriate for every relationship. A polynomial model can describe curvature over a limited range:

y^=b0+b1x+b2x2.\hat{y}=b_0+b_1x+b_2x^2.

An exponential model can describe growth or decay:

y^=aekx.\hat{y}=ae^{kx}.

Taking logarithms gives

ln⁡y=ln⁡a+kx,\ln y=\ln a+kx,

which can sometimes make linear regression applicable to transformed data. A power model has the form

y^=axk,\hat{y}=ax^k,

and transforming both variables gives

ln⁡y=ln⁡a+kln⁡x.\ln y=\ln a+k\ln x.

Transformations may make a relationship easier to fit, but they also change the interpretation of errors and predictions. Choose a form using the context, , behavior, and purpose of the analysis rather than numerical fit alone.

Takeaway: A transformed or nonlinear model is useful only when its structure and interpretation are appropriate for the problem.

Comparing candidate models

When several candidate models fit the same data, compare them using both statistical evidence and practical usefulness. Important questions include:

  1. Do residuals show less curvature or changing variance for one model?

  2. Are typical prediction errors small enough for the application?

  3. Does the model explain the data without unnecessary terms?

  4. Can its parameters and predictions be interpreted clearly?

  5. Does it predict new observations well?

  6. Does its behavior make sense according to subject-matter constraints?

Adding parameters often lowers training error, but it can produce : the model follows random fluctuations in the available data and performs poorly on new data. A more complex model is justified only when the added structure is supported and improves meaningful predictions.

When possible, use a training set for fitting and a validation or test set for evaluation. For small data sets, cross-validation or another resampling method can provide a more reliable comparison than in-sample fit alone.

Takeaway: Prefer the simplest model that captures meaningful structure and performs well on data not used to fit it.

Knowing when predictions are limited

A fitted model is an approximation whose reliability depends on the conditions represented by the data. Important limitations include:

  • Poor measurement quality or inconsistent definitions can produce misleading estimates.

  • A nonrepresentative sample may not describe the broader population.

  • can make an apparent association difficult to interpret.

  • beyond the observed range can be highly unreliable.

  • A relationship may change across time, locations, or other conditions.

  • Outliers and influential observations can strongly affect parameter estimates.

  • Regression assumptions, including approximately independent residuals centered around zero with reasonably stable variance, should be investigated rather than assumed.

  • Association alone does not establish causation.

A responsible interpretation states the population, units, data range, uncertainty, and intended use of the model. Predictions should be treated as supported only to the extent that the data and assumptions justify them.

Takeaway: Good modeling includes explicit limits on what the fitted relationship can support.