What is data fitting?
Data fitting estimates a mathematical function that describes the relationship between measured variables. An empirical model relies primarily on observed data rather than a complete theoretical explanation.
Study 6 Data Fitting and Empirical Models with 12 free online flashcards. Review key terms, definitions, and concepts with this interactive flashcard deck.
What is data fitting?
Data fitting estimates a mathematical function that describes the relationship between measured variables. An empirical model relies primarily on observed data rather than a complete theoretical explanation.
What do the parts of an empirical model represent?
In
y=fx;boldsymbol{beta}+\varepsilon
, f is the functional form, β contains unknown parameters, and ε represents measurement error, random variation, and omitted effects.
Why inspect a scatterplot before fitting?
A scatterplot reveals association, approximate linearity, curvature, clusters, outliers, changing variability, gaps, and limited ranges. It does not by itself establish causation.
Why prefer a sufficiently simple model?
Begin with the simplest function that can describe the main structure. Unnecessary complexity can fit random noise, increase uncertainty, and reduce interpretability.
How are the intercept and slope interpreted?
In y^=b0+b1x, b0 is the predicted response when x=0, if that value is meaningful and in range; b1 is the predicted response change for a one-unit increase in x.
What does least squares minimize?
Least squares chooses coefficients that minimize SSE=∑i=1n(yi−y^i)2. Squaring makes large errors count more heavily than small errors.
What is a residual?
A residual is ei=yi−y^i, the observed value minus the fitted value. A positive residual means underprediction; a negative residual means overprediction.
What patterns signal problems in a residual plot?
A reasonable residual plot generally looks like a random cloud centered around zero. Curves suggest nonlinearity, funnels suggest nonconstant variance, and runs may suggest dependence or change over time.
What does residual standard deviation measure?
The residual standard deviation describes the typical size of unexplained errors. Because it uses the response's units, it helps judge practical prediction accuracy.
What does R2 measure, and what can it not guarantee?
R2=1−SSTSSE is commonly the fraction of observed response variation accounted for by the fitted model. A high R2 alone does not validate a model.
When might nonlinear model forms be useful?
Polynomial models such as y^=b0+b1x+b2x2 describe curvature. Exponential models y^=aekx can describe growth or decay, while power models y^=axk relate multiplicatively.
What is a consequence of transforming nonlinear models?
For an exponential model, taking logarithms gives lny=lna+kx; for a power model, lny=lna+klnx. These transformations can enable linear regression but change error and prediction interpretations.