Read SAT scatterplots, identify positive, negative, or no association, and use a line of best fit to predict values, interpret slope and intercept, and reason about residuals.
A linear model summarizing the trend in a scatterplot.
predicted change in per one-unit increase in
State it with units, e.g. "about 6 more points per additional hour."
predicted when
Only meaningful if $x = 0$ is within or near the data range.
Above the line = positive (underpredicted); below = negative (overpredicted).
A scatterplot plots one point for each observation using two variables — one on the horizontal axis, one on the vertical axis. The cloud of points reveals whether the two variables tend to move together. The SAT expects you to describe that tendency, called the association, and to reason with a line of best fit drawn through the cloud.
There are three associations to recognize at a glance. A positive association means the points rise from lower-left to upper-right — as increases, tends to increase. A negative association means the points fall from upper-left to lower-right. No association means the points show no consistent direction. Association can also be described as strong (points hug an imaginary line tightly) or weak (points scatter loosely), and as linear or nonlinear (curved).
| Pattern of points | Association | Slope of best-fit line |
|---|---|---|
| rise left to right | positive | positive |
| fall left to right | negative | negative |
| no clear direction | none | near zero |
| tight around a line | strong | (clearer fit) |
The line of best fit (or trend line) is the straight line that comes closest to all the points at once. It is a model: it summarizes the trend so you can make predictions. Its equation has the familiar form , and the SAT loves to ask you to interpret the two constants in the context of the data.
The slope tells you the predicted change in for each one-unit increase in — always stated with the data's units. The -intercept is the predicted value of when ; sometimes it is meaningful and sometimes it is only a mathematical artifact outside the data's range.
Suppose a scatterplot relates hours studied () to a -point quiz score (), and the line of best fit is .
Interpret the slope. The slope means the model predicts each additional hour of study is associated with about a -point higher score.
Interpret the intercept. The intercept predicts a student who studied hours scores about .
Predict a value. For a student who studied hours:
The model predicts about an .
A residual is the vertical gap between an actual point and the line:
If an actual student who studied hours scored , the residual is : that student did points better than the model predicted. Points above the line have positive residuals — the model underpredicted them. Points below the line have negative residuals — the model overpredicted them. The best-fit line is precisely the line that makes these residuals as small as possible overall.
Many SAT questions never give the equation; they hand you the plotted line and ask you to read a value or count points meeting a condition — for example, "how many points lie above the line" or "for how many observations did the line overpredict." Overprediction means the actual point is below the line. Work carefully off the gridlines.
Finally, be wary of extrapolation: predictions far outside the plotted range of the data are unreliable, because there is no evidence the trend continues. Using to predict the score for hours of study would output an impossible on a -point quiz — a reminder that the model is only trustworthy near the data it summarizes. And a strong association never proves causation: two variables can move together because a third factor drives both.
On a scatterplot, the points fall steadily from the upper-left toward the lower-right. What association does this show, and what is the sign of the best-fit slope?
The line of best fit for a data set is . What does the model predict when ?
For used cars, the line of best fit relating age in years () to price in dollars () is . What does the number mean?
A line of best fit is . A student who studied hours actually scored . Find the residual and state whether the model over- or underpredicted.
A line of best fit passes through the points and on a scatterplot. Write its equation, then predict when .
A scatterplot has data points. The line of best fit passes below of the points and above of them. For how many points did the model overpredict the actual value, and does that seem consistent with a good fit?
Confusing correlation with causation. Seeing a strong trend and concluding one variable causes the other.
A scatterplot only describes how variables move together. A lurking third variable can drive both, so never infer cause from association alone.
Reading the slope as a total instead of a per-one-unit rate of change.
Slope answers "how much does change for each in ?" In , the is per hour, not a total; the intercept is the value at .
Mislabeling residual direction — calling a point above the line "overpredicted."
Above the line means actual predicted, so the model underpredicted (positive residual). Overprediction is when the actual point is below the line.
Trusting predictions far outside the data range (extrapolation).
A best-fit line is only reliable near the plotted data. Plugging in a far-away can yield impossible values (a negative price, a on a -point quiz).
A scatterplot of daily high temperature (x, in degrees) versus number of hot chocolates sold (y) at a cafe shows points falling from upper-left to lower-right. Which best describes the association?
For a set of used cars, the line of best fit relating a car's age in years (x) to its price in dollars (y) is y = −900x + 15000. Based on this model, what is the best interpretation of the number −900?
The line of best fit for a scatterplot is y = 2.5x + 8. According to this model, what is the predicted value of y when x = 6?
A line of best fit is y = 4x + 10. A particular data point has x = 5 and an actual y-value of 25. What is the residual for this point?
A line of best fit passes through the points (1, 14) and (5, 2) on a scatterplot. What is the slope of this line?
A scatterplot has 10 data points and a line of best fit. The line lies below 4 of the points and above the other 6. For how many points did the model UNDERpredict the actual y-value?
Read slope as a rate of change, move fluently between slope-intercept, standard, and point-slope form, and handle the parallel and perpendicular slope relationships the SAT rewards.
Compute mean, median, and mode, reason about range and standard deviation, and judge the effect of outliers on SAT statistics questions.