Statistical Questions and Data Distributions
Students distinguish statistical questions from nonstatistical questions and describe data sets by their center, spread, and overall shape.

Illustrations are auto-generated and may be placeholders. They can be refreshed to match the narration.
Identify Statistical Questions
A statistical question is a question that can be answered by collecting data and that anticipates different answers. For example, “How many minutes do sixth-grade students at our school read each evening?” is statistical because students will report different amounts of time. After gathering many responses, you can study the data as a group. In contrast, “How many minutes did Elena read last night?” is nonstatistical because it asks for one specific value. To decide whether a question is statistical, ask: Will the answers vary? Will I need data from several observations or individuals? Questions such as “What are the heights of students in our class?” and “How many pets do students own?” are statistical. A well-written statistical question names the group and the quantity being measured.
Collect and Organize Data
After writing a statistical question, collect data from the group named in the question. Suppose ten students answer, “How many books did you read last month?” Their responses are 1, 3, 2, 4, 2, 5, 3, 2, 6, and 4 books. Record every response using the same unit. Then organize the values in order: 1, 2, 2, 2, 3, 3, 4, 4, 5, 6. A frequency table can show how often each value occurs. For example, 2 books has a frequency of 3 because three students gave that response. You can also display the data on a dot plot, placing one dot above the correct number for each response. Organized displays make patterns, clusters, gaps, and unusual values easier to see.
Examine Distribution Shape
The shape of a data distribution describes how values are arranged across a graph. A distribution may be roughly symmetric, with similar patterns on both sides of its center. It may be skewed right, with most values on the lower end and a tail stretching toward greater values. It may be skewed left, with a tail stretching toward smaller values. A distribution can also have clusters, gaps, or peaks. Imagine a dot plot of quiz scores with most scores from 80 to 90 and only a few scores near 60. The data have a cluster from 80 to 90 and a left tail toward 60, so the distribution is skewed left. When describing shape, consider the entire display rather than focusing on only the highest or lowest value.
Describe Center
The center of a distribution is a value that represents what is typical for the data set. Two common measures of center are the mean and median. To find the mean, add all values and divide by the number of values. For 2, 3, 3, 4, and 8 books, the sum is 20, so the mean is 20 ÷ 5 = 4 books. To find the median, order the values and locate the middle value. The median is 3 books. These centers differ because the value 8 pulls the mean upward. The median is less affected by an unusually high or low value. Always report a measure of center with its unit, and use the distribution’s shape and unusual values to decide which measure best describes a typical result.
Describe Spread
Spread tells how much the data values differ from one another. A simple measure of spread is the range. Find it by subtracting the minimum value from the maximum value. For the data 2, 3, 3, 4, and 8 books, the minimum is 2 and the maximum is 8, so the range is 8 − 2 = 6 books. Two distributions can have the same center but different spreads. For example, 4, 5, 6 has a mean of 5 and a range of 2, while 1, 5, 9 also has a mean of 5 but a range of 8. The second set is more spread out. A graph shows spread through the distance covered by the values, while a numerical measure such as range summarizes that distance.
Interpret Variability
Variability means that data values are not all the same. It is expected when a statistical question is asked because people, objects, and repeated measurements differ. Suppose two sixth-grade classes each have a median homework time of 30 minutes. In Class A, most times are between 27 and 33 minutes. In Class B, times range from 10 to 55 minutes. Although the classes have the same median, Class B has greater variability because its values are more spread out. To interpret a distribution, describe its center, spread, and shape together. You might say, “Both classes have a median of 30 minutes, but Class B has a wider spread and is skewed right because a few students reported long times.” This description gives more information than a single measure alone.
