Cumulative Relative Frequency
Cumulative relative frequency is the running total of relative frequencies. Learn to compute it, read percentiles from it, and draw an ogive.
Cumulative Relative Frequency: What It Is and What It Isn’t
Most students assume cumulative relative frequency is just adding up percentages until you run out of numbers. That is not wrong, but it is dangerously incomplete. The real definition is a running total of relative frequencies for all values at or below a given point, and the one number that makes it click is the final entry in the column: it always equals 1.00 (or 100%) within rounding error. If your last cumulative relative frequency does not land on 1.00, you have either mis-added, mis-sorted, or rounded intermediate steps when you should not have. The formula, the table, the ogive, and the questions about percentiles that teachers love to ask are covered.
Definition and Formula for Cumulative Relative Frequency
The cumulative relative frequency of a value is the proportion of all observations in a dataset that are less than or equal to that value. The formula is straightforward: for any given value x, take the cumulative frequency (the total count of observations at or below x) and divide it by the total number of observations n. In symbols, cumulative relative frequency = (cumulative frequency) / n. This is not a new kind of math; it is the same relative frequency you already know, just accumulated as you move down the table. The last row always gives 1.0 because every observation is at or below the maximum value. If you multiply that 1.0 by 100, you get the cumulative percentage, which is the same column expressed as a percent. The key distinction from plain relative frequency is that plain relative frequency answers “what fraction of the data is this one value?” while cumulative relative frequency answers “what fraction is this value or anything below it?” That single difference is why the column never decreases as you scan downward.
Computing the Cumulative Relative Frequency Column
To compute the column, start with a frequency table that lists each unique value (or class interval) in ascending order. Add a column for cumulative frequency, which is the running sum of the frequencies: first row equals its own frequency, second row equals first plus second, and so on. Then add the cumulative relative frequency column by dividing each cumulative frequency by the total count n. For example, if you have 20 test scores and 5 of them are 70 or below, the cumulative relative frequency for 70 is 5/20 = 0.25. The next row, say 8 scores are 80 or below, gives 8/20 = 0.40. The column rises steadily, and the final row must read 20/20 = 1.00. If your data is grouped into bins like “10-20,” treat the bin as a single row and use the upper boundary for the cumulative total. The most common error is summing the relative frequencies themselves before dividing, which works only if you first compute all relative frequencies and then add them, but that invites rounding drift. The cleaner path is to sum the raw frequencies first, then divide once per row. If you use a spreadsheet, you can put the raw frequencies in one column, use a running total formula, and divide by the grand total.
Reading Percentiles and Answering ‘At Most’ Questions
Cumulative relative frequency is the tool for any question that uses the phrase “at most,” “or less,” or “up to.” If someone asks what proportion of students scored at most 75, you find the row for 75 and read its cumulative relative frequency. That number is the percentile rank of 75. Conversely, if you want the 60th percentile, you scan down the cumulative relative frequency column until you find the first value that reaches or exceeds 0.60, then read the corresponding data value. For a discrete distribution, the percentile is that data value; for grouped data, you interpolate within the bin. A common trap is confusing “less than” with “at most.” For continuous data, “less than 75” and “at most 75” are the same because the probability of hitting exactly 75 is zero. For discrete data, they differ by the count of exact 75s. The cumulative relative frequency column answers “at most” directly; for “less than,” subtract the relative frequency of the exact value. The same logic applies to cumulative percentage, which is just this column multiplied by 100, so a cumulative relative frequency of 0.60 is the 60th percentile and the 60th cumulative percentage point.
Drawing an Ogive from the Cumulative Relative Frequency Graph
An ogive is the line graph of the cumulative relative frequency column. To draw one, put the data values (or upper class boundaries) on the x-axis and the cumulative relative frequency, from 0 to 1.0, on the y-axis. Plot a point for each row: the x-coordinate is the upper boundary of the class, and the y-coordinate is the cumulative relative frequency for that class. Then connect the points with straight lines. The graph always starts at (lowest boundary, 0) and ends at (highest boundary, 1.0). The shape tells you the distribution: a steep middle section means most data clusters there, while a flat tail means sparse extremes. To read a percentile off the ogive, draw a horizontal line from 0.60 on the y-axis to the curve, then drop straight down to the x-axis. That x-value is the 60th percentile. The ogive is the visual twin of the cumulative relative frequency graph, and it is the single best way to compare two datasets because you can overlay two curves on the same axes. Do not use a bar chart for this; the ogive is a line plot, and the points must be connected to show the running total.
Ogive Example: Step-by-Step with Real Numbers
Suppose a class of 30 students reports the number of books read: 0, 1, 1, 2, 2, 2, 3, 3, 3, 3, 4, 4, 4, 4, 4, 5, 5, 5, 5, 5, 5, 6, 6, 6, 7, 7, 8, 9, 10, 12. The frequency table has values 0 through 10 plus 12. The cumulative frequency for value 3 is 0+1+2+3+4 = 10, because there is one 0, two 1s, three 2s, and four 3s. That gives cumulative relative frequency 10/30 = 0.333. For value 5, the cumulative count is 10 + 6 = 16, so 16/30 = 0.533. The last value, 12, has cumulative frequency 30, so 30/30 = 1.00. To draw the ogive, plot (0,0), then (1, 1/30), (2, 3/30), (3, 10/30), (4, 14/30), (5, 16/30), and so on up to (12, 1.00). Connect with straight lines. The curve rises slowly at first, steepens around 3-5, then flattens.50, which falls between 4 and 5, so you would interpolate to get roughly 4.75 books. That is the median. The ogive example shows exactly why the cumulative column must be sorted: if you jumbled the values, the running total would make no sense and the graph would zigzag backward.
Cumulative Percentage and When to Use It
Multiplying any cumulative relative frequency by 100 gives the cumulative percentage. So the column that read 0.333 becomes 33.3%, and 1.00 becomes 100%. The cumulative percentage is the same data, just in a unit that feels more natural for reports. It is particularly useful when you want to say “the top 20%” or “the bottom 15%” because those are percentage statements. For example, if the cumulative relative frequency for a score of 80 is 0.72, then 72% of the data is at or below 80, and 28% is above. The cumulative percentage is the answer to “what percent of observations are less than or equal to this value?” It is not a new computation, just a rescaling. Be careful with the phrase “percentage point”: if the cumulative relative frequency rises from 0.60 to 0.65, that is a 5 percentage point increase, not a 5% increase. The percent increase is (0.05/0.60)*100 = 8.33%. Mixing these up is a classic exam trap.
Common Mistakes and How to Catch Them
The most frequent error is a cumulative relative frequency column that does not end at 1.00. That almost always comes from rounding each relative frequency to two decimals before adding. Fix it by summing raw frequencies first, then dividing, and only round the final displayed value. The second mistake is applying cumulative logic to nominal data. If your categories are Red, Blue, Green, there is no natural order, so a cumulative column is meaningless. The ogive requires ordinal or numeric data; for nominal data, use a bar chart. The third mistake is the off-by-one bin edge: when data is continuous and a value lands exactly on a bin boundary like 20 in “10-20,” decide once whether the bin is inclusive of the upper or lower bound and apply it consistently. In a spreadsheet, the FREQUENCY function in Excel and Google Sheets requires array entry (Ctrl+Shift+Enter in classic Excel) and returns one more element than the number of bins, which trips up beginners. If you get a #N/A error, you selected the wrong number of cells. A final check: the cumulative relative frequency of the maximum value must be 1.0, and the ogive must start at zero on the left. If your graph starts at 0.5, you forgot to include the initial zero point.
When the Normal Route Fails: What to Do at 1am
You have a homework problem due at 8am and the cumulative relative frequency column will not reconcile. The normal route of checking your arithmetic has failed twice. The first thing to do is rebuild the frequency table from scratch, because a single miscounted frequency poisons every cumulative value below it. If that fails, check whether your data has ties you missed: two 5s counted as one, or a 0 counted as nothing. The second emergency fix is to stop hand-calculating and use a spreadsheet. Put the raw data in one column, use the FREQUENCY function to get bin counts, then build a running total. If the spreadsheet still will not cooperate, strip out all formatting and start a fresh sheet, because a hidden filter or a stray text cell can break the array formula. The third move is to find the raw data list and sort it by hand in a text editor, then count each value out loud. If you are still stuck, calculate the relative frequency of each value individually, then add them with a calculator that shows full precision; do not round until the final step. The cumulative relative frequency of the last row is your sanity check. If it is not 1.00 after all that, you have a data entry error, not a math error, and you need to re-enter the list.
Related Measures and When They Are Overrated
Cumulative relative frequency gets overrated as a catch-all. It is superb for percentiles and “at most” questions, but it is useless for nominal data and actively misleading if you apply it to a relative frequency table for categorical data without a natural order. The plain relative frequency table answers “what share is this exact value?” and that is often the better tool for a single category. The conditional relative frequency, which divides a cell by its row or column total rather than the grand total, answers “what proportion of this row is that column?” and is a different question entirely. Marginal relative frequency, the row or column total divided by the grand total, tells you the overall share of a category. Do not confuse any of these with cumulative relative frequency. If someone asks “what proportion of all data is in this cell?” they want the joint relative frequency, not the cumulative column. The ogive is a great graph, but it is not the right choice for comparing two nominal categories; a side-by-side bar chart does that better. For ordinal data, the cumulative relative frequency table is meaningful, but for nominal data it is nonsense, and any calculator that forces a cumulative column on nominal categories is doing you a disservice.
Spreadsheet Shortcuts: Excel and Google Sheets
In Excel and Google Sheets, the FREQUENCY function takes a data array and a bins array and returns a vertical array of counts. For example, =FREQUENCY(A2:A101, B2:B6) counts how many data points fall into each bin defined by B2:B6. The function returns one more count than the number of bins, which captures everything above the last bin. In classic Excel, you must select the output range, type the formula, and press Ctrl+Shift+Enter to make it an array formula; otherwise you get a #VALUE! error. Google Sheets handles it with a single Enter. To get the cumulative relative frequency, add a column that divides the running total of FREQUENCY results by the grand total =SUM(B2:B7). For a two-way table, the PivotTable option “Show Values As > % of Column Total” gives you conditional relative frequencies for columns, but it does not compute cumulative values. If you need a cumulative relative frequency graph, plot the upper boundaries against the cumulative column and choose the line chart type. The Excel FREQUENCY output is always vertical, one element per bin plus one, so plan your output range before you type. The most common spreadsheet failure is selecting too few cells for the array formula, which truncates the result and corrupts the cumulative column.
Why the Last Column Must Equal One
The cumulative relative frequency of the largest value in a dataset is always 1.0, because every observation is less than or equal to the maximum. That is not a coincidence; it is a definitional truth. If your final entry is 0.98, you have either dropped a data point, miscounted a frequency, or rounded intermediate steps. The fix is to carry full precision in the running total and round only the final display. Some textbooks allow the last entry to be 0.99 or 1.01 due to rounding, but the raw sum of all relative frequencies must be 1.0 exactly. When you build a cumulative relative frequency table, the last cumulative frequency equals the total count n, so dividing by n gives 1. If you see anything else, go back to the raw data list and recount. This one check catches more errors than any other diagnostic. It is also the reason the ogive ends at the top right corner: the graph cannot go above 1.0, because you cannot have more than 100% of the data at or below the maximum.
For Whom This Subject Works and for Whom It Does Not
Cumulative relative frequency suits anyone working with ordinal, interval, or ratio data who needs to answer questions about position, rank, or proportion. It is the right tool for a market researcher asking what share of customers spend at most $50, or a quality control engineer checking what proportion of parts are within tolerance. It also suits students in an introductory statistics course who have already mastered fractions and percentages, because the math is just division and addition. It does not suit anyone analyzing nominal categories like eye color or country of birth, where there is no order to accumulate. It also does not suit someone looking for a theoretical probability, because cumulative relative frequency describes the observed sample, not the long-run chance. If you need to know what fraction of a population lies below a cutoff and you have the data, this is your measure. If you have no data and are trying to predict the future, you are better served by a probability model, not an empirical table.
Common Questions
Why does my cumulative relative frequency column not end at exactly 1.00?
That happens when you round each relative frequency to two decimals before adding them. The correct method is to sum the raw frequencies first, then divide by the total, and round only the final displayed value. If you still get 0.98, you likely dropped a data point or miscounted a frequency; recount the raw list.
How do I calculate cumulative relative frequency when my data is a range like '10-20' instead of a single number?
Treat the range as one class. Use the upper boundary (20 in this case) as the x-value for the cumulative total. Count how many observations fall in that bin, add it to the running total of frequencies, then divide by the grand total. The cumulative relative frequency for that bin answers 'what proportion is at most 20?'
What is the difference between relative frequency and cumulative relative frequency?
Relative frequency of a value is its count divided by the total, answering 'what share is exactly this value?' Cumulative relative frequency adds up the relative frequencies for all values at or below the current one, answering 'what share is this value or less?' The last cumulative value is always 1.0.
Is cumulative relative frequency the same as percentile?
Closely related. The cumulative relative frequency of a value, multiplied by 100, is its cumulative percentage, which equals the percentile rank. The 60th percentile is the value at which the cumulative relative frequency reaches 0.60. For grouped data, you may need to interpolate within a bin.
Can I use cumulative relative frequency for nominal data like colors?
No. Nominal data has no natural order, so a running total is meaningless. A cumulative column is only valid for ordinal, interval, or ratio data. For nominal categories, use a simple relative frequency table or a bar chart instead of an ogive.
What does the ogive graph show me that the table does not?
The ogive shows the shape of the cumulative distribution at a glance. You can see where the data clusters (steep slopes) and where it thins out (flat regions). It also lets you read percentiles directly by drawing a horizontal line from the y-axis to the curve and dropping down. A table gives exact numbers, but the graph shows the pattern faster.
How many decimal places should I use for cumulative relative frequency?
Use one more decimal place than the raw data, or match the precision of the original measurements.333 or 0.33 is fine. The key is to carry full precision during calculation and round only the final answer to avoid cumulative rounding errors.