How to Make a Relative Frequency Table
Build a relative frequency table from raw data, including grouped data in class intervals, and turn it into a relative frequency histogram, with examples.
The Relative Frequency Table: How to Make One
You have a column of quiz scores and your instructor wants a relative frequency table by tomorrow. Turn those raw numbers into proportions that show what share of the class scored where. A relative frequency table does that. Build one by following fixed steps, not by guessing.
Layout of the Relative Frequency Table
Three Essential Columns
The skeleton is simple: one column for the value or category, one for the count of how often it occurs, and one for the relative frequency itself, which is that count divided by the total number of observations. For a single categorical variable, the table has three columns and a total row at the bottom. The total row must show the sum of the counts equaling n, and the sum of the relative frequencies equaling 1.00 (or 100% if you are working in percentages).
When your data is numerical and ungrouped, each distinct value gets its own row. That works fine when the values are few, like dice rolls or yes/no responses. But with something like heights or incomes, each value occurs once or twice, and the table becomes a long, useless list. That is when you move to grouping.
The Relative Frequency Column
The relative frequency column answers the question, “Out of all observations, what fraction falls into this row?” For a report, add a fourth column for cumulative relative frequency, which runs the proportions upward so that the last row shows 1.00. That column is only meaningful for ordinal or numerical data; for nominal categories like “Red” and “Blue,” it is nonsense, so leave it out.
A common mistake is to label the columns vaguely. Write “Relative Frequency” explicitly, and if you add the cumulative column, write “Cumulative Relative Frequency” in full. A reader should never have to guess what a column holds. The table should also state the total n somewhere, either in a caption or in the top-left cell, because without it the relative frequencies are just floating numbers.
Grouped Numerical Data: Choosing Class Width
How Many Bins?
When your data is continuous, like test scores or ages, bin it into class intervals. The choice of class width determines how much detail survives and how much is smoothed away. Aim for between 5 and 15 bins, depending on the size of the dataset and the spread of the values. For a dataset of 50 to 200 points, 7 to 10 bins usually works well.
To find a starting width, take the range of the data, which is the maximum minus the minimum, and divide by the number of bins you want. Round up to a convenient number, like 5, 10, or 20. For example, if scores run from 40 to 98, the range is 58. Dividing by 8 gives 7.25, so round to 8 or 10. Every class interval must have the same width, and the boundaries must be chosen so that no value falls into two bins.
The Boundary Trap
Here is the trap that catches most students: what happens to a value exactly on the boundary, like 50 when your bins are 40-49 and 50-59? Decide up front whether the interval is inclusive of the lower boundary and exclusive of the upper one, or vice versa. The standard convention is to include the lower boundary and exclude the upper, so a value of 50 goes into the 50-59 bin, not the 40-49 bin. Write this rule at the top of your table so there is no ambiguity.
For a grouped relative frequency distribution, the relative frequency is still the count in the bin divided by the total n. The class intervals are just labels for the rows; the arithmetic is unchanged. The cumulative relative frequency column then tells you what proportion of the data falls below the upper boundary of each bin, which is how you find percentiles later.
Relative Frequency Histogram and Bar Chart
Histogram for Numerical Data
Once the table is built, the visual is next. A relative frequency histogram is the standard plot for grouped numerical data. The x-axis shows the class intervals, and the y-axis shows the relative frequency, which is the proportion of observations in each bin. Unlike a frequency histogram, which uses raw counts, the relative version uses the proportions, so two datasets of different sizes can be compared on the same scale.
For categorical data, use a bar chart instead. The bars do not touch, because there is no continuity between categories. The height of each bar is the relative frequency, and the y-axis should be labeled “Relative Frequency” or “Proportion,” not “Count.” A frequent error is to plot the raw frequencies and then mislabel the axis as relative; check the numbers on the y-axis before you print or submit.
Equal Bin Widths
A histogram with equal bin widths has bars of equal width, and the area of each bar is proportional to the relative frequency of that bin. That is the property that makes the histogram a true density plot. If you use unequal bin widths, the y-axis must become “relative frequency per unit width,” which is more than most students need, so stick with equal bins unless you have a specific reason not to.
When you build the chart, resist the urge to use a pie chart for numerical data. A pie chart is for nominal categories only. For a relative frequency distribution of test scores, a histogram is the only defensible choice. The bar chart works for the raw categories but not for binned numbers.
Worked Examples
Discrete Data: Exercise Responses
Walk through a concrete case. Suppose you have 30 responses to a survey question, “How many times did you exercise last week?” and the values are: 3, 5, 0, 2, 4, 3, 6, 1, 2, 3, 5, 4, 0, 2, 3, 4, 5, 6, 1, 2, 3, 4, 0, 2, 3, 5, 4, 2, 3, 1. The counts are: 0 appears 3 times, 1 appears 3, 2 appears 6, 3 appears 8, 4 appears 5, 5 appears 4, 6 appears 2, and 7 appears 0. Total n is 30.
Divide each count by 30 to get the relative frequency: 0.10, 0.10, 0.20, 0.27, 0.17, 0.13, 0.07, and 0.00.For the cumulative relative frequency, start with the first row: 0.10. Add the next: 0.20, then 0.40, 0.67, 0.84, 0.97, 1.00. The last cumulative value must equal 1.00; if it does not, go back and recheck your division.
Continuous Data: Heights
Now take a continuous dataset: heights in inches of 20 people: 58, 60, 62, 62, 63, 64, 64, 65, 66, 66, 67, 68, 68, 69, 70, 70, 71, 72, 74, 75. The range is 75 minus 58, which is 17. Choose a class width of 4, giving bins 58-61, 62-65, 66-69, 70-73, and 74-77. The counts are 1, 5, 7, 5, and 2, totaling 20. Relative frequencies are 0.05, 0.25, 0.35, 0.25, and 0.10. The cumulative relative frequencies run 0.05, 0.30, 0.65, 0.90, and 1.00. The histogram would show a peak in the 66-69 bin, telling you most people are in that range.
Two-Way Table
For a two-way table, say gender by exercise category (none, some, regular), the joint relative frequency is each cell divided by the grand total. If 10 out of 100 respondents are women who exercise regularly, that joint relative frequency is 0.10. The conditional relative frequency would divide by the row or column total instead, such as the proportion of women who exercise regularly, which uses the total number of women as the denominator. Do not mix the two.
How to Calculate Relative Frequency
The formula is fixed: relative frequency equals the frequency of the value divided by the total number of observations, written as f divided by n. For a single category, count how many times that value appears, then divide by the grand total. Express the result as a decimal, a percentage, or a fraction; all three are correct, but pick one format and use it consistently throughout your table.
Sort First
When your data is a list of raw numbers, the first step is to sort them. This makes counting easier and reveals any outliers. Then tally each distinct value, write down the counts, and sum the counts to get n. The sum of all relative frequencies must be 1.00; if your rounded values add to 0.99 or 1.01, do not adjust the numbers to force the sum. Report the unrounded sum or round only the final total. This is the “rounding cascade” error, and it is the most common reason a table fails the 1.00 check.
For grouped data, the same formula applies, but the frequency is the count within a bin. The class intervals act as the categories. You cannot calculate a relative frequency for a single raw value unless it is a distinct category; with continuous data, the bin is the unit of analysis.
Watch Your Denominator
One warning: never calculate a relative frequency using a subtotal as the denominator when you mean the grand total. If you have a two-way table and you divide a cell by its row total, that is a conditional relative frequency, not a marginal one. The marginal relative frequency always divides by the grand total. Label your columns so the reader knows which denominator you used.
Cumulative Relative Frequency
The cumulative relative frequency column is a running total. Start with the relative frequency of the first row, then add the next relative frequency, and so on. The last entry must equal 1.00, because it represents all the data. This column lets you answer questions like, “What proportion of students scored 70 or below?” by finding the row where the cumulative value crosses that threshold.
For numerical data, the cumulative relative frequency is the empirical cumulative distribution function evaluated at each bin. It is the tool for percentiles: the 50th percentile is the value where the cumulative relative frequency first reaches 0.50. If your data is continuous and you have binned it, the cumulative column gives you the proportion below the upper boundary of each bin.
Fix the Rounding Cascade
Here is the failure case: you build the table, and the cumulative column ends at 0.99 or 1.01 instead of 1.00. That is the rounding cascade again. The fix is to compute the cumulative values from the unrounded relative frequencies, then round each cumulative total to two or three decimals. Do not round each relative frequency first and then add the rounded values.
Do not put a cumulative column on a nominal table. For categories like “Red,” “Blue,” and “Green,” there is no order, so a running total is meaningless. The calculator on a website may still show the column, but you should ignore it and remove it from your report.
Relative Frequency Excel and Spreadsheet Shortcuts
COUNTIF for Small Datasets
Spreadsheets have two built-in ways to build a relative frequency table, and choosing the wrong one wastes time. The first is the COUNTIF function, which takes a range and a criterion, returning the count of matching cells. Use it to count each value or bin, then divide by the total using a separate cell. This is the simple, single-cell approach that works for small datasets.
FREQUENCY for Large Datasets
The second is the FREQUENCY function, an array formula that returns multiple counts at once. The syntax is FREQUENCY(data_array, bins_array), where bins_array is a range of bin boundaries. This is faster for large datasets but has a catch: you must select the output range, type the formula, and press Ctrl+Shift+Enter (on a PC) or Command+Shift+Enter (on a Mac). If you press only Enter, you get a single value or a #N/A error. Also, the bins_array must contain the upper boundaries for each bin, and the function returns one more count than the number of bins, so select one extra cell.
Once you have the counts, the relative frequency is a simple division. Create a column next to the counts and write a formula like =B2/$B$10, where B2 is the first count and B10 is the total. Use absolute references (the dollar signs) so you can copy the formula down without the denominator shifting. The cumulative relative frequency column then uses a running sum, which you can do with the SUM function and a relative reference that grows as you drag down.
Avoid the Off-by-One Error
The most common spreadsheet mistake is the off-by-one error with FREQUENCY. If you select too few cells for the output, you miss the last bin; if you select too many, you get a #N/A in the extra cell. Always select one more output cell than the number of bins, and if you see #N/A, delete the extra cell or adjust the selection. For a quick check, sum the frequency output; it must equal n.
Building the Table Without a Spreadsheet
By Hand
You are on a school computer with no spreadsheet software, and the Wi-Fi is down. The relative frequency table can still be built by hand. Start with a clean sheet of paper, draw three columns, and label them “Value,” “Frequency,” and “Relative Frequency.” Sort your data by hand, tally marks work fine, and count each value. Sum the frequencies to get n, then divide each count by n using long division or a calculator.
The failure case is when you have 200 values and no calculator. In that situation, round each relative frequency to two decimals, but keep a running total on a scrap of paper. The sum may drift from 1.00 by a cent; that is acceptable in a draft, but for a final report you must reconcile the last value. The cleanest fix is to adjust the largest relative frequency by the difference, so the total equals 1.00. This is called “prorating,” and it is standard practice.
Grouped Table by Hand
For a grouped table, you must first decide the bin width. Write the minimum and maximum values, subtract to get the range, and divide by the desired number of bins. Round up to a convenient number. Then create bins that cover every value exactly once. A value on the boundary goes to the higher bin if you use the “lower inclusive” rule, which is the default. Write that rule at the top of the table.
The final check is the cumulative column. Add the relative frequencies from the top down, and the last entry must be 1.00. If it is not, you made an arithmetic error in the division, not in the addition. Go back and redo the division for each row.
Common Pitfalls and How to Avoid Them
The denominator drift is the first trap. You are working on a two-way table, and you divide a cell by its row total because it feels natural. That gives you a conditional relative frequency, not a marginal one. The marginal relative frequency always uses the grand total. If your instructor asks for the relative frequency table of gender, the denominator is the total number of people, not the number of males or females.
The empty bin error shows up with the FREQUENCY function. You select the output range, press Ctrl+Shift+Enter, and get a #N/A in one cell. That is because you selected one cell too many or one too few. The fix is to select exactly one more cell than the number of bins, because FREQUENCY returns the count of values above the last bin boundary as well.
The off-by-one bin edge happens with continuous data. If your bin is 10-20 and a value is exactly 20, where does it go? The rule is to include the lower boundary and exclude the upper, so 20 goes into the next bin. If you forget, your histogram will be shifted, and the cumulative column will be wrong. State the rule on the table so anyone reading it knows.
The rounding cascade is the most frustrating. You round each relative frequency to two decimals, sum them, and get 0.99. The correct practice is to compute the cumulative from the unrounded values and round only the final sum. If you must round each value, note the total may not equal 1.00, and adjust the largest value by the difference.
The chart mislabel is a quiet error. You make a bar chart of frequencies, then label the y-axis “Relative Frequency” without dividing by n. The shape looks identical, but the numbers on the axis are counts, not proportions. Check the maximum value on the y-axis: if it is a whole number like 15, it is a frequency chart. If it is a decimal like 0.15, it is relative.
When the normal route fails, such as a calculator that throws a division by zero because your total n is zero, stop. Relative frequency is undefined for an empty dataset. You cannot divide by zero, and no amount of rounding will fix it. Go back and verify your data entry; you likely have a missing value or an empty column.
Common Questions
Why does my cumulative relative frequency column not end at exactly 1.00?
This is almost always the rounding cascade. You rounded each relative frequency to two decimals, then added the rounded values. The fix is to compute the cumulative from the unrounded proportions and round only the final cumulative total. If you must round each value, adjust the last cumulative entry to force 1.00, and note the adjustment in a footnote.
How do I calculate relative frequency when my data is a range like '10-20' instead of a single number?
A range like 10-20 is a class interval, or bin. You count how many observations fall into that bin, then divide by the total n. The relative frequency is the count in the bin divided by n. The bin label is just a name; the arithmetic does not change.
What is the difference between the relative frequency I get here and the probability I see in my textbook?
Relative frequency describes what happened in your sample; it is observed data. Probability is a theoretical long-run proportion. As the number of trials grows, relative frequency tends to approach the theoretical probability, a principle called the law of large numbers, but for a small sample they can differ.
How do I use the calculator for a two-way table with three rows and four columns?
A two-way table of counts is not a single list, so a basic relative frequency calculator cannot handle it directly. You must enter the counts as raw data or use a matrix-capable tool. If you have only the summary table, compute each cell divided by the grand total yourself.
Is it okay to use a relative frequency as a percentage in a pie chart?
A pie chart works for nominal categories, where order does not matter. For ordinal or numerical data, a bar chart or histogram is better because it preserves the order. If your categories have no natural sequence, a pie chart is acceptable, but a bar chart is usually clearer for comparing proportions.
What should I do if my total n is zero?
Relative frequency is undefined when n equals zero because you cannot divide by zero. Go back to your data; you likely have an empty column or a filtering error. If the dataset is genuinely empty, report that no relative frequencies can be computed.
How many decimal places should I use?
Use one more decimal place than the raw data has. If your counts are whole numbers and n is under 100, two decimals (like 0.23) is standard. For larger n, three decimals may be needed to show differences. Match the precision of your measurements; do not report 0.333333 when the data only has two significant digits.