What you will learn
- Tell the difference between categorical and numerical variables
- Tell the difference between discrete and continuous numerical variables
- Recognize ordinal categories that have a natural order
- Explain why the type of variable decides which statistics make sense
Not all data is the same kind. Some data is made of numbers you can add and average, and some is made of labels you cannot. Before you compute anything, you have to know which type of variable you are dealing with, because the type decides which statistics are even meaningful. Trying to average a set of labels, for example, does not make sense. This short lesson gives you the vocabulary.
Variables and types of variables (MarinStatsLectures)
A clean breakdown of numerical versus categorical and discrete versus continuous. Watch this first for the core distinctions.
Categorical versus numerical
The first split is between categorical and numerical variables. A categorical variable puts each observation into a group or label, like eye color, the sector a company belongs to, or whether a day was an up day or a down day. A numerical variable is an actual number you can do arithmetic with, like a person's height, a stock's return, or the number of trades in a day. The simple test: if it makes sense to take an average of it, it is numerical. You can average returns, but you cannot average sectors.
Key terms
- Categorical variable
- A label or group, like sector or up-day versus down-day. You cannot average it.
- Numerical variable
- An actual number you can do arithmetic with, like a return or a height.
- Discrete variable
- A numerical variable that takes separate, countable values, like the number of trades.
- Continuous variable
- A numerical variable that can take any value in a range, like a return of 1.37 percent.
Discrete versus continuous
- A discrete numerical variable takes separate, countable values, usually whole numbers. The number of up days in a month is discrete, because it can only be a whole number like 12 or 13. You cannot have 12.5 up days.
- A continuous numerical variable can take any value within a range. A stock's return is essentially continuous, because it could be 1.37 percent, or negative 0.42 percent, or any value in between.
In finance, returns and prices are usually treated as continuous, while counts of events, such as the number of winning trades, are discrete. This distinction returns in the next lesson, where discrete and continuous random variables are described with slightly different tools.
A note on ordered categories
One subtlety is worth knowing. Some categorical variables have a natural order, and these are called ordinal. A bond credit rating like AAA, AA, A, BBB is categorical, because it is a label, but the labels clearly rank from safest to riskiest. Eye color, by contrast, has no order and is called nominal. Ordinal data can be ranked but the gaps between categories are not necessarily equal, so you still cannot simply average them like true numbers.
| Type | Can you average it? | Everyday example | Finance example |
|---|---|---|---|
| Numerical, continuous | Yes | Height | A daily return |
| Numerical, discrete | Yes | Number of siblings | Number of winning trades |
| Categorical, nominal | No | Eye color | A company's sector |
| Categorical, ordinal | No, but it can be ranked | Small/medium/large | A bond credit rating |
Classify the variable
Match each example to its variable type.
Ask one question of any variable: does it make sense to average it? If yes, it is numerical. If no, it is a category.
Types of data: nominal, ordinal, interval, ratio (Dr Nic's Maths and Stats)
A friendly tour of data types, including the nominal and ordinal distinction for categories. Good reinforcement of the table above.
Why the type decides the statistics
This all matters because the type of variable decides which statistics you are allowed to compute. You can find the average and standard deviation of a numerical variable like returns, which is exactly what the rest of this unit does. For categorical variables, averaging is meaningless, and instead you count how often each category appears, or find the most common one. Using the wrong kind of statistic for the wrong kind of variable produces numbers with no meaning. Since this unit is about the statistics of returns, and returns are continuous numerical data, everything ahead will assume you are working with numbers you can add and average.
Which statistic makes sense?
You have a dataset where each row is a company, with its sector (tech, energy, finance) and its yearly return. Which of these is a meaningful calculation?
Return is numerical, so you can average it. Sector is categorical, so the meaningful thing is to count how many companies fall in each category. Matching the statistic to the variable type is essential.Why you cannot average a category
Explain in a sentence or two why it makes no sense to take the average of a categorical variable like eye color or sector, even though you can easily average a numerical variable like return.
Write an answer before comparing it with the model response.
Model answer
An average requires adding values and dividing, which only works when the values are real numbers on a scale. Categories like eye color or sector are just labels with no numerical value and no meaningful spacing, so adding them produces nonsense. Returns are actual numbers on a scale, so their average is a real, meaningful summary. The type of data determines whether arithmetic like averaging is even valid.
Now that you can classify data, we can put the pieces together. The next lesson introduces the random variable, the idea that turns an uncertain future outcome, like tomorrow's return, into something you can describe with a probability distribution.