Five support requests arrived, but only three wait times were counted. Did two requests disappear? groupby().count() counts non-missing values in the selected column, while size() counts rows in each group. When measurements are missing, these answer different questions.
size counts rows; count counts populated values
The diagram distinguishes received rows from measured wait times. Email's zero-minute wait is a valid value and is counted; a None measurement is excluded by count().
Create five illustrative support requests. Two have wait_min set to None: the requests arrived, but their wait times were not recorded. 0 is a valid measurement, not a missing value.
import pandas as pd
calls = pd.DataFrame({
"channel": ["email", "email", "email", "chat", "chat"],
"wait_min": [12, None, 0, 9, None],
})
summary = calls.groupby("channel").agg(
received=("wait_min", "size"),
timed=("wait_min", "count"),
average=("wait_min", "mean"),
)
print(summary.to_string()) received timed average
channel
chat 2 1 9.0
email 3 2 6.0Of three email rows, the one with None contributes to received but not timed. The zero-minute row contributes to both. The mean is (12 + 0) / 2 = 6.0, not 4.0 from treating the missing measurement as another observed zero.
Find missing cells when the numbers differ
Calculate the gap and inspect the corresponding rows:
summary["missing_time"] = summary["received"] - summary["timed"]
print(summary["missing_time"].to_dict())
print(calls.loc[calls["wait_min"].isna(), ["channel", "wait_min"]])missing_time is one for chat and one for email. The second output identifies the row with no wait time in each channel. Showing received and measured counts together prevents readers from treating the difference as a calculation error. A useful test mixes zero, None, and a positive value and checks received == timed + missing_time.
What if the grouping column is missing too?
This dataset has a channel on every row. If a row has no channel, decide separately whether that row should form a group. Changing count to size cannot recover a row that was excluded when groups were formed. Compare input row count with grouped row count before interpreting the result.
Name each metric for the question it answers
For “requests received by channel,” row count from size is natural. For “requests with a measured wait time,” count fits. Report the number of missing measurements beside the average so readers know its sample size. Filling missing waits with zero without evidence can lower the average; a measured zero and an absent value are different facts.
Before optimizing a large aggregation, define the metric contract: which rows count as received, why measurements are missing, and what to do with rows lacking a channel. Duplicate input rows can inflate both size and count, so also check the input key for duplicates.
Key takeaways: the counts answer different questions
groupby().size() counts rows in a group; a column's count() counts its non-missing values. Zero is a value and remains in the count. When totals differ, inspect the missing cells and the intended meaning of each metric before changing the function.

