Advanced NLP Interview Questions #22 – The Inter-Annotator Agreement Trap
Why accuracy collapses under class imbalance - and why Cohen’s Kappa is the metric that actually matters.
You’re in a Senior ML Engineer interview at Google and the interviewer asks:
“We’re building a toxicity detection dataset where only 1% of comments are actually toxic. We hired two annotators. Their inter-annotator agreement is 99%. Are we good to go?”
Most candidates smile and say: “Yes! 99% agreement is amazing. The data is clearly high quality.”
They just failed the interview since they confused 𝐜𝐨𝐧𝐜𝐮𝐫𝐫𝐞𝐧𝐜𝐞 with 𝐬𝐢𝐠𝐧𝐚𝐥.
The reality? 𝘐𝘯 𝘪𝘮𝘣𝘢𝘭𝘢𝘯𝘤𝘦𝘥 𝘥𝘢𝘵𝘢𝘴𝘦𝘵𝘴, “𝘢𝘤𝘤𝘶𝘳𝘢𝘤𝘺” 𝘪𝘴 𝘢 𝘭𝘪𝘢𝘳.
If your dataset is 99% safe and 1% toxic, an annotator could be asleep at the wheel, mark every single comment as “Safe,” and still achieve 99% agreement with another lazy annotator. They found 0% of the toxicity, but on paper, they look perfect.
The Senior Engineer knows you are fighting 𝐂𝐡𝐚𝐧𝐜𝐞 𝐀𝐠𝐫𝐞𝐞𝐦𝐞𝐧𝐭.
When one class dominates (like “Safe” comments), the statistical probability of two people agreeing by accident skyrockets. You aren’t measuring quality, you’re measuring the class imbalance.
To fix this, you need to normalize for that baseline probability.


