AI Interview Prep

AI Interview Prep

Advanced NLP Interview Questions #22 – The Inter-Annotator Agreement Trap

Why accuracy collapses under class imbalance - and why Cohen’s Kappa is the metric that actually matters.

Hao Hoang's avatar
Hao Hoang
Dec 27, 2025
∙ Paid

You’re in a Senior ML Engineer interview at Google and the interviewer asks:

“We’re building a toxicity detection dataset where only 1% of comments are actually toxic. We hired two annotators. Their inter-annotator agreement is 99%. Are we good to go?”

Most candidates smile and say: “Yes! 99% agreement is amazing. The data is clearly high quality.”

They just failed the interview since they confused 𝐜𝐨𝐧𝐜𝐮𝐫𝐫𝐞𝐧𝐜𝐞 with 𝐬𝐢𝐠𝐧𝐚𝐥.

The reality? 𝘐𝘯 𝘪𝘮𝘣𝘢𝘭𝘢𝘯𝘤𝘦𝘥 𝘥𝘢𝘵𝘢𝘴𝘦𝘵𝘴, “𝘢𝘤𝘤𝘶𝘳𝘢𝘤𝘺” 𝘪𝘴 𝘢 𝘭𝘪𝘢𝘳.

If your dataset is 99% safe and 1% toxic, an annotator could be asleep at the wheel, mark every single comment as “Safe,” and still achieve 99% agreement with another lazy annotator. They found 0% of the toxicity, but on paper, they look perfect.

The Senior Engineer knows you are fighting 𝐂𝐡𝐚𝐧𝐜𝐞 𝐀𝐠𝐫𝐞𝐞𝐦𝐞𝐧𝐭.

When one class dominates (like “Safe” comments), the statistical probability of two people agreeing by accident skyrockets. You aren’t measuring quality, you’re measuring the class imbalance.

To fix this, you need to normalize for that baseline probability.

AI Interview Prep is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

User's avatar

Continue reading this post for free, courtesy of Hao Hoang.

Or purchase a paid subscription.
© 2026 Hao Hoang · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture