You’re in an AI Engineer interview at OpenAI and the interviewer asks:
“You’ve successfully fine-tuned a model with RL. It’s now excellent at following instructions, but it’s become ‘dumber’ at general knowledge and creative writing. What is this phenomenon called, and what specific term would you add to your loss function to prevent this?”


