Youโre in a Senior ML Engineer interview at Google DeepMind and the interviewer asks:
โYou have just launched a new Mixture of Experts (MoE) training run. After a few thousand steps, you check the logs and see the validation loss has flatlined. What is the ๐ฆ๐จ๐ฌ๐ญ ๐ฅ๐ข๐ค๐๐ฅ๐ฒ ๐๐๐ฎ๐ฌ๐ specific to an MoE, and how do you fix it?โ


