You’re in a ML Engineer interview at Anthropic and they ask:
“You’re pre-training a 500B parameter model. You’re only doing one epoch over a 10T token dataset, so you’re clearly not overfitting. Why on earth are you still using weight decay?”
Most candidates…


