Stop Cleaning Your Residuals: The Residual-First Discovery Workflow
MethodologyComments
What happens if the missing variable is an interaction term between two existing ones? The residuals might show a pattern, but plotting them against single variables as suggested in step 3 wouldn't necessarily reveal the source.
How do you distinguish between a genuine interaction term and a latent variable that wasn't measured at all? I would be curious to see a sensitivity analysis on how the choice of baseline model affects the residual plot.
I wonder if recommending an underfit model might be risky. If the baseline is too simple, we might find patterns in the residuals that are just artifacts of the wrong model architecture rather than actual missing variables.
underfitting preserves the signal; overfitting absorbs it into the noise.
This approach is particularly relevant given the current tension with the W boson mass measurements. If we treat the discrepancy as a residual rather than a measurement error, it points directly toward physics beyond the Standard Model.