If a churn model suddenly starts overfitting after retraining, I’d focus less on the model itself and more on what changed in the data between training cycles.
A few things I’d investigate first:
1. Compare feature distributions before and after retraining
Even small shifts in customer behavior can cause a model to latch onto patterns that don’t generalize well. Look at your top features and check whether their distributions have changed significantly.
2. Check class balance and churn rate changes
If churn has increased or decreased materially in the new training window, the model may be learning patterns that are specific to that period rather than broadly predictive.
3. Review feature importance
Compare feature importance from the previous model against the retrained version. If a handful of variables suddenly dominate predictions, that can be a sign the model is overfitting to recent noise.
4. Validate your training window
A six-month window can sometimes be too narrow, especially if there are seasonal effects or temporary business changes. Try training on a longer history and compare results.
5. Look for data drift
Training and validation performance diverging after retraining is often a symptom of drift. Check whether the validation data represents the same population and behavior patterns as the new training data.
6. Increase regularization
If you’re using a gradient boosting model, experiment with:
- Lower tree depth
- Higher minimum child weight / leaf size
- Lower learning rate
- Stronger regularization parameters
If the issue appeared only after retraining, I’d rank my investigation priorities as:
Data drift → Feature distribution changes → Training window → Regularization
One question: did the validation performance drop immediately after retraining, or did it degrade gradually over subsequent weeks? That detail can help narrow down whether you’re dealing with overfitting or a changing customer population.