FP&A Today: How Microsoft Uses AI to Win at Forecasts
FP&A Today Podcast · 45 min
In this landmark episode of the FP&A Today podcast, Microsoft data scientist Daniel Sousa-Lennox reveals how Microsoft's finance team uses machine learning to achieve up to 99% forecasting accuracy for Azure revenues — and shares the key lessons any FP&A team can apply, regardless of company size.
What makes Daniel Sousa-Lennox's perspective uniquely valuable is that he bridges two worlds: he has the ML engineering depth to build production forecasting systems and the finance domain knowledge to understand what those systems are actually predicting. His key insight is that finance domain expertise doesn't diminish in importance as ML models get more sophisticated — it becomes more important, because the questions the models can answer get harder.
The Microsoft Forecasting Challenge
Microsoft forecasts revenue across dozens of product lines, geographies, and customer segments. The Azure business alone requires thousands of individual forecasts that feed into the consolidated company view. Before AI, this required enormous teams building manual models that were inevitably out of date by the time they were reviewed.
- Azure revenue spans 60+ countries and thousands of SKUs
- Traditional manual approach required 200+ analyst hours per quarter
- Forecast variance was consistently 8-12% from actuals
- Senior leaders were making pricing and investment decisions based on outdated data
The Machine Learning Approach
Daniel explains how Microsoft built an ML forecasting pipeline that ingests real-time usage data, billing data, customer signals, and macroeconomic indicators to generate daily updated forecasts. The models use ensemble methods — combining multiple ML approaches to reduce single-model failure risk.
Achieving 99% Accuracy
The podcast reveals that 99% accuracy isn't achieved by a single breakthrough model — it's the result of iterative improvement over three years. Key factors: (1) continuous feature engineering using domain knowledge from finance team members, (2) rapid retraining when model drift is detected, (3) human-in-the-loop review cycles that catch edge cases.
- Daily model retraining catches regime changes faster than weekly cycles
- Finance domain expertise is critical for feature engineering
- Uncertainty quantification — knowing when the model is less confident — is as important as accuracy
- The models are most accurate 30-45 days out; beyond 90 days, scenario ranges replace point estimates
What Finance Teams of Any Size Can Learn
Daniel emphasises that Microsoft's approach isn't just for large companies. The core principles — driver-based modelling, continuous model evaluation, combining AI with human expertise — apply at any scale. Start small, measure relentlessly, and iterate. Even a 10% accuracy improvement compounds significantly over time.
Key Tools and Technologies
The conversation covers the Microsoft internal tooling (Azure ML, Power BI, Excel) that underlies the forecasting pipeline, as well as the organisational changes required to make AI forecasting work — including how to get finance leaders to trust AI outputs and act on them.
We didn't wake up one day and have 99% accuracy. We woke up every day for three years and made the model 0.1% better. That's how you get to 99%.— Daniel Sousa-Lennox, Data Scientist, Microsoft Finance (FP&A Today Podcast, 2023)
Practical Implementation Checklist
- Start with uncertainty quantification before chasing accuracy — knowing when your model is less confident is as valuable as the predictions themselves
- Involve finance professionals directly in feature engineering: the operational knowledge that makes a variable a useful ML input lives in the finance team, not the data science team
- Set up daily model retraining if your business has fast-moving data — weekly retraining misses regime changes that daily training catches within hours
- Establish a 'model drift' monitoring process: track model accuracy metrics weekly and trigger a model review whenever accuracy drops more than 2% from baseline
- Use ensemble methods (combining multiple ML models) to reduce the risk of single-model failure — Microsoft's approach uses this explicitly for its Azure forecasts
- Replace point estimate forecasts with confidence intervals for horizons beyond 45 days — beyond that range, ranges are more honest and more useful than false-precision single numbers
Microsoft's forecasting success story is not a technology story — it's a discipline story. Daily retraining, continuous feature engineering, domain expertise integration, and iterative improvement over three years produced 99% accuracy. Any finance team willing to invest in those disciplines, at whatever scale, can capture the same compounding returns.
Key Takeaways
99% accuracy comes from iterative improvement, not a single breakthrough
Domain expertise from finance teams is critical for ML feature engineering
Daily model retraining outperforms weekly cycles for fast-moving businesses
Uncertainty quantification — knowing when the model is less confident — is essential
The approach is scalable: core principles apply to teams of any size
Ensemble methods (multiple models combined) reduce single-model failure risk significantly
For horizons beyond 45–90 days, scenario ranges outperform point estimates for decision utility

