Large Language Models in FP&A: An Architecture for Grounded Variance Commentary and Driver-Based Narrative Generation
Main Article Content
Abstract
FP&A organisations spend a disproportionate share of each close cycle producing variance commentary, flux narrative, and board-pack discussion that explain how reported results moved against budget, forecast, and prior period. This paper describes an enterprise architecture that pairs a deterministic numerical engine with retrieval-augmented generation, an explicit driver-relationship graph, and a human-in-the-loop review step to produce this narrative without offloading arithmetic or attribution to a general-purpose large language model. In a three-business-unit pilot, commentary cycle time fell from roughly six hours to under one hour per unit per cycle; numerical accuracy ranged from 97.9% to 98.7% against source-of-truth ledger figures; and analyst reviewer ratings ranged from 4.5 to 4.7 on a five-point scale. The architecture is designed for SOX-controlled environments and runs against ERP and EPM systems already in place at most enterprises.