MFU = Model FLOP Utilization: fraction of the GPU’s peak FLOP/s actually realized. 40% is typical for well-tuned dense training; MoE routing and small batches push it lower.