
Nonuniform Tensor Parallelism: Improving Goodput in Large-Scale Language Model Training
This article explores Nonuniform Tensor Parallelism, a technique proposed to improve effective throughput ('goodput') in large-scale language model (LLM) training. As training jobs span thousands of GPUs over extended periods, even infrequent hardware disruptions can cause disproportionate delays. Nonuniform Tensor Parallelism aims to mitigate the impact of device unavailability and resource fluctuations on overall training efficiency.
Source: NVIDIA Robotics Blog








