Part 3: Multi-GPU training with DDP (code walkthrough)

Pubblicato il: 20 settembre 2022
sul canale di: PyTorch
68,849
714

In the third video of this series, Suraj Subramanian walks through the code required to implement distributed training with DDP on multiple GPUs. The video starts with a training script that runs on a single GPU, and migrates that step-by-step to training on 4 GPUs.

Chapters:
00:00 Intro
00:19 Non-distributed training code walkthrough
01:40 Migrating single-GPU code to DDP
02:17 Constructing the Process Group
04:14 Wrapping the model with DDP
04:41 Saving and loading DDP checkpoints
05:33 Distributing the input batch
06:02 Updating the training job’s main() function and entrypoint
08:06 Recap of all code updates
09:13 Run the DDP training job
09:53 Outro

Tutorial page for this video → https://bit.ly/3E219BM
Code used in this video → https://bit.ly/3So1KSO

Like this video and subscribe to the PyTorch channel for more videos like this!


In questa pagina del sito puoi guardare il video online Part 3: Multi-GPU training with DDP (code walkthrough) della durata di ore minuti seconda in buona qualità , che l'utente ha caricato PyTorch 20 settembre 2022, condividi il link con amici e conoscenti, su youtube questo video è già stato visto 68,849 volte e gli è piaciuto 714 spettatori. Buona visione!