In this session we cover ways to optimize PySpark code. This includes descriptions of situations where slowness may occur, for example, uneven partitions and skewed joins. To combat these issues I explain repartitioning/coalescing and broadcast joins. I also explain how to place your data in memory or on disk to cache commonly used data sets. Finally, I show the interface where you can monitor memory and CPU usage to make sure you are using the optimal cluster size.
Lastly, I show how to use multiple languages inside of one Databricks notebooks including SQL and R code.
To gain access to code, data, and course materials visit https://kelseyemnett.com/2021/05/30/o....
Sur cette page du site, vous pouvez voir la vidéo en ligne Optimizing PySpark Code durée heure minute seconde en bonne qualité , qui a été Téléchargé par l'utilisateur Data Analysis Lab 30 mai 2021, Partagez le lien avec vos amis et connaissances, sur youtube cette vidéo a déjà été regardée 1,180 fois et il a aimé like téléspectateurs. Bon visionnage!