Scaling Python for Data Science Using Apache Spark (Garren Staubli)

Publicado el: 27 septiembre 2018
en el canal de: Databricks
329
4

Garren Staubli, a Senior Data Engineer at Blueprint Consulting Services, discusses how python is the de facto language of data science and engineering, which affords it an outsized community of users. However, when many data scientists and engineers come to Spark with a Python background, unexpected performance potholes can stand in the way of progress. These “Performance Potholes” include PySpark’s ease of integration with existing packages (e.g. Pandas, SciPy, Scikit Learn, etc), using Python UDFs, and utilizing the RDD APIs instead of Spark SQL DataFrames without understanding the implications..

Learn more here: https://databricks.com/session/apache...

Article you might like: https://databricks.com/session/analyz...
About: Databricks provides a unified data analytics platform, powered by Apache Spark™, that accelerates innovation by unifying data science, engineering and business.
Read more here: https://databricks.com/product/unifie...

Connect with us:
Website: https://databricks.com
Facebook:   / databricksinc  
Twitter:   / databricks  
LinkedIn:   / databricks  
Instagram:   / databricksinc   Databricks is proud to announce that Gartner has named us a Leader in both the 2021 Magic Quadrant for Cloud Database Management Systems and the 2021 Magic Quadrant for Data Science and Machine Learning Platforms. Download the reports here. https://databricks.com/databricks-nam...


En esta página del sitio puede ver el video en línea Scaling Python for Data Science Using Apache Spark (Garren Staubli) de Duración hora minuto segunda en buena calidad , que subió el usuario Databricks 27 septiembre 2018, comparta el enlace con amigos y conocidos, en youtube este video ya ha sido visto 329 veces y le gustó 4 a los espectadores. Disfruta viendo!