Data Engineer PySpark Data Bricks Session Day 6

Опубликовано: 20 Октябрь 2024
на канале: Code with Kristi
274
5

📍 𝐒𝐭𝐚𝐫𝐭 𝐚 𝐒𝐩𝐚𝐫𝐤 𝐒𝐞𝐬𝐬𝐢𝐨𝐧 : Set up the PySpark environment.
🧣 𝐂𝐫𝐞𝐚𝐭𝐞 𝐚 𝐋𝐢𝐬𝐭 : Define the list with three elements.
📢 𝐏𝐚𝐫𝐚𝐥𝐥𝐞𝐥𝐢𝐳𝐞 𝐭𝐡𝐞 𝐋𝐢𝐬𝐭 : Distribute the list across the cluster nodes.
🔔 𝐂𝐨𝐧𝐯𝐞𝐫𝐭 𝐭𝐨 𝐃𝐚𝐭𝐚𝐅𝐫𝐚𝐦𝐞 : Convert the distributed RDD to a DataFrame.
🔋 𝐏𝐞𝐫𝐟𝐨𝐫𝐦 𝐏𝐫𝐨𝐜𝐞𝐬𝐬𝐢𝐧𝐠 : Show the contents and perform any desired operations.

📍 This video will explain how to write first program in PySpark.
Don't miss out on this opportunity to excel!
🚀 *Course:* Master Azure Data Engineering
📅 *Last Date:* 15 Jan 2025

Course Registration: https://tinyurl.com/5n7aatdm

Don't miss out on this opportunity to upscale your skills and dive deep into the realm of data engineering! Reserve your spot now! 🎉

#hiring #career #databricks #azure #career #hiring #databricksanalytics


📢 Video Link:    • Data Engineer PySpark Data Bricks Session ...  

LinkedIn Profile of author:
  / sachin-saxena-graphic-designer  

Code Source Link:
https://lnkd.in/g67a4kY3

𝐄𝐱𝐩𝐥𝐚𝐧𝐚𝐭𝐢𝐨𝐧 𝐨𝐟 𝐭𝐡𝐞 𝐂𝐨𝐝𝐞 :

𝟏. 𝐒𝐩𝐚𝐫𝐤 𝐒𝐞𝐬𝐬𝐢𝐨𝐧 : The SparkSession is created to provide an entry point for Spark functionality.
𝟐. 𝐋𝐢𝐬𝐭 𝐂𝐫𝐞𝐚𝐭𝐢𝐨𝐧 : A list of three elements is defined.
𝟑. 𝐏𝐚𝐫𝐚𝐥𝐥𝐞𝐥𝐢𝐳𝐞 : The list is parallelized with numSlices=3, which ensures that each element is assigned to a different partition in the RDD. This is how we can distribute it across the three nodes.
𝟒. 𝐂𝐨𝐧𝐯𝐞𝐫𝐭 𝐭𝐨 𝐃𝐚𝐭𝐚𝐅𝐫𝐚𝐦𝐞 : The RDD is mapped to a tuple format to convert it into a DataFrame. The column is named "element".
𝟓. 𝐃𝐢𝐬𝐩𝐥𝐚𝐲 𝐃𝐚𝐭𝐚𝐅𝐫𝐚𝐦𝐞 : The contents of the DataFrame are printed using df.show(), which will display each element as a separate row.
𝟔. 𝐂𝐨𝐮𝐧𝐭 : The total number of elements is counted and printed.
𝟕. 𝐅𝐮𝐫𝐭𝐡𝐞𝐫 𝐏𝐫𝐨𝐜𝐞𝐬𝐬𝐢𝐧𝐠 : An optional step is included to filter the DataFrame for elements containing "1" and display the result.
𝟖. 𝐒𝐭𝐨𝐩 𝐒𝐩𝐚𝐫𝐤 𝐒𝐞𝐬𝐬𝐢𝐨𝐧 Finally, the Spark session is stopped to release resources.


1:22 # Databricks notebook source
2:56 Upload CSV file over Workspace
3:54 Databricks source
6:00 Show the number of students in the file
6:55 withcolomn in PySpark
7:38 schema Databricks notebook source
11:00 create custom data
20:11 lit command in PySpark
26:00 renamed multiple columns in single line using withColumnRenamed
27:55 alias name of any column
31:00 Filter rows as SQL query in PySpark
32:00 select * from student where course in ['DB', 'Cloud','OOP'] is in method
35:00 select * from student where course in ['DB', 'Cloud','OOP']
36:00 course_value= ['DB', 'Cloud','OOP']
39:00 In SQL like operators
41:11 Search course with particular String Pattern
44:11 startswith in PySpark
44:33 endswith in PySpark
46:10 contains in PySpark
47:11 df.filter(df.name.like('%s%e%')).show() in PySpark


На этой странице сайта вы можете посмотреть видео онлайн Data Engineer PySpark Data Bricks Session Day 6 длительностью часов минут секунд в хорошем качестве, которое загрузил пользователь Code with Kristi 20 Октябрь 2024, поделитесь ссылкой с друзьями и знакомыми, на youtube это видео уже посмотрели 274 раз и оно понравилось 5 зрителям. Приятного просмотра!