Data Engineer PySpark Data Bricks Session Day 6

Published: 20 October 2024
on channel: Code with Kristi
274
5

📍 𝐒𝐭𝐚𝐫𝐭 𝐚 𝐒𝐩𝐚𝐫𝐤 𝐒𝐞𝐬𝐬𝐢𝐨𝐧 : Set up the PySpark environment.
🧣 𝐂𝐫𝐞𝐚𝐭𝐞 𝐚 𝐋𝐢𝐬𝐭 : Define the list with three elements.
📢 𝐏𝐚𝐫𝐚𝐥𝐥𝐞𝐥𝐢𝐳𝐞 𝐭𝐡𝐞 𝐋𝐢𝐬𝐭 : Distribute the list across the cluster nodes.
🔔 𝐂𝐨𝐧𝐯𝐞𝐫𝐭 𝐭𝐨 𝐃𝐚𝐭𝐚𝐅𝐫𝐚𝐦𝐞 : Convert the distributed RDD to a DataFrame.
🔋 𝐏𝐞𝐫𝐟𝐨𝐫𝐦 𝐏𝐫𝐨𝐜𝐞𝐬𝐬𝐢𝐧𝐠 : Show the contents and perform any desired operations.

📍 This video will explain how to write first program in PySpark.
Don't miss out on this opportunity to excel!
🚀 *Course:* Master Azure Data Engineering
📅 *Last Date:* 15 Jan 2025

Course Registration: https://tinyurl.com/5n7aatdm

Don't miss out on this opportunity to upscale your skills and dive deep into the realm of data engineering! Reserve your spot now! 🎉

#hiring #career #databricks #azure #career #hiring #databricksanalytics


📢 Video Link:    • Data Engineer PySpark Data Bricks Session ...  

LinkedIn Profile of author:
  / sachin-saxena-graphic-designer  

Code Source Link:
https://lnkd.in/g67a4kY3

𝐄𝐱𝐩𝐥𝐚𝐧𝐚𝐭𝐢𝐨𝐧 𝐨𝐟 𝐭𝐡𝐞 𝐂𝐨𝐝𝐞 :

𝟏. 𝐒𝐩𝐚𝐫𝐤 𝐒𝐞𝐬𝐬𝐢𝐨𝐧 : The SparkSession is created to provide an entry point for Spark functionality.
𝟐. 𝐋𝐢𝐬𝐭 𝐂𝐫𝐞𝐚𝐭𝐢𝐨𝐧 : A list of three elements is defined.
𝟑. 𝐏𝐚𝐫𝐚𝐥𝐥𝐞𝐥𝐢𝐳𝐞 : The list is parallelized with numSlices=3, which ensures that each element is assigned to a different partition in the RDD. This is how we can distribute it across the three nodes.
𝟒. 𝐂𝐨𝐧𝐯𝐞𝐫𝐭 𝐭𝐨 𝐃𝐚𝐭𝐚𝐅𝐫𝐚𝐦𝐞 : The RDD is mapped to a tuple format to convert it into a DataFrame. The column is named "element".
𝟓. 𝐃𝐢𝐬𝐩𝐥𝐚𝐲 𝐃𝐚𝐭𝐚𝐅𝐫𝐚𝐦𝐞 : The contents of the DataFrame are printed using df.show(), which will display each element as a separate row.
𝟔. 𝐂𝐨𝐮𝐧𝐭 : The total number of elements is counted and printed.
𝟕. 𝐅𝐮𝐫𝐭𝐡𝐞𝐫 𝐏𝐫𝐨𝐜𝐞𝐬𝐬𝐢𝐧𝐠 : An optional step is included to filter the DataFrame for elements containing "1" and display the result.
𝟖. 𝐒𝐭𝐨𝐩 𝐒𝐩𝐚𝐫𝐤 𝐒𝐞𝐬𝐬𝐢𝐨𝐧 Finally, the Spark session is stopped to release resources.


1:22 # Databricks notebook source
2:56 Upload CSV file over Workspace
3:54 Databricks source
6:00 Show the number of students in the file
6:55 withcolomn in PySpark
7:38 schema Databricks notebook source
11:00 create custom data
20:11 lit command in PySpark
26:00 renamed multiple columns in single line using withColumnRenamed
27:55 alias name of any column
31:00 Filter rows as SQL query in PySpark
32:00 select * from student where course in ['DB', 'Cloud','OOP'] is in method
35:00 select * from student where course in ['DB', 'Cloud','OOP']
36:00 course_value= ['DB', 'Cloud','OOP']
39:00 In SQL like operators
41:11 Search course with particular String Pattern
44:11 startswith in PySpark
44:33 endswith in PySpark
46:10 contains in PySpark
47:11 df.filter(df.name.like('%s%e%')).show() in PySpark


On this page of the site you can watch the video online Data Engineer PySpark Data Bricks Session Day 6 with a duration of hours minute second in good quality, which was uploaded by the user Code with Kristi 20 October 2024, share the link with friends and acquaintances, this video has already been watched 274 times on youtube and it was liked by 5 viewers. Enjoy your viewing!