Data Analyst Portfolio Project (Exploratory Data Analysis With Python Pandas)

Publicado em: 10 Julho 2023
no canal de: Ryan & Matt Data Science
99,582
3.3k

🧠 Don’t miss out! Get FREE access to my Skool community — packed with resources, tools, and support to help you with Data, Machine Learning, and AI Automations! 📈 https://www.skool.com/data-and-ai-aut...

In this video, we take a look at an Exploratory Data Analysis (EDA) portfolio project within Python Pandas. Everything is coded within Jupyter Notebook and the data is sourced from Kaggle.

Python Libraries needed: Pandas, Seaborn

Kaggle Data: https://www.kaggle.com/datasets/aiaia...


🚀 Hire me for Data Work: https://ryanandmattdatascience.com/da...
👨‍💻 Mentorships: https://ryanandmattdatascience.com/me...
📧 Email: ryannolandata@gmail.com
🌐 Website & Blog: https://ryanandmattdatascience.com/
🖥️ Discord:   / discord  
📚 *Practice SQL & Python Interview Questions: https://stratascratch.com/?via=ryan
📖 *SQL and Python Courses: https://datacamp.pxf.io/XYD7Qg

🍿 WATCH NEXT
Python Pandas Playlist:    • Python Pandas for Beginners  
Python Data Cleaning:    • Real World Data Cleaning in Python Pandas ...  
Python Pandas Groupby:    • The Complete Guide to Python Pandas Groupby  
Python Pandas Melt:    • Python Pandas Melt Tutorial: Transform Wid...  

In this Python data analytics portfolio project, I analyze over 7 million ultra marathon race records from 1798 to 2022 using exploratory data analysis techniques. We work with a massive dataset from Kaggle, cleaning and transforming the data to answer interesting questions about ultra running performance across different distances, ages, genders, and seasons.

Starting with data cleaning and preparation in Jupyter Notebook, I demonstrate how to filter 26,000 USA race records from 2020 for 50k and 50-mile events. Throughout the project, I use pandas for data manipulation, apply lambda functions for creating season categories from race dates, and leverage Seaborn for data visualization with histogram plots, distribution plots, violin plots, and lm plots.

Key analysis includes comparing male versus female performance differences, identifying optimal age groups for ultra running (spoiler: 29 is peak performance), examining how race seasons impact speed (summer races are significantly slower), and even locating my own race results in the dataset. I also demonstrate advanced pandas techniques like query functions, group by operations, string manipulation, and multi-level filtering to extract meaningful insights from messy real-world data.

Whether you're building your data analyst portfolio or want to learn practical Python data analysis skills, this project shows you how to handle large datasets, perform data cleaning, create insightful visualizations, and answer business questions through exploratory data analysis. All code is available on GitHub for you to practice with.

TIMESTAMPS
00:00 Introduction & Project Overview
02:07 Data Source & Download from Kaggle
03:50 Importing Data into Jupyter Notebook
05:17 Importing Libraries & Creating DataFrame
07:32 Initial Data Exploration & Understanding
10:20 Identifying Race Distance Formats
13:00 Filtering Data (2020, USA, 50K/50Mi)
17:36 Grabbing USA Events with String Split
20:05 Cleaning Data - Removing USA from Event Names
21:52 Creating Athlete Age Column
23:55 Handling Null Values & Data Types
26:00 Dropping Unnecessary Columns
28:00 Renaming Columns for Better Analysis
32:00 Reordering Columns
34:02 Finding Personal Race Results
38:20 Creating Visualizations with Seaborn
42:00 Violin Plots - Gender & Speed Analysis
44:15 LM Plots - Age vs Speed
46:20 Group By Analysis - Gender & Distance
48:30 Age Group Performance Analysis
50:20 Creating Race Season Column with Lambda
54:00 Season Performance Analysis
56:00 Final Analysis & Conclusions

OTHER SOCIALS:
Ryan’s LinkedIn:   / ryan-p-nolan  
Matt’s LinkedIn:   / matt-payne-ceo  
Twitter/X: https://x.com/RyanMattDS

Who is Ryan
Ryan is a Data Scientist at a fintech company, where he focuses on fraud prevention in underwriting and risk. Before that, he worked as a Data Analyst at a tax software company. He holds a degree in Electrical Engineering from UCF.

Who is Matt
Matt is the founder of Width.ai, an AI and Machine Learning agency. Before starting his own company, he was a Machine Learning Engineer at Capital One.

*This is an affiliate program. We receive a small portion of the final sale at no extra cost to you.


Nesta página do site você pode assistir ao vídeo on-line Data Analyst Portfolio Project (Exploratory Data Analysis With Python Pandas) duração hora minuto segundo em boa qualidade , que foi baixado pelo usuário Ryan & Matt Data Science 10 Julho 2023, compartilhe o link com seus amigos e conhecidos, no youtube este vídeo já foi visto 99,582 vezes e gostou 3.3 mil espectadores. Boa visualização!