Data Science
•8 min read•298 words

What I Learned from Exploratory Data Analysis on the Titanic Dataset

Comprehensive exploratory data analysis examining survival dynamics, imputation strategies for missing demographics, and multidimensional categorical encoding.

Md Adil Iftekhar
Md Adil IftekharB.Tech CS Student at JBIT | Data Science & ML Enthusiast

Why the Titanic Dataset Remains a Gold Standard for EDA

The Kaggle Titanic dataset is often a developer's first foray into structured data science. While seemingly compact, it provides exceptional real-world complexity: missing age distributions, categorical embarkation points, high-cardinality ticket labels, and nested family relationships.

In this project, my objective was to extract actionable survival patterns and establish clean data preprocessing pipelines.

Handling Missing Values & Feature Engineering

Approximately 20% of the Age column and over 75% of the Cabin column were missing. Rather than dropping rows indiscriminately, systematic imputation was implemented:

  • Title Extraction: Extracted honorific titles (Mr., Mrs., Miss., Master., Dr.) from passenger names.
  • Median Age Imputation: Imputed missing ages using the median age of the passenger's specific title group.
  • Family Size Feature: Combined SibSp (siblings/spouses) and Parch (parents/children) to calculate overall family size.
  • python
    import pandas as pd
    import seaborn as sns
    import matplotlib.pyplot as plt
    
    df = pd.read_csv('titanic.csv')
    
    # Extract title from passenger name
    df['Title'] = df['Name'].str.extract(' ([A-Za-z]+)\.', expand=False)
    title_age_medians = df.groupby('Title')['Age'].median()
    
    # Impute age based on title
    df['Age'] = df.apply(
        lambda row: title_age_medians[row['Title']] if pd.isnull(row['Age']) else row['Age'],
        axis=1
    )
    
    # Engineer FamilySize
    df['FamilySize'] = df['SibSp'] + df['Parch'] + 1
    df['IsAlone'] = (df['FamilySize'] == 1).astype(int)

    Key Visual Discoveries

    Using Seaborn boxplots, heatmaps, and faceted bar plots, key demographic drivers emerged:

  • Gender Disparity: Female passengers achieved an overall survival rate of approximately 74%, compared to just 19% for male passengers.
  • Class Stratification: Passengers in 1st class survived at a rate of 63%, whereas 3rd class passengers experienced only 24% survival.
  • Solo Travelers: Solo travelers in 3rd class suffered the highest mortality rate.
  • This exercise proved that deep exploratory data analysis and domain-aware feature engineering are more influential in predictive workflows than algorithm selection alone.

    Related Topics:#Data Science#EDA#Pandas#Seaborn#Data Visualization#Python