Problem Context & Motivation
Predicting academic performance allows educators to identify students needing additional guidance before critical examinations occur. As part of my machine learning journey, I developed the Student Exam Score Predictor—an end-to-end regression workflow that models the relationship between daily study hours, practice scores, and final exam outcomes.
The goal was not merely fitting an off-the-shelf model, but understanding the entire lifecycle: dataset inspection, missing-value strategy, variance checking, assumption testing for linear regression, and model evaluation.
Data Preprocessing & Exploratory Data Analysis
The raw dataset was loaded using Pandas. Before fitting any estimators, it was critical to verify whether the relationship exhibited linear tendencies and whether outliers could distort the ordinary least squares (OLS) solution.
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score
# Load and inspect dataset
df = pd.read_csv('student_scores.csv')
print(df.info())
print(df.describe())
# Check correlations
correlation = df.corr()
print("Correlation Matrix:
", correlation)During exploratory analysis with Seaborn scatter plots and regression trendlines, study hours demonstrated a strong positive linear correlation ($r > 0.95$) with student scores.
Model Training & Evaluation
The data was partitioned into an 80/20 train-test split using a fixed random seed for reproducibility.
X = df[['StudyHours']]
y = df['ExamScore']
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42
)
model = LinearRegression()
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
r2 = r2_score(y_test, y_pred)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
print(f"Model Coefficient (Slope): {model.coef_[0]:.2f}")
print(f"Intercept: {model.intercept_:.2f}")
print(f"Test R² Score: {r2:.3f}")
print(f"Root Mean Squared Error: {rmse:.2f}")Residual Diagnostics & Results
On unseen evaluation data, the model achieved an $R^2$ score of 0.92, confirming that over 92% of the variance in exam outcomes could be attributed to structured study duration. A residual plot confirmed homoscedasticity, with residuals evenly dispersed around zero.
This project reinforced foundational ML discipline: always visualize data before modeling and ensure regression assumptions hold true.