← Back to Projects

At a Glance

Predicting Medical Insurance Costs with Deep Learning demo
01

Problem

Medical insurance costs in the US vary enormously — from under $2,000 to over $60,000 per year — and yet the factors driving that spread are predictable: smoking status, age, BMI, number of dependents, and region all have documented relationships with healthcare utilization. The challenge is that these relationships are nonlinear. A standard linear regression can approximate the overall trend but systematically fails at the extremes — precisely where the highest-cost individuals live. A neural network can learn these nonlinear interactions end-to-end, from raw features to dollar predictions, without requiring any manual feature engineering.

The objective was clear: build a Keras regression model that predicts individual charges from the six input features, with a well-balanced train/test performance indicating real generalization rather than memorization.


02

Solution

This project built a full end-to-end deep learning regression pipeline: one-hot encoding for categorical features, separate MinMaxScaler instances for features and the target variable, a Sequential Keras model with two hidden layers (64 → 32 neurons, ReLU activation, 20% Dropout), and early stopping to find the optimal epoch count automatically. The model was evaluated on a held-out validation set with both quantitative metrics (MSE, MAE) and a qualitative prediction-vs-actual comparison. Final diagnosis: well-balanced — test MAE was actually lower than training MAE (an expected effect of Dropout being active only during training), indicating clean generalization with no overfitting.


03

Skills Acquired


04

Deep Dive

Health insurance pricing is one of the clearest examples of machine learning applied to a real-world regression problem. Every insurer does some version of this calculation — how much will this person cost us? — and the features that drive the answer (smoking, age, BMI) are well understood. The interesting challenge is learning how they interact without hand-encoding those interactions.

The dataset has 1,338 records and 7 columns. No missing values. The target — charges — spans roughly $1,122 to $63,770. Smokers, on average, cost three to four times more than non-smokers. Age and BMI add additional nonlinear curvature to predictions.

A linear regression on this data would underfit the high-cost individuals — the tail of the distribution where interactions between smoking status, age, and BMI compound each other. A neural network learns these compound effects directly from data.

Why This Project?

This was Sprint 12 of my TripleTen AI and Machine Learning Bootcamp, the first sprint dedicated to neural networks and deep learning. After nine previous sprints focused on classical ML (decision trees, random forests, gradient boosting, linear models), this sprint introduced Keras as the framework for building trainable multi-layer networks — moving from feature engineering to architecture design.

Insurance cost prediction is a canonical regression problem that appears across actuarial science, healthtech, and insuretech. The skills applied here — sequential model design, dropout regularization, target normalization, and training stability analysis — transfer directly to time series forecasting, demand prediction, and any regression task where the relationship is nonlinear and the dataset is tabular.


What You'll Learn from This


Key Takeaways


The Dataset

A US health insurance dataset with 1,338 beneficiary records and no missing values. The six input features cover demographics, lifestyle, and family structure. The target variable is the annual medical charge billed by the insurer.

FeatureTypeDescription
ageNumericAge of primary beneficiary (18–64)
sexCategoricalGender (female / male)
bmiNumericBody mass index (kg/m²)
childrenNumericNumber of dependents covered (0–5)
smokerCategoricalSmoking status (yes / no)
regionCategoricalUS residential region (4 values)
chargesNumericTARGET — annual medical costs ($1,122–$63,770)

⚡ Try It — Interactive Chart

Min-to-max range with median markers, per numeric feature — generated with Amazon Quick
After one-hot encoding the three categorical columns, the input layer receives 8 features. The dataset is small (1,338 rows), which informed the conservative architecture choice: 2,689 total parameters keeps the model from overfitting a limited training set.

Model Architecture

The network is a Sequential regression model with two hidden layers. ReLU was selected for its effectiveness in deep networks and resistance to the vanishing gradient problem. Dropout at 0.2 after each hidden layer provides regularization — important given the small dataset size.

Input (8 features)
  ↓
Dense(64, ReLU) ← 576 params (8×64 + 64 biases)
  ↓
Dropout(0.2) ← silences 20% of neurons during training
  ↓
Dense(32, ReLU) ← 2,080 params (64×32 + 32 biases)
  ↓
Dropout(0.2) ← silences 20% of neurons during training
  ↓
Dense(1, linear) ← 33 params (32×1 + 1 bias) — regression output

Total trainable params: 2,689  |  Optimizer: Adam  |  Loss: MSE  |  Metric: MAE

My Process

Part 1

Data Loading & Exploration

Loaded the dataset and ran a first-pass inspection: shape, column types, statistical summary, and missing value counts. 1,338 rows, 7 columns, zero missing values — clean data, no imputation needed.

import pandas as pd
import matplotlib.pyplot as plt
from sklearn.model_selection import train_test_split
from sklearn.preprocessing   import MinMaxScaler
from sklearn.linear_model    import LinearRegression
from tensorflow.keras.models    import Sequential
from tensorflow.keras.layers    import Dense, Dropout
from tensorflow.keras.callbacks import EarlyStopping

df = pd.read_csv('insurance.csv')
print("Shape:", df.shape)        # (1338, 7)
print(df.isnull().sum())      # all zeros — no missing values

Part 2

Data Preprocessing

Three steps: encode categoricals, split the data (before scaling to prevent leakage), then normalize features and target separately. The train/test split was performed before fitting the scalers — a critical ordering that the reviewer explicitly flagged as correct.

# One-hot encode categorical features (sex, smoker, region)
data_ohe = pd.get_dummies(df, drop_first=True)

features = data_ohe.drop("charges", axis=1)
targets  = data_ohe["charges"]

# Split BEFORE scaling to prevent data leakage
features_train, features_valid, target_train, target_valid = train_test_split(
    features, targets, test_size=0.2, random_state=12345
)

# Scale features: fit on train, transform both
numerical_cols = ["age", "bmi", "children"]
scaler = MinMaxScaler()
features_train.loc[:, numerical_cols] = scaler.fit_transform(features_train[numerical_cols])
features_valid.loc[:, numerical_cols] = scaler.transform(features_valid[numerical_cols])

# Scale target separately (improves training stability)
target_scaler = MinMaxScaler()
target_train = target_scaler.fit_transform(target_train.values.reshape(-1, 1))
target_valid = target_scaler.transform(target_valid.values.reshape(-1, 1))

Part 3

Build the Neural Network

Designed the Sequential model architecture: two hidden layers with ReLU activation and Dropout(0.2), followed by a single linear output neuron for regression. Compiled with Adam optimizer, MSE loss, and MAE as the tracking metric.

model = Sequential()

# Hidden layer 1: 64 neurons, ReLU, 20% dropout
model.add(Dense(64, input_shape=(features_train.shape[1],), activation="relu"))
model.add(Dropout(0.2))

# Hidden layer 2: 32 neurons, ReLU, 20% dropout
model.add(Dense(32, activation="relu"))
model.add(Dropout(0.2))

# Output: 1 neuron, linear activation (regression)
model.add(Dense(1))

model.compile(optimizer="adam", loss="mse", metrics=["mae"])
model.summary()   # Total params: 2,689

⚡ Try It — Interactive Chart

Trainable parameters per layer, 2,689 total — generated with Amazon Quick

Part 4

Train with Early Stopping

Set a maximum of 100 epochs with a batch size of 32. EarlyStopping monitored validation loss with patience=10 and restore_best_weights=True — automatically reverting to the epoch with the lowest validation loss when training stalled. The model converged between epoch 45 and 97 depending on the run.

early_stop = EarlyStopping(
    monitor='val_loss',
    patience=10,
    restore_best_weights=True
)

history = model.fit(
    features_train, target_train,
    validation_data=(features_valid, target_valid),
    epochs=100,
    batch_size=32,
    callbacks=[early_stop]
)

Training and validation loss both declined together and plateaued cleanly — no divergence, no spikes. The validation loss (orange) stayed slightly below training loss throughout, a signature of Dropout being active only during training.

⚡ Try It — Interactive Chart

Training vs. validation MSE over 59 epochs — generated with Amazon Quick

⚡ Try It — Interactive Chart

Epoch 1 vs. final epoch loss, same training run — generated with Amazon Quick

Part 5

Evaluate & Diagnose

Evaluated the model on the validation set, then compared training and test metrics to diagnose underfitting, overfitting, or a well-balanced model.

test_loss, test_mae = model.evaluate(features_valid, target_valid)
print(f"Test Loss (MSE): {test_loss:.4f}")   # 0.0051
print(f"Test MAE:        {test_mae:.4f}")    # 0.0422

final_train_mae = history.history['mae'][-1]  # 0.0552

# Diagnosis logic
if   final_train_mae > 0.1 and test_mae > 0.1: print("UNDERFITTING")
elif test_mae > final_train_mae * 1.5:          print("OVERFITTING")
else:                                             print("WELL BALANCED")
MetricTrainingTestDiagnosis
Loss (MSE) 0.0073 0.0051 Test lower — dropout effect
MAE 0.0552 0.0422 WELL BALANCED ✓

The test MAE being lower than training MAE is expected and correct — Dropout randomly deactivates 20% of neurons during training, adding noise to the training loss. During inference, Dropout is disabled and all neurons contribute, producing cleaner, lower-error predictions.

⚡ Try It — Interactive Chart

Training vs. validation MAE over 59 epochs — generated with Amazon Quick

Part 6

Visualization & Qualitative Analysis

Two visualizations confirmed the model's behavior. The training history plot showed both curves declining together and plateauing cleanly after ~20 epochs. The predicted vs. actual scatter plot showed tight clustering along the perfect prediction line for low-cost individuals, with wider scatter in the USD 15,000–35,000 range — where high-cost outliers are underrepresented in training data.

# Inverse-transform predictions back to dollar scale
predictions       = model.predict(features_valid)
predictions_actual = target_scaler.inverse_transform(predictions)
target_actual      = target_scaler.inverse_transform(target_valid.reshape(-1, 1))

# Scatter: Predicted vs Actual
plt.scatter(target_actual, predictions_actual, alpha=0.5)
plt.plot([target_actual.min(), target_actual.max()],
         [target_actual.min(), target_actual.max()], 'r--', label='Perfect Prediction')
plt.title('Predicted vs Actual Insurance Charges')
plt.show()

⚡ Try It — Interactive Chart

Predicted vs. actual charges, 10 test samples — generated with Amazon Quick

⚡ Try It — Interactive Chart

Actual minus predicted charges, same 10 test samples — generated with Amazon Quick

Final Results

The model achieved consistent, low-error performance across multiple runs. Because training involves stochastic elements (weight initialization, dropout, batch shuffling), the exact epoch of convergence varies — but the final metrics remained stable:

MetricValueNotes
Test Loss (MSE) 0.0051 On normalized target (0–1 scale)
Test MAE 0.0422 Dropout disabled during eval
Training MAE 0.0552 Slightly higher due to active dropout
Convergence epoch 45 – 97 Varies per run; EarlyStopping adapts
Model diagnosis WELL BALANCED ✓ No underfitting, no overfitting
Reviewer status APPROVED — Iteration 1 ✓ No red-flag issues raised

Main Takeaways


What I Learned & Why It Matters to Employers

Sprint 12 was where classical ML thinking met deep learning — same data pipeline discipline (split before scale, no leakage), same evaluation rigor (train vs. test comparison, diagnosis), but now applied to a multi-layer trainable architecture. Understanding why Dropout makes test loss lower than training loss, and why normalizing the target matters for gradient stability, separates someone who can run a Keras tutorial from someone who can reason about training behavior. These are the intuitions that matter when a model isn't converging in production and you need to diagnose why.

Conclusion & Reflections

This project confirmed something I had suspected: the habits that make classical ML work — clean splits, preventing leakage, choosing metrics that reflect business goals, diagnosing the bias-variance tradeoff — transfer directly to deep learning. The framework changes; the thinking doesn't.

If I had more time, I would add inverse-transformed MAE reporting in dollar units (so the metric is interpretable without knowing the scaler's range), run a formal feature importance analysis to quantify whether smoker status dominates as expected, and experiment with batch normalization for more stable training across runs with different random seeds.

Project RequirementStatus
End-to-end Keras regression pipelineCOMPLETE ✓
Train/test split before scalingCORRECT — data leakage prevented ✓
Target variable normalizedYES — separate MinMaxScaler ✓
Early stopping implementedYES — patience=10, restore_best_weights ✓
Training history visualizationYES — loss + MAE over epochs ✓
Predicted vs. actual scatter plotYES — inverse-transformed to dollar scale ✓
Model diagnosisWELL BALANCED — approved Iteration 1 ✓

Want to Explore the Full Notebook?

The complete Keras pipeline — all 7 parts, training history, and prediction analysis — is on GitHub.