At a Glance
- Built a Keras regression network to predict annual medical insurance charges from demographics and lifestyle features
- Designed a 2-hidden-layer architecture (64→32 neurons, ReLU, Dropout), trained with early stopping — approved on first submission
- Python, TensorFlow/Keras, scikit-learn, Pandas, Matplotlib
Problem
Medical insurance costs in the US vary enormously — from under $2,000 to over $60,000 per year — and yet the factors driving that spread are predictable: smoking status, age, BMI, number of dependents, and region all have documented relationships with healthcare utilization. The challenge is that these relationships are nonlinear. A standard linear regression can approximate the overall trend but systematically fails at the extremes — precisely where the highest-cost individuals live. A neural network can learn these nonlinear interactions end-to-end, from raw features to dollar predictions, without requiring any manual feature engineering.
The objective was clear: build a Keras regression model that predicts individual charges from the six input features, with a well-balanced train/test performance indicating real generalization rather than memorization.
Solution
This project built a full end-to-end deep learning regression pipeline: one-hot encoding for categorical features, separate MinMaxScaler instances for features and the target variable, a Sequential Keras model with two hidden layers (64 → 32 neurons, ReLU activation, 20% Dropout), and early stopping to find the optimal epoch count automatically. The model was evaluated on a held-out validation set with both quantitative metrics (MSE, MAE) and a qualitative prediction-vs-actual comparison. Final diagnosis: well-balanced — test MAE was actually lower than training MAE (an expected effect of Dropout being active only during training), indicating clean generalization with no overfitting.
Skills Acquired
- TensorFlow / Keras — used
Sequential,Dense, andDropoutlayers to build a regression neural network;EarlyStoppingcallback to halt training when validation loss stopped improving. - Python — end-to-end implementation language for data loading, preprocessing, model training, evaluation, and visualization.
- Pandas — data loading, shape inspection, and applying
get_dummiesfor one-hot encoding of categorical features (sex,smoker,region). - scikit-learn —
train_test_splitfor the 80/20 split before any scaling, andMinMaxScalerapplied separately to features and the target variable. - Matplotlib — training history plots (training vs. validation loss over epochs) and a predicted vs. actual scatter plot to qualitatively verify predictions.
- Neural network design — selecting layer sizes, activation functions (ReLU), dropout rate, optimizer (Adam), loss function (MSE), and evaluation metric (MAE) for a regression task on a small tabular dataset.
- Data leakage prevention — train/test split was performed before fitting the scalers, ensuring no information from the validation set influenced preprocessing.
Deep Dive
Health insurance pricing is one of the clearest examples of machine learning applied to a real-world regression problem. Every insurer does some version of this calculation — how much will this person cost us? — and the features that drive the answer (smoking, age, BMI) are well understood. The interesting challenge is learning how they interact without hand-encoding those interactions.
The dataset has 1,338 records and 7 columns. No missing values. The target — charges — spans roughly $1,122 to $63,770. Smokers, on average, cost three to four times more than non-smokers. Age and BMI add additional nonlinear curvature to predictions.
Why This Project?
This was Sprint 12 of my TripleTen AI and Machine Learning Bootcamp, the first sprint dedicated to neural networks and deep learning. After nine previous sprints focused on classical ML (decision trees, random forests, gradient boosting, linear models), this sprint introduced Keras as the framework for building trainable multi-layer networks — moving from feature engineering to architecture design.
Insurance cost prediction is a canonical regression problem that appears across actuarial science, healthtech, and insuretech. The skills applied here — sequential model design, dropout regularization, target normalization, and training stability analysis — transfer directly to time series forecasting, demand prediction, and any regression task where the relationship is nonlinear and the dataset is tabular.
What You'll Learn from This
- Why the train/test split must happen before fitting the scaler — and what data leakage looks like when it doesn't
- Why normalizing the target variable (not just features) improves training stability for regression networks
- How Dropout works during training vs. inference — and why test loss can be lower than training loss
- How EarlyStopping with
restore_best_weights=Truefinds the optimal epoch without manual tuning - How to read a training history plot to diagnose underfitting, overfitting, or a clean convergence
Key Takeaways
- Model converged in 45–97 epochs (varies by run) — early stopping with patience=10 found the right stopping point automatically every time
- Test MSE: 0.0051, Test MAE: 0.0422 on the normalized target — lower than training metrics due to Dropout being disabled during evaluation
- Well-balanced diagnosis — no underfitting, no overfitting, training and test metrics consistently close
- Predictions in realistic dollar ranges (USD 1,000–50,000) — scatter plot showed tight clustering for low-cost individuals, wider spread for high-cost outliers (a known challenge with underrepresented extreme cases)
- Initial architecture was sufficient — the first design required no iteration, approved on first submission
The Dataset
A US health insurance dataset with 1,338 beneficiary records and no missing values. The six input features cover demographics, lifestyle, and family structure. The target variable is the annual medical charge billed by the insurer.
| Feature | Type | Description |
|---|---|---|
age | Numeric | Age of primary beneficiary (18–64) |
sex | Categorical | Gender (female / male) |
bmi | Numeric | Body mass index (kg/m²) |
children | Numeric | Number of dependents covered (0–5) |
smoker | Categorical | Smoking status (yes / no) |
region | Categorical | US residential region (4 values) |
charges | Numeric | TARGET — annual medical costs ($1,122–$63,770) |
⚡ Try It — Interactive Chart
Model Architecture
The network is a Sequential regression model with two hidden layers. ReLU was selected for its effectiveness in deep networks and resistance to the vanishing gradient problem. Dropout at 0.2 after each hidden layer provides regularization — important given the small dataset size.
↓
Dense(64, ReLU) ← 576 params (8×64 + 64 biases)
↓
Dropout(0.2) ← silences 20% of neurons during training
↓
Dense(32, ReLU) ← 2,080 params (64×32 + 32 biases)
↓
Dropout(0.2) ← silences 20% of neurons during training
↓
Dense(1, linear) ← 33 params (32×1 + 1 bias) — regression output
Total trainable params: 2,689 | Optimizer: Adam | Loss: MSE | Metric: MAE
My Process
Part 1
Data Loading & Exploration
Loaded the dataset and ran a first-pass inspection: shape, column types, statistical summary, and missing value counts. 1,338 rows, 7 columns, zero missing values — clean data, no imputation needed.
import pandas as pd import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.preprocessing import MinMaxScaler from sklearn.linear_model import LinearRegression from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Dropout from tensorflow.keras.callbacks import EarlyStopping df = pd.read_csv('insurance.csv') print("Shape:", df.shape) # (1338, 7) print(df.isnull().sum()) # all zeros — no missing values
Part 2
Data Preprocessing
Three steps: encode categoricals, split the data (before scaling to prevent leakage), then normalize features and target separately. The train/test split was performed before fitting the scalers — a critical ordering that the reviewer explicitly flagged as correct.
# One-hot encode categorical features (sex, smoker, region) data_ohe = pd.get_dummies(df, drop_first=True) features = data_ohe.drop("charges", axis=1) targets = data_ohe["charges"] # Split BEFORE scaling to prevent data leakage features_train, features_valid, target_train, target_valid = train_test_split( features, targets, test_size=0.2, random_state=12345 ) # Scale features: fit on train, transform both numerical_cols = ["age", "bmi", "children"] scaler = MinMaxScaler() features_train.loc[:, numerical_cols] = scaler.fit_transform(features_train[numerical_cols]) features_valid.loc[:, numerical_cols] = scaler.transform(features_valid[numerical_cols]) # Scale target separately (improves training stability) target_scaler = MinMaxScaler() target_train = target_scaler.fit_transform(target_train.values.reshape(-1, 1)) target_valid = target_scaler.transform(target_valid.values.reshape(-1, 1))
Part 3
Build the Neural Network
Designed the Sequential model architecture: two hidden layers with ReLU activation and Dropout(0.2), followed by a single linear output neuron for regression. Compiled with Adam optimizer, MSE loss, and MAE as the tracking metric.
model = Sequential() # Hidden layer 1: 64 neurons, ReLU, 20% dropout model.add(Dense(64, input_shape=(features_train.shape[1],), activation="relu")) model.add(Dropout(0.2)) # Hidden layer 2: 32 neurons, ReLU, 20% dropout model.add(Dense(32, activation="relu")) model.add(Dropout(0.2)) # Output: 1 neuron, linear activation (regression) model.add(Dense(1)) model.compile(optimizer="adam", loss="mse", metrics=["mae"]) model.summary() # Total params: 2,689
⚡ Try It — Interactive Chart
Part 4
Train with Early Stopping
Set a maximum of 100 epochs with a batch size of 32. EarlyStopping monitored validation loss with patience=10 and restore_best_weights=True — automatically reverting to the epoch with the lowest validation loss when training stalled. The model converged between epoch 45 and 97 depending on the run.
early_stop = EarlyStopping( monitor='val_loss', patience=10, restore_best_weights=True ) history = model.fit( features_train, target_train, validation_data=(features_valid, target_valid), epochs=100, batch_size=32, callbacks=[early_stop] )
Training and validation loss both declined together and plateaued cleanly — no divergence, no spikes. The validation loss (orange) stayed slightly below training loss throughout, a signature of Dropout being active only during training.
⚡ Try It — Interactive Chart
⚡ Try It — Interactive Chart
Part 5
Evaluate & Diagnose
Evaluated the model on the validation set, then compared training and test metrics to diagnose underfitting, overfitting, or a well-balanced model.
test_loss, test_mae = model.evaluate(features_valid, target_valid) print(f"Test Loss (MSE): {test_loss:.4f}") # 0.0051 print(f"Test MAE: {test_mae:.4f}") # 0.0422 final_train_mae = history.history['mae'][-1] # 0.0552 # Diagnosis logic if final_train_mae > 0.1 and test_mae > 0.1: print("UNDERFITTING") elif test_mae > final_train_mae * 1.5: print("OVERFITTING") else: print("WELL BALANCED")
| Metric | Training | Test | Diagnosis |
|---|---|---|---|
| Loss (MSE) | 0.0073 | 0.0051 | Test lower — dropout effect |
| MAE | 0.0552 | 0.0422 | WELL BALANCED ✓ |
The test MAE being lower than training MAE is expected and correct — Dropout randomly deactivates 20% of neurons during training, adding noise to the training loss. During inference, Dropout is disabled and all neurons contribute, producing cleaner, lower-error predictions.
⚡ Try It — Interactive Chart
Part 6
Visualization & Qualitative Analysis
Two visualizations confirmed the model's behavior. The training history plot showed both curves declining together and plateauing cleanly after ~20 epochs. The predicted vs. actual scatter plot showed tight clustering along the perfect prediction line for low-cost individuals, with wider scatter in the USD 15,000–35,000 range — where high-cost outliers are underrepresented in training data.
# Inverse-transform predictions back to dollar scale predictions = model.predict(features_valid) predictions_actual = target_scaler.inverse_transform(predictions) target_actual = target_scaler.inverse_transform(target_valid.reshape(-1, 1)) # Scatter: Predicted vs Actual plt.scatter(target_actual, predictions_actual, alpha=0.5) plt.plot([target_actual.min(), target_actual.max()], [target_actual.min(), target_actual.max()], 'r--', label='Perfect Prediction') plt.title('Predicted vs Actual Insurance Charges') plt.show()
⚡ Try It — Interactive Chart
⚡ Try It — Interactive Chart
Final Results
The model achieved consistent, low-error performance across multiple runs. Because training involves stochastic elements (weight initialization, dropout, batch shuffling), the exact epoch of convergence varies — but the final metrics remained stable:
| Metric | Value | Notes |
|---|---|---|
| Test Loss (MSE) | 0.0051 | On normalized target (0–1 scale) |
| Test MAE | 0.0422 | Dropout disabled during eval |
| Training MAE | 0.0552 | Slightly higher due to active dropout |
| Convergence epoch | 45 – 97 | Varies per run; EarlyStopping adapts |
| Model diagnosis | WELL BALANCED ✓ | No underfitting, no overfitting |
| Reviewer status | APPROVED — Iteration 1 ✓ | No red-flag issues raised |
Main Takeaways
- Split before you scale. The train/test split must come before fitting any scaler. Fitting on the full dataset leaks validation distribution information into training — a subtle but critical form of data leakage.
- Normalize the target, not just the features. Charges range from $1,122 to $63,770 — a span that would make gradient updates unstable. Scaling to [0,1] stabilizes training significantly.
- Dropout explains the "paradox" of lower test loss. Dropout is a training-only regularizer. It deactivates neurons stochastically during training, adding noise to training metrics. At inference, all neurons are active — so test performance is cleaner.
- EarlyStopping replaces manual epoch tuning.
restore_best_weights=Truemeans the model reverts to its best checkpoint automatically — no post-hoc selection required. - Small datasets reward small models. 2,689 parameters for 1,070 training rows keeps the model from memorizing the training set. The conservative architecture was the right call.
What I Learned & Why It Matters to Employers
Sprint 12 was where classical ML thinking met deep learning — same data pipeline discipline (split before scale, no leakage), same evaluation rigor (train vs. test comparison, diagnosis), but now applied to a multi-layer trainable architecture. Understanding why Dropout makes test loss lower than training loss, and why normalizing the target matters for gradient stability, separates someone who can run a Keras tutorial from someone who can reason about training behavior. These are the intuitions that matter when a model isn't converging in production and you need to diagnose why.
Conclusion & Reflections
This project confirmed something I had suspected: the habits that make classical ML work — clean splits, preventing leakage, choosing metrics that reflect business goals, diagnosing the bias-variance tradeoff — transfer directly to deep learning. The framework changes; the thinking doesn't.
If I had more time, I would add inverse-transformed MAE reporting in dollar units (so the metric is interpretable without knowing the scaler's range), run a formal feature importance analysis to quantify whether smoker status dominates as expected, and experiment with batch normalization for more stable training across runs with different random seeds.
| Project Requirement | Status |
|---|---|
| End-to-end Keras regression pipeline | COMPLETE ✓ |
| Train/test split before scaling | CORRECT — data leakage prevented ✓ |
| Target variable normalized | YES — separate MinMaxScaler ✓ |
| Early stopping implemented | YES — patience=10, restore_best_weights ✓ |
| Training history visualization | YES — loss + MAE over epochs ✓ |
| Predicted vs. actual scatter plot | YES — inverse-transformed to dollar scale ✓ |
| Model diagnosis | WELL BALANCED — approved Iteration 1 ✓ |
Want to Explore the Full Notebook?
The complete Keras pipeline — all 7 parts, training history, and prediction analysis — is on GitHub.