← Back to Projects

At a Glance

End-to-End Machine Learning on AWS with S3 and EC2 demo
01

Problem

In most ML tutorials, everything runs on your laptop. The dataset is local. The model is local. If your machine crashes, you lose everything. That setup breaks down the moment you need to train on data too large for one disk, collaborate with a team, or deploy predictions to an application. The real ML workflow — the one used in production — looks different: data lives in durable cloud storage, training happens on elastic compute, and the trained model is persisted independently of any single server.

The challenge here was to build that workflow from scratch using AWS. Starting from nothing — no data, no server, no permissions — the goal was to arrive at a trained diabetes prediction model stored in S3, accessible from anywhere, independent of whether the EC2 instance is even running.


02

Solution

The pipeline wires together four AWS services into a complete ML workflow: S3 for durable object storage (dataset in, model out), IAM for credential-free access between services, VPC for a private, internet-connected network, and EC2 for on-demand compute. The training itself — a scikit-learn Random Forest classifier on the Pima Indians Diabetes dataset — runs entirely on the cloud server via the AWS CLI. No local training. No manual file transfers. Every piece is wired by AWS services.

S3 Bucket (vietnguyen-ml-project-3421)
  ├── datasets/ diabetes.csv ← source data uploaded first
  └── models/ diabetes_model.pkl ← trained model saved here

IAM Role EC2-S3-Access-Role (AmazonS3FullAccess)
  ↓ attached to
VPC my-project-vpc (10.0.0.0/16)
  └── Public Subnet my-project-subnet-public1-us-east-1a
       └── EC2 ml-training-instance (t2.micro · Amazon Linux 2023)
            ├── Downloads diabetes.csv from S3
            ├── Trains RandomForestClassifier
            └── Uploads diabetes_model.pkl back to S3

Internet Gateway + Route Table connect EC2 to the internet

03

Skills Acquired


04

Deep Dive

The Pima Indians Diabetes dataset is a classic binary classification benchmark: 768 patient records, 8 clinical features, and a binary target indicating diabetes diagnosis. It's small enough to run anywhere — but training it on EC2 with data sourced from S3 is the point. The dataset is just the payload. The real lesson is the infrastructure: how do you move data securely between services? How does a server prove it has permission to read from a storage bucket without a password?

The IAM role is what makes this secure. Rather than embedding AWS credentials in code (a well-known anti-pattern that leaks secrets into version control), the EC2 instance is assigned a role at launch time. AWS automatically provides temporary, rotated credentials to the instance through its metadata service. The application — in this case, the AWS CLI — retrieves those credentials transparently. The result: full S3 access with zero credentials in code.

This is the same pattern used in every production AWS ML pipeline. Whether you're training on a t2.micro or a p4d.24xlarge GPU instance, the architecture is identical: an IAM role grants the instance access to S3, the instance pulls data, trains, and pushes the model artifact back. Swapping the instance type is a one-line change. The security and data flow don't change at all.


Why This Project?

This was Sprint 13 of the AWS Cloud Institute program — the capstone lab that ties together S3, IAM, VPC, and EC2 into a single coherent workflow. After individual labs on each service in isolation, this sprint required wiring them all together: storage, permissions, networking, and compute operating as one system.

The skills built here — S3 data management, IAM role configuration, VPC setup, EC2 provisioning — are foundational to any cloud-deployed ML system. Understanding how these services interact is what separates someone who can run a Jupyter notebook from someone who can deploy and maintain an ML system that actually stays running.


What You'll Learn from This


Key Takeaways


The Dataset

The Pima Indians Diabetes dataset is a widely used binary classification benchmark from the UCI Machine Learning Repository. It contains 768 patient records, all female patients of Pima Indian heritage, with 8 clinical features and a binary outcome variable.

FeatureTypeDescription
PregnanciesNumericNumber of times pregnant
GlucoseNumericPlasma glucose concentration (2-hour oral glucose tolerance test)
BloodPressureNumericDiastolic blood pressure (mm Hg)
SkinThicknessNumericTriceps skin fold thickness (mm)
InsulinNumeric2-hour serum insulin (mu U/ml)
BMINumericBody mass index (kg/m²)
DiabetesPedigreeFunctionNumericDiabetes pedigree function (genetic risk score)
AgeNumericAge in years
OutcomeBinaryTARGET — 0: no diabetes, 1: diabetes (268 positive / 500 negative)

My Process

Phase 1

Create the S3 Bucket & Folder Structure

Created an S3 bucket named vietnguyen-ml-project-3421 and organized it with two folders: datasets/ for raw training data and models/ for trained model artifacts. Good folder structure in S3 matters — in a real ML system, you'd also add versioned subfolders for experiment tracking.

Phase 2

Upload the Dataset to S3

Uploaded diabetes.csv to s3://vietnguyen-ml-project-3421/datasets/. The data is now in durable cloud storage — accessible by any authorized AWS service in any region, independent of any compute instance.

Phase 3

Create the IAM Role for EC2

Created EC2-S3-Access-Role in IAM: trusted entity type AWS service, use case EC2, permission AmazonS3FullAccess. This role will be attached to the EC2 instance, granting it transparent S3 access through AWS's internal credential vending mechanism — no keys in code.

Phase 4

Create the VPC

Used the "VPC and more" wizard to create my-project-vpc (CIDR: 10.0.0.0/16) with 1 Availability Zone, 1 public subnet, an Internet Gateway (my-project-igw), and a route table (my-project-rtb-public) with a 0.0.0.0/0 → IGW route. No NAT gateway, no private subnets — the EC2 instance is public-facing for SSH access.

Phase 5

Launch EC2 with IAM Role Attached

Launched a t2.micro instance running Amazon Linux 2023, placed inside my-project-vpc with auto-assign public IP enabled. Under Advanced Details, attached EC2-S3-Access-Role as the IAM instance profile. This single step is what gives the instance permission to access S3 — without it, every aws s3 command returns Access Denied.

Phase 6

Connect via EC2 Instance Connect & Verify S3 Access

Connected to the running instance via EC2 Instance Connect in the browser — no SSH key required. Verified S3 access with:

# List bucket contents — confirms IAM role is working
aws s3 ls s3://vietnguyen-ml-project-3421/

# Expected output:
#   PRE datasets/
#   PRE models/

Installed Python dependencies on the fresh instance:

sudo dnf install -y python3 python3-pip
pip3 install --upgrade pip
pip3 install pandas scikit-learn

Phase 7

Download Dataset from S3 to EC2

Created a working directory structure and pulled the dataset from S3 to the instance's local storage:

mkdir -p ~/ml-project/data ~/ml-project/models

aws s3 cp s3://vietnguyen-ml-project-3421/datasets/diabetes.csv \
  ~/ml-project/data/

# Verify — should show diabetes.csv
ls ~/ml-project/data/
head ~/ml-project/data/diabetes.csv

Phase 8

Train the Random Forest Model

Created train.py on the EC2 instance — a complete ML pipeline: load data, prepare features and target, split 80/20, train a Random Forest, evaluate, print feature importances, and save the model with pickle:

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, classification_report
import pickle

# Load dataset
df = pd.read_csv('/home/ec2-user/ml-project/data/diabetes.csv')
print(f"Dataset shape: {df.shape[0]} rows, {df.shape[1]} columns")

# Features and target
X = df.iloc[:, :-1]   # all columns except Outcome
y = df.iloc[:,  -1]   # Outcome column

# 80/20 split
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

# Train Random Forest
model = RandomForestClassifier(
    n_estimators=100, max_depth=10, random_state=42
)
model.fit(X_train, y_train)

# Evaluate
y_pred   = model.predict(X_test)
accuracy = accuracy_score(y_test, y_pred)
print(f"Test Accuracy: {accuracy:.2%}")
print(classification_report(y_test, y_pred,
      target_names=['No Diabetes', 'Diabetes']))

# Feature importance rankings
importances   = model.feature_importances_
feature_names = df.columns[:-1].tolist()
for name, imp in sorted(zip(feature_names, importances),
                         key=lambda x: x[1], reverse=True):
    print(f"  {name}: {imp:.4f}")

# Save model
model_path = '/home/ec2-user/ml-project/models/diabetes_model.pkl'
with open(model_path, 'wb') as f:
    pickle.dump(model, f)
print(f"Model saved to: {model_path}")

Ran the training script directly on the cloud server:

python3 ~/ml-project/train.py

⚡ Try It — Interactive Chart

Illustrative feature importance rankings, typical of a Random Forest on this dataset — not the measured output of this training run — generated with Amazon Quick

⚡ Try It — Interactive Chart

Illustrative precision, recall, and F1-score, typical for this dataset and hyperparameters — not the measured output of this training run — generated with Amazon Quick

Phase 9

Upload Trained Model Back to S3

With the model serialized locally, one aws s3 cp command moved it to permanent storage. The model now lives in S3 independently of the EC2 instance — the instance can be terminated, the model stays:

aws s3 cp ~/ml-project/models/diabetes_model.pkl \
  s3://vietnguyen-ml-project-3421/models/

# Verify from EC2 — confirms model landed in S3
aws s3 ls s3://vietnguyen-ml-project-3421/models/
# Output: 2026-05-21 ... 646390 diabetes_model.pkl

  Auto-Grader · Full Pipeline Check (VPC → IAM → EC2 → SSH → S3) · us-east-1
1) VPC RESOURCES
PASS VPC exists (my-project-vpc) — vpc-0985530cde96d5c6c
PASS CIDR is 10.0.0.0/16
PASS Tenancy is default
PASS Exactly 1 AZ and 1 subnet
PASS Subnet: my-project-subnet-public1-us-east-1a
PASS IGW my-project-igw exists and is attached
PASS RTB my-project-rtb-public has 0.0.0.0/0 → IGW route
PASS No NAT gateways / No VPC endpoints

2) IAM ROLE
PASS Role exists: EC2-S3-Access-Role
PASS Trust policy includes ec2.amazonaws.com
PASS AmazonS3FullAccess attached

3) EC2 INSTANCE
PASS Instance: ml-training-instance (i-032d9334018728a23)
PASS State: running · Type: t2.micro · AMI: Amazon Linux 2023
PASS No key pair · In my-project-vpc · Public IP assigned
PASS IAM profile attached: EC2-S3-Access-Role

4) SSH CHECKS (EC2 Instance Connect)
PASS SSH connected and checks executed
PASS python3 · pip3 · pandas · scikit-learn all available
PASS ~/ml-project/data/diabetes.csv exists
PASS ~/ml-project/train.py exists
PASS ~/ml-project/models/diabetes_model.pkl exists

5) S3 BUCKET + UPLOADS
PASS Bucket: vietnguyen-ml-project-3421
PASS datasets/ and models/ prefixes exist
PASS datasets/diabetes.csv uploaded
PASS models/diabetes_model.pkl — Size: 646,390 bytes

══════════════════════════════
Raw points: 274 / 274 · Percentage: 100%

⚡ Try It — Interactive Chart

274/274 checks passed across all four pipeline layers — generated with Amazon Quick

Final Results

CheckStatusDetails
VPC created and configuredPASS ✓my-project-vpc · 10.0.0.0/16 · 1 AZ · 1 public subnet
IAM role with S3 permissionsPASS ✓EC2-S3-Access-Role · AmazonS3FullAccess
EC2 instance runningPASS ✓ml-training-instance · t2.micro · Amazon Linux 2023
SSH via EC2 Instance ConnectPASS ✓Connected and checks executed
Python environmentPASS ✓python3 · pip3 · pandas · scikit-learn all available
Dataset on EC2PASS ✓~/ml-project/data/diabetes.csv
Training script on EC2PASS ✓~/ml-project/train.py
Trained model on EC2PASS ✓~/ml-project/models/diabetes_model.pkl
Dataset in S3PASS ✓s3://vietnguyen-ml-project-3421/datasets/diabetes.csv
Trained model in S3PASS ✓s3://…/models/diabetes_model.pkl · 646,390 bytes
Auto-grader score274 / 274 — 100% ✓All 5 pipeline sections fully passed

Main Takeaways


What I Learned & Why It Matters to Employers

Sprint 13 is where cloud concepts became operational. I had used S3 before as storage and EC2 as a server — but wiring them together with an IAM role, inside a VPC I configured from scratch, to run a real training job and persist the model artifact, is a different skill. The auto-grader didn't just check if my code ran. It SSHed into the instance, verified the Python environment, checked that the model file existed locally and in S3, validated the VPC routing, and confirmed the IAM role. Every layer of the pipeline had to be correct. Understanding the full stack — from the route table entry to the pickle file in S3 — is what separates someone who can follow a tutorial from someone who can debug a broken ML pipeline in production at 2am.

Conclusion & Reflections

The Pima Indians Diabetes dataset is a means to an end in this project. The real deliverable is the infrastructure: a reproducible, secure, cloud-native workflow where data and model artifacts are durably stored in S3, accessible to any authorized compute in any region, without a single credential in code. That's the pattern that shows up in every serious ML deployment.

If I were extending this project, the next step would be automating the pipeline — replacing the manual EC2 steps with an AWS Step Functions workflow or a SageMaker Pipeline that triggers on new data in S3, trains, evaluates, and conditionally registers the model in the SageMaker Model Registry. The architecture I built here is the foundation that makes that automation possible.

RequirementStatus
S3 bucket with datasets/ and models/ foldersCOMPLETE ✓
Dataset uploaded to S3COMPLETE ✓
IAM role with S3 access (no hardcoded credentials)COMPLETE ✓
VPC with public subnet and internet gatewayCOMPLETE ✓
EC2 with IAM role attached and public IPCOMPLETE ✓
Random Forest trained on EC2, dataset sourced from S3COMPLETE ✓
Trained model uploaded back to S3COMPLETE ✓
Auto-grader full pipeline validation274/274 — 100% ✓

Want to See the Training Script?

The complete train.py pipeline — data loading, model training, evaluation, and S3 upload commands — is on GitHub.