At a Glance
- Provisioned a full AWS cloud environment — VPC, public subnet, EC2 instance, and IAM role — to run a complete ML training pipeline entirely in the cloud
- Trained a Random Forest classifier on the Pima Indians Diabetes Dataset using scikit-learn on EC2; stored dataset and model artifact in S3 — scored 274/274 (100%) on the auto-grader
- AWS EC2, S3, IAM, VPC, EC2 Instance Connect, AWS CLI, Python, scikit-learn, pickle
Problem
In most ML tutorials, everything runs on your laptop. The dataset is local. The model is local. If your machine crashes, you lose everything. That setup breaks down the moment you need to train on data too large for one disk, collaborate with a team, or deploy predictions to an application. The real ML workflow — the one used in production — looks different: data lives in durable cloud storage, training happens on elastic compute, and the trained model is persisted independently of any single server.
The challenge here was to build that workflow from scratch using AWS. Starting from nothing — no data, no server, no permissions — the goal was to arrive at a trained diabetes prediction model stored in S3, accessible from anywhere, independent of whether the EC2 instance is even running.
Solution
The pipeline wires together four AWS services into a complete ML workflow: S3 for durable object storage (dataset in, model out), IAM for credential-free access between services, VPC for a private, internet-connected network, and EC2 for on-demand compute. The training itself — a scikit-learn Random Forest classifier on the Pima Indians Diabetes dataset — runs entirely on the cloud server via the AWS CLI. No local training. No manual file transfers. Every piece is wired by AWS services.
├── datasets/ diabetes.csv ← source data uploaded first
└── models/ diabetes_model.pkl ← trained model saved here
IAM Role EC2-S3-Access-Role (AmazonS3FullAccess)
↓ attached to
VPC my-project-vpc (10.0.0.0/16)
└── Public Subnet my-project-subnet-public1-us-east-1a
└── EC2 ml-training-instance (t2.micro · Amazon Linux 2023)
├── Downloads diabetes.csv from S3
├── Trains RandomForestClassifier
└── Uploads diabetes_model.pkl back to S3
Internet Gateway + Route Table connect EC2 to the internet
Skills Acquired
- Amazon S3 — created a bucket, organized it with
datasets/andmodels/prefixes, uploaded data, and stored a trained model — understanding S3 as durable, cloud-native object storage independent of compute. - AWS IAM — created an EC2-specific IAM role (
EC2-S3-Access-Role) withAmazonS3FullAccess, attached it at instance launch time — no access keys, no hardcoded credentials. - Amazon VPC — provisioned a private network with 1 AZ, 1 public subnet, an Internet Gateway, and route tables using the "VPC and more" wizard, understanding what each resource does and why it's needed.
- Amazon EC2 — launched a
t2.microAmazon Linux 2023 instance inside the VPC, attached the IAM role, enabled a public IP, and connected via EC2 Instance Connect — no SSH key required. - AWS CLI — used
aws s3 ls,aws s3 cpfrom the EC2 terminal to move data between S3 and the instance — the bridge between compute and storage. - scikit-learn — trained a
RandomForestClassifier(100 estimators, max depth 10) on the Pima Indians Diabetes dataset, evaluated with accuracy score and classification report, and serialized the model withpickle. - Python on EC2 — installed Python packages (
pandas,scikit-learn) on a fresh Amazon Linux instance and ran a full ML training pipeline from the command line. - Cloud ML architecture — internalized the separation of concerns: S3 is permanent storage, EC2 is temporary compute, IAM is the glue that connects them securely.
Deep Dive
The Pima Indians Diabetes dataset is a classic binary classification benchmark: 768 patient records, 8 clinical features, and a binary target indicating diabetes diagnosis. It's small enough to run anywhere — but training it on EC2 with data sourced from S3 is the point. The dataset is just the payload. The real lesson is the infrastructure: how do you move data securely between services? How does a server prove it has permission to read from a storage bucket without a password?
This is the same pattern used in every production AWS ML pipeline. Whether you're training on a t2.micro or a p4d.24xlarge GPU instance, the architecture is identical: an IAM role grants the instance access to S3, the instance pulls data, trains, and pushes the model artifact back. Swapping the instance type is a one-line change. The security and data flow don't change at all.
Why This Project?
This was Sprint 13 of the AWS Cloud Institute program — the capstone lab that ties together S3, IAM, VPC, and EC2 into a single coherent workflow. After individual labs on each service in isolation, this sprint required wiring them all together: storage, permissions, networking, and compute operating as one system.
The skills built here — S3 data management, IAM role configuration, VPC setup, EC2 provisioning — are foundational to any cloud-deployed ML system. Understanding how these services interact is what separates someone who can run a Jupyter notebook from someone who can deploy and maintain an ML system that actually stays running.
What You'll Learn from This
- How IAM roles grant EC2 instances access to S3 without embedding credentials — and why this is the correct approach
- What a VPC actually is and why EC2 instances need one to communicate with the internet
- The difference between S3 (permanent storage) and EC2 (temporary compute) — and why this separation matters for ML pipelines
- How to use the AWS CLI to move data between services from a running EC2 terminal
- How to install Python packages, run a training script, and serialize a model on a cloud server
- Why saving the model to S3 after training is critical — and what happens if you don't
Key Takeaways
- S3 + EC2 + IAM is the canonical AWS ML pattern. Every production ML workflow on AWS uses some version of this — even SageMaker uses it under the hood.
- Model artifacts live in S3, not on EC2. Terminating the EC2 instance doesn't lose the model. S3 is 99.999999999% durable; EC2 instance storage is ephemeral.
- IAM roles eliminate credential management. No access keys, no secret rotation, no leaked credentials in git history. The role is the credential.
- VPC gives you network isolation. The EC2 instance ran in a private network with a controlled path to the internet — the internet gateway and route table define that path explicitly.
- EC2 Instance Connect removes the key pair requirement. Browser-based SSH via the AWS console, gated by IAM — no PEM file to manage or lose.
- Auto-grader validated every layer: 274/274 (100%). VPC, IAM, EC2 config, SSH access, Python environment, and both S3 artifacts all passed.
The Dataset
The Pima Indians Diabetes dataset is a widely used binary classification benchmark from the UCI Machine Learning Repository. It contains 768 patient records, all female patients of Pima Indian heritage, with 8 clinical features and a binary outcome variable.
| Feature | Type | Description |
|---|---|---|
Pregnancies | Numeric | Number of times pregnant |
Glucose | Numeric | Plasma glucose concentration (2-hour oral glucose tolerance test) |
BloodPressure | Numeric | Diastolic blood pressure (mm Hg) |
SkinThickness | Numeric | Triceps skin fold thickness (mm) |
Insulin | Numeric | 2-hour serum insulin (mu U/ml) |
BMI | Numeric | Body mass index (kg/m²) |
DiabetesPedigreeFunction | Numeric | Diabetes pedigree function (genetic risk score) |
Age | Numeric | Age in years |
Outcome | Binary | TARGET — 0: no diabetes, 1: diabetes (268 positive / 500 negative) |
My Process
Phase 1
Create the S3 Bucket & Folder Structure
Created an S3 bucket named vietnguyen-ml-project-3421 and organized it with two folders: datasets/ for raw training data and models/ for trained model artifacts. Good folder structure in S3 matters — in a real ML system, you'd also add versioned subfolders for experiment tracking.
Phase 2
Upload the Dataset to S3
Uploaded diabetes.csv to s3://vietnguyen-ml-project-3421/datasets/. The data is now in durable cloud storage — accessible by any authorized AWS service in any region, independent of any compute instance.
Phase 3
Create the IAM Role for EC2
Created EC2-S3-Access-Role in IAM: trusted entity type AWS service, use case EC2, permission AmazonS3FullAccess. This role will be attached to the EC2 instance, granting it transparent S3 access through AWS's internal credential vending mechanism — no keys in code.
Phase 4
Create the VPC
Used the "VPC and more" wizard to create my-project-vpc (CIDR: 10.0.0.0/16) with 1 Availability Zone, 1 public subnet, an Internet Gateway (my-project-igw), and a route table (my-project-rtb-public) with a 0.0.0.0/0 → IGW route. No NAT gateway, no private subnets — the EC2 instance is public-facing for SSH access.
Phase 5
Launch EC2 with IAM Role Attached
Launched a t2.micro instance running Amazon Linux 2023, placed inside my-project-vpc with auto-assign public IP enabled. Under Advanced Details, attached EC2-S3-Access-Role as the IAM instance profile. This single step is what gives the instance permission to access S3 — without it, every aws s3 command returns Access Denied.
Phase 6
Connect via EC2 Instance Connect & Verify S3 Access
Connected to the running instance via EC2 Instance Connect in the browser — no SSH key required. Verified S3 access with:
# List bucket contents — confirms IAM role is working aws s3 ls s3://vietnguyen-ml-project-3421/ # Expected output: # PRE datasets/ # PRE models/
Installed Python dependencies on the fresh instance:
sudo dnf install -y python3 python3-pip
pip3 install --upgrade pip
pip3 install pandas scikit-learn
Phase 7
Download Dataset from S3 to EC2
Created a working directory structure and pulled the dataset from S3 to the instance's local storage:
mkdir -p ~/ml-project/data ~/ml-project/models aws s3 cp s3://vietnguyen-ml-project-3421/datasets/diabetes.csv \ ~/ml-project/data/ # Verify — should show diabetes.csv ls ~/ml-project/data/ head ~/ml-project/data/diabetes.csv
Phase 8
Train the Random Forest Model
Created train.py on the EC2 instance — a complete ML pipeline: load data, prepare features and target, split 80/20, train a Random Forest, evaluate, print feature importances, and save the model with pickle:
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import accuracy_score, classification_report import pickle # Load dataset df = pd.read_csv('/home/ec2-user/ml-project/data/diabetes.csv') print(f"Dataset shape: {df.shape[0]} rows, {df.shape[1]} columns") # Features and target X = df.iloc[:, :-1] # all columns except Outcome y = df.iloc[:, -1] # Outcome column # 80/20 split X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # Train Random Forest model = RandomForestClassifier( n_estimators=100, max_depth=10, random_state=42 ) model.fit(X_train, y_train) # Evaluate y_pred = model.predict(X_test) accuracy = accuracy_score(y_test, y_pred) print(f"Test Accuracy: {accuracy:.2%}") print(classification_report(y_test, y_pred, target_names=['No Diabetes', 'Diabetes'])) # Feature importance rankings importances = model.feature_importances_ feature_names = df.columns[:-1].tolist() for name, imp in sorted(zip(feature_names, importances), key=lambda x: x[1], reverse=True): print(f" {name}: {imp:.4f}") # Save model model_path = '/home/ec2-user/ml-project/models/diabetes_model.pkl' with open(model_path, 'wb') as f: pickle.dump(model, f) print(f"Model saved to: {model_path}")
Ran the training script directly on the cloud server:
python3 ~/ml-project/train.py
⚡ Try It — Interactive Chart
⚡ Try It — Interactive Chart
Phase 9
Upload Trained Model Back to S3
With the model serialized locally, one aws s3 cp command moved it to permanent storage. The model now lives in S3 independently of the EC2 instance — the instance can be terminated, the model stays:
aws s3 cp ~/ml-project/models/diabetes_model.pkl \ s3://vietnguyen-ml-project-3421/models/ # Verify from EC2 — confirms model landed in S3 aws s3 ls s3://vietnguyen-ml-project-3421/models/ # Output: 2026-05-21 ... 646390 diabetes_model.pkl
1) VPC RESOURCES PASS VPC exists (my-project-vpc) — vpc-0985530cde96d5c6c PASS CIDR is 10.0.0.0/16 PASS Tenancy is default PASS Exactly 1 AZ and 1 subnet PASS Subnet: my-project-subnet-public1-us-east-1a PASS IGW my-project-igw exists and is attached PASS RTB my-project-rtb-public has 0.0.0.0/0 → IGW route PASS No NAT gateways / No VPC endpoints 2) IAM ROLE PASS Role exists: EC2-S3-Access-Role PASS Trust policy includes ec2.amazonaws.com PASS AmazonS3FullAccess attached 3) EC2 INSTANCE PASS Instance: ml-training-instance (i-032d9334018728a23) PASS State: running · Type: t2.micro · AMI: Amazon Linux 2023 PASS No key pair · In my-project-vpc · Public IP assigned PASS IAM profile attached: EC2-S3-Access-Role 4) SSH CHECKS (EC2 Instance Connect) PASS SSH connected and checks executed PASS python3 · pip3 · pandas · scikit-learn all available PASS ~/ml-project/data/diabetes.csv exists PASS ~/ml-project/train.py exists PASS ~/ml-project/models/diabetes_model.pkl exists 5) S3 BUCKET + UPLOADS PASS Bucket: vietnguyen-ml-project-3421 PASS datasets/ and models/ prefixes exist PASS datasets/diabetes.csv uploaded PASS models/diabetes_model.pkl — Size: 646,390 bytes ══════════════════════════════ Raw points: 274 / 274 · Percentage: 100%
⚡ Try It — Interactive Chart
Final Results
| Check | Status | Details |
|---|---|---|
| VPC created and configured | PASS ✓ | my-project-vpc · 10.0.0.0/16 · 1 AZ · 1 public subnet |
| IAM role with S3 permissions | PASS ✓ | EC2-S3-Access-Role · AmazonS3FullAccess |
| EC2 instance running | PASS ✓ | ml-training-instance · t2.micro · Amazon Linux 2023 |
| SSH via EC2 Instance Connect | PASS ✓ | Connected and checks executed |
| Python environment | PASS ✓ | python3 · pip3 · pandas · scikit-learn all available |
| Dataset on EC2 | PASS ✓ | ~/ml-project/data/diabetes.csv |
| Training script on EC2 | PASS ✓ | ~/ml-project/train.py |
| Trained model on EC2 | PASS ✓ | ~/ml-project/models/diabetes_model.pkl |
| Dataset in S3 | PASS ✓ | s3://vietnguyen-ml-project-3421/datasets/diabetes.csv |
| Trained model in S3 | PASS ✓ | s3://…/models/diabetes_model.pkl · 646,390 bytes |
| Auto-grader score | 274 / 274 — 100% ✓ | All 5 pipeline sections fully passed |
Main Takeaways
- S3 is for storage; EC2 is for compute — never conflate them. Data and models live in S3. Computation happens on EC2. When compute is done, you terminate it. The data persists regardless.
- IAM roles are the right way to grant access between AWS services. Hardcoding access keys is a security anti-pattern that surfaces credentials in code and requires manual rotation. Roles provide temporary, automatically rotated credentials through the metadata service — no keys, no risk.
- The VPC is the network fabric. Without the Internet Gateway and public subnet route, the EC2 instance has no path to download packages or connect via EC2 Instance Connect — the VPC configuration isn't optional boilerplate; it's load-bearing infrastructure.
- EC2 Instance Connect eliminates key pair management. For an IAM-authenticated team, browser-based SSH gated by IAM permissions is operationally cleaner than distributing PEM files and managing key rotation.
- This pattern scales linearly. Swap the t2.micro for a p4d.24xlarge and the architecture is identical. The IAM role, the S3 paths, the AWS CLI commands — nothing changes. The compute scales; the data pipeline doesn't need to be redesigned.
What I Learned & Why It Matters to Employers
Sprint 13 is where cloud concepts became operational. I had used S3 before as storage and EC2 as a server — but wiring them together with an IAM role, inside a VPC I configured from scratch, to run a real training job and persist the model artifact, is a different skill. The auto-grader didn't just check if my code ran. It SSHed into the instance, verified the Python environment, checked that the model file existed locally and in S3, validated the VPC routing, and confirmed the IAM role. Every layer of the pipeline had to be correct. Understanding the full stack — from the route table entry to the pickle file in S3 — is what separates someone who can follow a tutorial from someone who can debug a broken ML pipeline in production at 2am.
Conclusion & Reflections
The Pima Indians Diabetes dataset is a means to an end in this project. The real deliverable is the infrastructure: a reproducible, secure, cloud-native workflow where data and model artifacts are durably stored in S3, accessible to any authorized compute in any region, without a single credential in code. That's the pattern that shows up in every serious ML deployment.
If I were extending this project, the next step would be automating the pipeline — replacing the manual EC2 steps with an AWS Step Functions workflow or a SageMaker Pipeline that triggers on new data in S3, trains, evaluates, and conditionally registers the model in the SageMaker Model Registry. The architecture I built here is the foundation that makes that automation possible.
| Requirement | Status |
|---|---|
| S3 bucket with datasets/ and models/ folders | COMPLETE ✓ |
| Dataset uploaded to S3 | COMPLETE ✓ |
| IAM role with S3 access (no hardcoded credentials) | COMPLETE ✓ |
| VPC with public subnet and internet gateway | COMPLETE ✓ |
| EC2 with IAM role attached and public IP | COMPLETE ✓ |
| Random Forest trained on EC2, dataset sourced from S3 | COMPLETE ✓ |
| Trained model uploaded back to S3 | COMPLETE ✓ |
| Auto-grader full pipeline validation | 274/274 — 100% ✓ |
Want to See the Training Script?
The complete train.py pipeline — data loading, model training, evaluation, and S3 upload commands — is on GitHub.