Data Science and Machine Learning in Chemical Engineering#
06-642 · Spring 2026 · Carnegie Mellon University
Welcome to the course! This half-semester course covers practical applications of data science and machine learning techniques to problems in chemical engineering and related sciences.
Course Goals#
By the end of this course, you will be able to:
Manipulate data using NumPy and Pandas for analysis and visualization
Build predictive models using scikit-learn for regression and classification
Apply advanced techniques including ensemble methods, clustering, and dimensionality reduction
Quantify uncertainty in model predictions using pycse and Gaussian processes
Interpret models to understand what drives predictions
Apply these skills to real chemical engineering problems
Course Structure#
The course consists of 12 lectures covering:
Module |
Topic |
|---|---|
00 |
Introduction & Python Environment |
01 |
NumPy Fundamentals |
02 |
Pandas Introduction |
03 |
Intermediate Pandas |
04 |
Dimensionality Reduction |
05 |
Linear Regression |
06 |
Regularization & Model Selection |
07 |
Nonlinear Methods |
08 |
Ensemble Methods |
09 |
Clustering |
10 |
Uncertainty Quantification |
11 |
Model Interpretability |
Plus optional lectures on databases, deep learning, symbolic regression, and LLMs.
Prerequisites#
Basic Python programming
Undergraduate mathematics (linear algebra, calculus, statistics)
No prior machine learning experience required
Interactive Learning#
This course includes an interactive learning system to help reinforce concepts:
Data Academy Honors#
Earn badges and ranks as you progress:
5 Badges: Data Wrangler, Pattern Seeker, Model Builder, Ensemble Master, Uncertainty Expert
4 Ranks: Data Apprentice → Data Analyst → Data Scientist → Senior Data Scientist
See Data Academy Honors for details.
Games & Activities#
Trivia (
/trivia): Test your knowledge with character-hosted questionsAdventure (
/adventure): “The Data Detective Agency” - solve ChemE data mysteriesFlashcards (
/flashcards): Spaced repetition for key conceptsPuzzles (
/puzzle): Metric matching, confusion matrices, code completionScavenger Hunts (
/hunt): Explore datasets, documentation, and papers
Course Characters#
Meet your guides including Nadia Null, Reggie Regression, Val Validation, and more! See characters.md.
Quick Reference#
The reference sheet provides concise summaries of pandas, sklearn, metrics, and common patterns.
Repository Setup#
Initial Setup#
# Clone the course repository
git clone <course-repo-url>
cd dsmles
# Install notebook merge tool (handles Jupyter notebook conflicts)
pip install nbdime
nbdime config-git --enable --global
# Create your working branch (keep main clean for updates)
git checkout -b my-work
Getting Course Updates#
When the instructor announces updates:
# Switch to main and pull updates
git checkout main
git pull origin main
# Return to your working branch
git checkout my-work
# Optionally merge updates into your branch
git merge main
If you encounter notebook merge conflicts, nbdime provides a visual merge tool:
git mergetool --tool=nbdime
Why This Workflow?#
main branch stays clean and always pulls without conflicts
your branch contains all your work, notes, and completed exercises
you control when to bring in updates
nbdime handles notebook merges intelligently when needed
Getting Started#
Review the Syllabus for course policies
Set up your Python environment (see Module 00: Introduction)
Clone the repo and create your working branch (see above)
Complete assignments as they are released
Try
/triviato test your knowledge as you learn!
Let’s begin!