CS-4125 Machine Learning
Course overview, syllabus, and week-by-week topic breakdown for CS-4125 Machine Learning.
CS-4125 Machine Learning
A university course on machine learning covering supervised learning, unsupervised learning, and practical ML engineering — primarily following Andrew Ng's materials and Stanford CS229, supplemented with ISR and other resources.
Broad Topic Areas
- General topics on machine learning — motivation, settings, and the ML pipeline
- Supervised learning — regression, classification, neural networks, SVMs
- Unsupervised learning — clustering, dimensionality reduction, generative models
- Practical aspects of machine learning — evaluation, engineering, bias-variance analysis
Algorithms Covered
| Category | Algorithm |
|---|---|
| Regression | Linear Regression, Polynomial Regression |
| Classification | Perceptron, Logistic Regression, Naive Bayes, GDA |
| Neural Networks | Feedforward NN – Representation, Backpropagation & Learning |
| Kernel Methods | Support Vector Machine (SVM) |
| Clustering | K-Means Clustering, Gaussian Mixture Model (GMM) |
| Generative Models | EM Algorithm |
| Dimensionality Reduction | Principal Component Analysis (PCA), Autoencoders |
| Anomaly Detection | Univariate & Multivariate Gaussian |
| Recommender Systems | Collaborative Filtering |
Weekly Syllabus
Week 1 — Overview and Linear Regression
Topics: Human intelligence vs machine learning. Settings of ML (supervised, unsupervised). Evolution from an algorithmic perspective. Linear regression: assumption, evaluation, refinement. Gradient descent.
Reading: PAN Lecture Notes 01 · SUCS229 pp. 1–13 · ISR pp. 1–6, 9–11, 26–29
📄 Notes: Lecture 01 & 02: Introduction, Regression Analysis and Gradient Descent
Week 2 — Linear and Polynomial Regression
Topics: Closed-form solution. Validation set, model selection, hyperparameter tuning. Minibatch and stochastic gradient descent. Feature scaling. Multivariate linear regression.
Polynomial regression. Overfitting and underfitting, learning curves. Bias and variance, generalisation error. Cross-validation. Regularisation (Ridge vs Lasso).
Reading: PAN Lecture Notes 02, 04, 07 · ISR pp. 15–17, 59–63, 71–75, 87–94, 197–205, 237–242 · SUCS229 pp. 113–119, 135–137 · ML Yearning pp. 1–19
📄 Notes: Lecture 03: Linear Algebra Review · Lecture 04: Linear Regression with Multiple Variables · Lecture 07: Regularization
Week 3 — Classification: Problem Formulation
Topics: Problem formulation, hyperplane, decision boundary. Hypothesis, relation with linear regression, target label encoding. Non-linear decision boundary.
📄 Notes: Lecture 06: Logistic Regression (see decision boundary section)
Week 4 — Classification: Perceptron and Logistic Regression
Topics: Perceptron hypothesis, loss function, update rule, prediction, limitations. Logistic regression: hypothesis, loss, update rule. Multiclass: one-vs-all. Multinomial logistic regression. Maximum Likelihood Estimation (MLE).
Reading: PAN Lecture Notes 06 · SUCS229 pp. 20–23 · ISR pp. 129–140
📄 Notes: Lecture 06: Logistic Regression
Week 5 — Generative vs Discriminative Models
Topics: MLE in detail for linear and logistic regression. Naive Bayes: hypothesis, learning, prediction. Generative vs discriminative comparison. MAP as regularisation (L1 and L2).
Reading: SUCS229 Chapter 4 · Jurafsky Ch. 4 pp. 1–7
📄 Notes: Lecture 06: Logistic Regression (covers MLE and classification foundations)
Week 6 — Evaluation Metrics & ML Engineering
Topics: Naive Bayes: text encoding, spam classifier, continuous features. Confusion matrix, accuracy, precision, recall, F1-score. Macro and micro averaging.
ML Engineering for real-life projects.
Reading: Jurafsky Ch. 4 pp. 1–7, 11–14 · Kubat Ch. 11 · ML Yearning Sections 1–13
📄 Notes: Lecture 10: Advice for Applying Machine Learning · Lecture 11: Machine Learning System Design
Week 7 — Neural Networks (Introduction)
Topics: History and motivation. Hypothesis, forward propagation, vectorised computation. Relation with logistic regression.
Reading: PAN Lecture 08
📄 Notes: Lecture 08: Neural Networks – Representation
Week 8 — Neural Networks (Training)
Topics: Backpropagation. Non-convexity of loss, landscape visualisation. Activation functions, hyperparameters, regularisation, weight initialisation. Full forward + backward propagation. Image encoding, multiclass and multilabel classification, softmax, cross-entropy loss. Visualising weights and features.
Reading: PAN Lecture 09 · CS229 Notes
📄 Notes: Lecture 09: Neural Networks – Learning
Week 9 — Incourse Exam Week
Week 10 — SVM and Kernel Methods
Topics: Large and optimum margin. Functional margin vs geometric margin. Primal SVM optimisation. Soft margin SVM, slack variable, hinge loss.
Dual form via inner product. Kernel machine: polynomial kernel derivation, kernel trick.
Reading: cs229-notes3.pdf · PAN Lecture 12
📄 Notes: Lecture 12: Support Vector Machines
Week 11 — ML Engineering (Continued)
Topics: Single-point evaluation metric, optimising/satisfying metrics, train-validation split, iterative development. Empirical bias-variance analysis, reduction techniques, learning curve analysis. Manual error analysis. Comparing with human-level performance, mislabelled examples, data mismatch.
Reading: ML Yearning Sections 1–43, 47–49 · Andrew Ng Deep Learning Specialisation Course 3
📄 Notes: Lecture 10: Advice for Applying Machine Learning · Lecture 11: Machine Learning System Design
Week 12 — Clustering and Anomaly Detection
Topics: Unsupervised learning. K-means: algorithm, cost function, convergence, choosing K, initialisation. Anomaly detection: density-based methods — univariate and multivariate Gaussian, use of labelled cross-validation set.
Reading: PAN Lectures 13, 15 · CS229 Resource · WUSTL Lecture Note 12
📄 Notes: Lecture 13: Clustering · Lecture 15: Anomaly Detection
Week 13 — Gaussian Mixture Models and EM Algorithm
Topics: Gaussian Discriminant Analysis (GDA) — linear and quadratic discriminant analysis. From GDA to GMMs. EM algorithm for GMM. K-means vs GMM (soft) clustering comparison. GMM for anomaly detection.
📄 Notes: Lecture 13: Clustering (covers clustering foundations)
Week 14 — Dimensionality Reduction and Non-Parametric Methods
Topics: Feature selection and engineering, data compression. PCA: motivation, formulation, algorithm, reconstruction, choosing . Projected variance and reconstruction error. Autoencoders: motivation, loss function of encoder-decoder architecture, feature learning.
Non-parametric supervised learning: k-nearest neighbour. Decision Trees (regression and classification): training, prediction, node impurity measures, pruning, benefits and drawbacks.
Reading: ISR pp. 39–42 (k-NN), pp. 329–348 (trees and ensembles)
📄 Notes: Lecture 14: Dimensionality Reduction
Week 15 — Ensemble Methods
Topics: Bagging and Random Forest for variance reduction: from bagging to random forest, why it works. Boosting for bias reduction: motivation and pseudocode of AdaBoost and Gradient Boost.
Also Available
These lectures are available in the notes but not part of the core weekly schedule above:
| Lecture | Topic |
|---|---|
| Lecture 05: Octave | MATLAB/Octave programming for ML |
| Lecture 16: Recommender Systems | Collaborative filtering, low-rank matrix factorisation |
| Lecture 17: Large Scale Machine Learning | Stochastic & minibatch gradient descent, MapReduce |
| Lecture 18: Application Example – Photo OCR | ML pipeline, sliding window, ceiling analysis |
| Lecture 19: Course Summary | End-to-end course recap |
Exam Preparation & Past Year Questions
- 📝 2024 Final Examination (27th Batch) Questions & Solutions — Complete question paper with detailed step-by-step mathematical proofs and solutions covering Regularized Regression, Classification metrics (ROC/AUC), Neural Networks & Backprop, SVM Primal/Dual & Kernels, Decision Trees, K-Means Clustering, Anomaly Detection, GMM vs GDA, and PCA.
Reference Books and Materials
| Reference | Description |
|---|---|
| [ISR] Introduction to Statistical Learning | James, Witten, Hastie & Tibshirani, 2023 |
| [PAN] Andrew Ng's Course Materials | Lecture notes and slides |
| [SUCS229] Stanford CS229, 2025 | Official Stanford ML course notes |
| [IML] Introduction to Machine Learning | Miroslav Kubat |
| [PRML] Pattern Recognition and Machine Learning | Bishop, 2006 |
| [CMU] CMU 10-701 Slides | Classical ML course with strong unsupervised learning content |
In-Course Examinations & Solutions
CSE-4101 Artificial Intelligence In-Course Examination Questions and Detailed Step-by-Step Solutions
Lecture 01 & 02: Introduction, Regression Analysis and Gradient Descent
An introduction to machine learning concepts, supervised vs unsupervised learning, univariate linear regression, cost functions, and gradient descent optimization.