CSE-41XX
CS-4125 ML

CS-4125 Machine Learning

Course overview, syllabus, and week-by-week topic breakdown for CS-4125 Machine Learning.

CS-4125 Machine Learning

A university course on machine learning covering supervised learning, unsupervised learning, and practical ML engineering — primarily following Andrew Ng's materials and Stanford CS229, supplemented with ISR and other resources.


Broad Topic Areas

  1. General topics on machine learning — motivation, settings, and the ML pipeline
  2. Supervised learning — regression, classification, neural networks, SVMs
  3. Unsupervised learning — clustering, dimensionality reduction, generative models
  4. Practical aspects of machine learning — evaluation, engineering, bias-variance analysis

Algorithms Covered

CategoryAlgorithm
RegressionLinear Regression, Polynomial Regression
ClassificationPerceptron, Logistic Regression, Naive Bayes, GDA
Neural NetworksFeedforward NN – Representation, Backpropagation & Learning
Kernel MethodsSupport Vector Machine (SVM)
ClusteringK-Means Clustering, Gaussian Mixture Model (GMM)
Generative ModelsEM Algorithm
Dimensionality ReductionPrincipal Component Analysis (PCA), Autoencoders
Anomaly DetectionUnivariate & Multivariate Gaussian
Recommender SystemsCollaborative Filtering

Weekly Syllabus

Week 1 — Overview and Linear Regression

Topics: Human intelligence vs machine learning. Settings of ML (supervised, unsupervised). Evolution from an algorithmic perspective. Linear regression: assumption, evaluation, refinement. Gradient descent.

Reading: PAN Lecture Notes 01 · SUCS229 pp. 1–13 · ISR pp. 1–6, 9–11, 26–29

📄 Notes: Lecture 01 & 02: Introduction, Regression Analysis and Gradient Descent


Week 2 — Linear and Polynomial Regression

Topics: Closed-form solution. Validation set, model selection, hyperparameter tuning. Minibatch and stochastic gradient descent. Feature scaling. Multivariate linear regression.

Polynomial regression. Overfitting and underfitting, learning curves. Bias and variance, generalisation error. Cross-validation. Regularisation (Ridge vs Lasso).

Reading: PAN Lecture Notes 02, 04, 07 · ISR pp. 15–17, 59–63, 71–75, 87–94, 197–205, 237–242 · SUCS229 pp. 113–119, 135–137 · ML Yearning pp. 1–19

📄 Notes: Lecture 03: Linear Algebra Review · Lecture 04: Linear Regression with Multiple Variables · Lecture 07: Regularization


Week 3 — Classification: Problem Formulation

Topics: Problem formulation, hyperplane, decision boundary. Hypothesis, relation with linear regression, target label encoding. Non-linear decision boundary.

📄 Notes: Lecture 06: Logistic Regression (see decision boundary section)


Week 4 — Classification: Perceptron and Logistic Regression

Topics: Perceptron hypothesis, loss function, update rule, prediction, limitations. Logistic regression: hypothesis, loss, update rule. Multiclass: one-vs-all. Multinomial logistic regression. Maximum Likelihood Estimation (MLE).

Reading: PAN Lecture Notes 06 · SUCS229 pp. 20–23 · ISR pp. 129–140

📄 Notes: Lecture 06: Logistic Regression


Week 5 — Generative vs Discriminative Models

Topics: MLE in detail for linear and logistic regression. Naive Bayes: hypothesis, learning, prediction. Generative vs discriminative comparison. MAP as regularisation (L1 and L2).

Reading: SUCS229 Chapter 4 · Jurafsky Ch. 4 pp. 1–7

📄 Notes: Lecture 06: Logistic Regression (covers MLE and classification foundations)


Week 6 — Evaluation Metrics & ML Engineering

Topics: Naive Bayes: text encoding, spam classifier, continuous features. Confusion matrix, accuracy, precision, recall, F1-score. Macro and micro averaging.

ML Engineering for real-life projects.

Reading: Jurafsky Ch. 4 pp. 1–7, 11–14 · Kubat Ch. 11 · ML Yearning Sections 1–13

📄 Notes: Lecture 10: Advice for Applying Machine Learning · Lecture 11: Machine Learning System Design


Week 7 — Neural Networks (Introduction)

Topics: History and motivation. Hypothesis, forward propagation, vectorised computation. Relation with logistic regression.

Reading: PAN Lecture 08

📄 Notes: Lecture 08: Neural Networks – Representation


Week 8 — Neural Networks (Training)

Topics: Backpropagation. Non-convexity of loss, landscape visualisation. Activation functions, hyperparameters, regularisation, weight initialisation. Full forward + backward propagation. Image encoding, multiclass and multilabel classification, softmax, cross-entropy loss. Visualising weights and features.

Reading: PAN Lecture 09 · CS229 Notes

📄 Notes: Lecture 09: Neural Networks – Learning


Week 9 — Incourse Exam Week


Week 10 — SVM and Kernel Methods

Topics: Large and optimum margin. Functional margin vs geometric margin. Primal SVM optimisation. Soft margin SVM, slack variable, hinge loss.

Dual form via inner product. Kernel machine: polynomial kernel derivation, kernel trick.

Reading: cs229-notes3.pdf · PAN Lecture 12

📄 Notes: Lecture 12: Support Vector Machines


Week 11 — ML Engineering (Continued)

Topics: Single-point evaluation metric, optimising/satisfying metrics, train-validation split, iterative development. Empirical bias-variance analysis, reduction techniques, learning curve analysis. Manual error analysis. Comparing with human-level performance, mislabelled examples, data mismatch.

Reading: ML Yearning Sections 1–43, 47–49 · Andrew Ng Deep Learning Specialisation Course 3

📄 Notes: Lecture 10: Advice for Applying Machine Learning · Lecture 11: Machine Learning System Design


Week 12 — Clustering and Anomaly Detection

Topics: Unsupervised learning. K-means: algorithm, cost function, convergence, choosing K, initialisation. Anomaly detection: density-based methods — univariate and multivariate Gaussian, use of labelled cross-validation set.

Reading: PAN Lectures 13, 15 · CS229 Resource · WUSTL Lecture Note 12

📄 Notes: Lecture 13: Clustering · Lecture 15: Anomaly Detection


Week 13 — Gaussian Mixture Models and EM Algorithm

Topics: Gaussian Discriminant Analysis (GDA) — linear and quadratic discriminant analysis. From GDA to GMMs. EM algorithm for GMM. K-means vs GMM (soft) clustering comparison. GMM for anomaly detection.

📄 Notes: Lecture 13: Clustering (covers clustering foundations)


Week 14 — Dimensionality Reduction and Non-Parametric Methods

Topics: Feature selection and engineering, data compression. PCA: motivation, formulation, algorithm, reconstruction, choosing kk. Projected variance and reconstruction error. Autoencoders: motivation, loss function of encoder-decoder architecture, feature learning.

Non-parametric supervised learning: k-nearest neighbour. Decision Trees (regression and classification): training, prediction, node impurity measures, pruning, benefits and drawbacks.

Reading: ISR pp. 39–42 (k-NN), pp. 329–348 (trees and ensembles)

📄 Notes: Lecture 14: Dimensionality Reduction


Week 15 — Ensemble Methods

Topics: Bagging and Random Forest for variance reduction: from bagging to random forest, why it works. Boosting for bias reduction: motivation and pseudocode of AdaBoost and Gradient Boost.


Also Available

These lectures are available in the notes but not part of the core weekly schedule above:

LectureTopic
Lecture 05: OctaveMATLAB/Octave programming for ML
Lecture 16: Recommender SystemsCollaborative filtering, low-rank matrix factorisation
Lecture 17: Large Scale Machine LearningStochastic & minibatch gradient descent, MapReduce
Lecture 18: Application Example – Photo OCRML pipeline, sliding window, ceiling analysis
Lecture 19: Course SummaryEnd-to-end course recap

Exam Preparation & Past Year Questions

  • 📝 2024 Final Examination (27th Batch) Questions & Solutions — Complete question paper with detailed step-by-step mathematical proofs and solutions covering Regularized Regression, Classification metrics (ROC/AUC), Neural Networks & Backprop, SVM Primal/Dual & Kernels, Decision Trees, K-Means Clustering, Anomaly Detection, GMM vs GDA, and PCA.

Reference Books and Materials

ReferenceDescription
[ISR] Introduction to Statistical LearningJames, Witten, Hastie & Tibshirani, 2023
[PAN] Andrew Ng's Course MaterialsLecture notes and slides
[SUCS229] Stanford CS229, 2025Official Stanford ML course notes
[IML] Introduction to Machine LearningMiroslav Kubat
[PRML] Pattern Recognition and Machine LearningBishop, 2006
[CMU] CMU 10-701 SlidesClassical ML course with strong unsupervised learning content

On this page