This is the module most people quit on. I need to be honest about that.
The sequence I designed — Shoup, then Nield, then Deisenroth — is deliberate. Shoup holds your hand. Nield connects the math to data. Deisenroth demands rigor. If you go straight to Deisenroth without the warm-up, you will hit Chapter 2 (Linear Algebra), see the formal definition of a vector space with eight axioms, and your brain will say "I am not smart enough for this." That thought will be wrong, but it will feel true, and feelings drive behavior more than facts.
Here is what I actually think about math and ML:
You need less math than Twitter tells you, but more than bootcamps admit. The influencer consensus is either "you need a PhD in math" (gatekeeping) or "you don't need any math, just use sklearn" (dangerous). The truth is in between. You need enough math to understand why your model is failing, not just that it's failing. When your gradient is exploding, you need to know what a gradient is. When your model overfits, you need to understand the bias-variance tradeoff geometrically, not just as a phrase you memorized.
The hardest part is not the math itself. It's the shame. You're an engineer with years of experience. You've built production systems. And now you're sitting with a book about derivatives feeling like a beginner. That gap between your professional identity and your current ability is painful. It's the same feeling that drives imposter syndrome, except here it's inverted — you're an expert who genuinely is a beginner in this one domain. The emotional management of that gap is harder than any equation.
My advice: watch 3Blue1Brown before reading anything. Grant Sanderson's animations do something no textbook can — they make you see the math. When you watch a matrix transformation rotate and scale a vector space in real-time, something clicks that no amount of reading notation achieves. Those 9 hours of YouTube are the highest-ROI investment in this entire curriculum.
About linear algebra specifically: this is the math that matters most for ML, and fortunately, it's the most visual and intuitive branch of mathematics. Vectors, matrices, and transformations have geometric meaning you can literally draw. It's not like number theory where you're manipulating abstract symbols with no physical intuition. Every concept in linear algebra has a picture. Use that.
About calculus: you need less than you think. You need derivatives (how fast is something changing?), partial derivatives (how fast is it changing in this direction?), and the chain rule (how do I propagate changes through a pipeline?). That's backpropagation. That's gradient descent. You do not need to integrate weird functions or solve differential equations for ML. If a resource is spending a lot of time on integration techniques, it's wasting your time for ML purposes.
About probability: this is where it gets genuinely hard, and I won't pretend otherwise. Bayesian reasoning is counterintuitive. Conditional probability breaks people's brains. The reason is that human beings are fundamentally bad at probabilistic thinking — we evolved to think in certainties and stories, not distributions. You will need to sit with the discomfort of "I don't intuitively get this" for longer than you'd like. Harvard Stat 110 (Blitzstein) is the best resource here because he teaches probability through stories and puzzles, not formulas.
Conclusion #
Module 1 is the gatekeeper. Not because the math is impossibly hard, but because it requires a kind of patience that professional software engineers aren't used to exercising. In Rails, you can build something in an afternoon and see it work. In math, you spend a week understanding one concept and have nothing to show for it except a slightly different way of thinking. That's the investment. It pays off in Module 3 when you understand why your model works, not just that it works — and that understanding is what separates ML engineers from ML users.
Predictions #
-
You will want to skip this module. You will feel it's taking too long. You will see people on Twitter building cool ML projects who claim they "never learned the math." Resist. Those people hit a ceiling they can't see yet.
-
3Blue1Brown's linear algebra series will be a revelation. You'll watch it and think "why didn't anyone teach me this in school?"
-
Deisenroth Chapter 5 (Vector Calculus) will be the hardest chapter in this entire curriculum. You'll read it three times. The third time it will click.
-
You'll start seeing matrices everywhere — in database queries, in CSS transforms, in Rails routing tables. Once you have the vocabulary, you can't unsee it.
-
MIT 18.06 (Strang) will either be your favorite course or feel too slow. There's no middle ground with Strang. If it feels slow, switch to the Imperial College Coursera specialization — same content, faster pace, and it's literally made by the author of your textbook.
-
Six months after finishing this module, you'll look back and realize it was the most valuable thing you did. Not the most fun. The most valuable.