ITiCSE 2026 · Madrid
University of Toronto Mississauga
Delft University of Technology
Delft University of Technology

Lisa Zhang
University of Toronto Mississauga

Gosia Migut
Delft University of Technology

Jesse Krijthe
Delft University of Technology
28 years of combined ML teaching · North America + Europe
undergrad and graduate · educator and researcher

“…akin to a portal, opening up a new and previously inaccessible way of thinking about something. It represents a transformed way of understanding, or interpreting, or viewing something without which the learner cannot progress.” — (Meyer and Land 2003)
Transformative fundamentally shifts the learner’s perspective
Troublesome difficult to grasp; conflict with prior understanding
Integrative unifies disparate concepts in a domain
Irreversible once understood, unlikely to be unlearned
Bounded limited to specific disciplinary boundaries
Concept granularity: ours will be broad, field-level ideas
…like the AI4K12 “five big ideas” (Touretzky et al. 2019)
Threshold concepts are usually identified through consensus (Barradell 2013; Timmermans and Meyer 2019)
Our approach:
1. Brainstorm Independent brainstorming from our teaching + research experience
2. Discuss Several rounds of discussion on these candidate concepts
3. Ground Ground the arguments in prior work in ML research, ML education & ML/AI misconceptions
Whether TCs can be identified with empirical rigor at all is contested (Rowbottom 2007; Rountree and Rountree 2009)
To describe an ML method, we describe:
Supervised Learning: minimize empirical risk
\[ \min_{f \in \mathcal{F}} \; \frac{1}{n}\sum_{i=1}^{n} \ell\big(f(x_i),\, y_i\big) \]
Reinforcement Learning: maximize expected discounted return
\[ \max_{\pi} \; \mathbb{E}_{\tau \sim \pi}\Big[\textstyle\sum_{t=0}^{T} \gamma^t r_t\Big] \]
Counters novices’ beliefs that:
🧠
ML works similarly to the human brain (Marx et al. 2024)
⚙️
ML behaviour is programmed rather than learned (Marx et al. 2024)
💾
Training data is stored inside the model (Bewersdorff et al. 2023)
This shift lets researchers construct and communicate ML methods by specifying the optimization problem.
⋮ 
⋮ 
Unsupervised learning: k-means clustering
Block coordinate descent on \(J = \sum_{i=1}^{n} \lVert x_i - \mu_{c(i)} \rVert^2\)
Systematic reasoning: “What is this process implicitly optimizing?”
AI alignment: does the system do what its designers intended, not just what the objective literally rewards?

Goodhart’s Law: “When a measure becomes a target, it ceases to be a good measure”
“Empirical”: grounded in observation and experimentation
🔍
“Best algorithm”
→
🎯
“Accuracy”
→
😭
Theory alone isn’t enough Accepting that theory can’t determine the “best” model is unsatisfying
🧪
Good evaluation is a skill Not always taught explicitly; reproducibility is a field-wide concern
One lens for many failure modes:
⚖️
Overfitting & underfitting
🎛️
Lack of hyperparameter exploration
🗂️
Bias in training data
💧
Data leakage
“Natural images lie on a manifold within \(\mathbb{R}^D\)”
“The earlier layers of a Multi-Layer Perceptron learn features, and the final layer is a linear classifier on these features”
Data as geometric objects Points in \(\mathbb{R}^D\) with meaningful distances, symmetries, invariances
Models as geometric processes that distorts this space
Models as geometric objects …with meaningful distances, symmetries, invariances
Optimizers as geometric processes navigating this space
🌀
Connecting Algebraic and Geometric Views Students struggle connecting algebraic and geometric views (Sibia et al. 2025)
📚
Counters prior CS training Earlier courses emphasize data type (chars, pixels, waveforms) and time/space complexity.
🔢
Data geometry Shapes what can be learned: why one-hot encode categorical variables? distributed representation?
🕸️
Model geometry Shapes what models are possible: CNNs & GNNs exploit symmetries; embeddings & transfer learning
🏔️
Loss geometry Shapes how we find good models: gradient descent + momentum, step sizes, clipping, ravines
🔄
Transformative concepts Rather than exhaustive topic coverage
⏳
Liminal time Build in time to work through liminal phases
✅
Real assessment Design assessments that reveal transformation, not mimicry
Not every course needs every concept, with equal weight
Decide the role (user, builder, researcher?), then the transformations we want the learners to undergo
| Week | Lecture Topic |
|---|---|
| 1 | Supervised Learning; Nearest Neighbours |
| 2 | Decision Trees |
| 3 | Linear Regression |
| 4 | Feature Mapping; Classification |
| 5 | Multi-Class Classification; Multi-Layer Perceptrons |
| 6 | Neural Networks; Backpropagation |
| 7 | Bias-Variance Decomposition; Probabilistic Modeling |
| 8 | Algorithmic Fairness |
| 9 | Naive Bayes |
| 10 | Gaussian Discriminant Analysis |
| 11 | Clustering; Mixture Models; Expectation Maximization |
| 12 | Principal Component Analysis |
Revisit the TCs in every unit
Math maturity is an important pre-liminal variation
Learners can cross the threshold with minimal math
Math self-efficacy remains a barrier; careful instruction can help, but transformations may not stick
Thank you!
m.a.migut@tudelft.nl
j.h.krijthe@tudelft.nl