The practitioner's pipeline, from a raw table to a model you can trust in production: missing data and what imputation costs, feature engineering, categorical encoding, scaling, and the leakage that makes a validation score lie. Then evaluation done properly: the four cells of a confusion matrix, precision against recall, ROC and AUC, class imbalance, and the baseline that makes any of those numbers readable. It closes with kNN, k-means, decision trees, random forests, naive Bayes, hyperparameter tuning and model drift. Scoped against mlmath, statistics and dataliteracy, which own the linear algebra behind learning, cross-validation and model selection, and reading other people's charts respectively.
Free to start · adaptive placement finds your level · reviews timed so it stays learned.
Every idea is taught with motivation and a worked example before the drills, and an FSRS spaced-repetition engine schedules each review for the moment just before you'd forget it. A short placement check finds what you already know, so you start Applied Data Science exactly where it's useful.