Machine Learning System Design Interview
Paolo Perrone · System Design Newsletter issue #146 · plain-English field guide to ML system design — free portion below covers concepts 1–11; full 38-concept version is subscriber-only · read
What it is
Issue #146 of systemdesign.one (“Shipping Production AI”), guest-written by Paolo Perrone — ML engineer with 8+ years of production AI experience, read by 1M+ engineers across LinkedIn/Substack (runs The AI Engineer newsletter).
Premise: most ML-interview prep offers algorithm drills or LeetCode, but what you actually need first is vocabulary + building blocks + how they fit together — no TensorFlow or gradient-descent math required. It traces how real systems work throughout: Netflix recommendations, Uber ETA prediction, Spotify Discover Weekly.
Core frame: an ML system is just a pipeline — “data flows in, patterns get extracted, predictions flow out, and the system keeps learning from its own results.” Every concept maps to a stage of that pipeline.
Free concepts (1–11)
Raw material: how data becomes intelligence
- Features — single measurable inputs the model predicts from (“columns in a spreadsheet”); quality matters more than model sophistication. Spotify doesn’t “hear” music — it reads features like top genres, time of day, skip rate.
- Feature engineering — reshaping raw data into signal: Uber’s raw timestamp
2024-12-25 08:47:32 UTCbecomes day-of-week/morning-rush/holiday/rainy. Best features are often aggregates that don’t exist in raw data (“avg rides per week”). - Labels & ground truth — label = what you train on, ground truth = what actually happened; the latter is surprisingly ambiguous (“did the user enjoy this movie?” — watched to end? rated? similar next?). 💡 Your label definition shapes everything: pick the wrong definition of success and the model optimizes for the wrong thing.
- Training/validation/test sets — typical splits 70/15/15 or 80/10/10; non-negotiable rule: test set must never influence development. Netflix-style time-based splits (train Jan–Oct, validate Nov, test Dec) matter because real-world data drifts.
- Class imbalance & data leakage — fraud where 1-in-10,000 cases means “predict not fraud always” scores 99.99% and is useless; leakage = answer hiding inside training data (a “customer complaint filed” column that exists only because delivery was late). 💡 Rule of thumb: >95% accuracy on first try → check leakage; >99% → check imbalance too.
- Data pipelines & ETL — extract/transform/load plumbing; breaks are silent (dropped row, null, late run — no errors, just worse predictions). Uber runs thousands daily with monitoring and quality checks. Mentioning this in interviews signals you think about practice, not papers.
- Data versioning — Git-for-data so you can answer “did performance drop because of the model or the data?” Also required for reproducibility in finance/healthcare audits. Tools: DVC, MLflow.
Pattern machines: how models learn
- Model training & loss functions — loss measures how wrong predictions are; the model nudges internal settings to shrink it millions of times. MSE penalizes big misses (regression), cross-entropy punishes confident wrongness (classification); choice of loss shapes behavior downstream.
- Parameters vs hyperparameters — parameters are learned (model memory; GPT-4-scale hundreds of billions), hyperparameters you set before training (learning rate/batch size/epochs/architecture — the study plan vs knowledge metaphor).
- Embeddings — represent anything (word/user/product) as vectors where meaning = proximity on a map (“comedy” and “rom-com” neighbors, “horror” distant same galaxy, “accounting” another universe). Spotify song embeddings, Netflix movie×user embeddings power recommendations.
- Overfitting & regularization — (section begins, then paywall) the danger of a model learning its training data too well.
What’s behind the paywall
Concepts 11–38 continue into serving/production territory per the pitch (feature stores, real-world architectures, scale/reliability/performance) as subscriber-only “golden member” content. The visible page also carries a partner ad read (AgentField — open-source harness orchestration to call coding agents programmatically from Python/TS/Go) — sponsored insert, not part of the guide.
Interview takeaways worth stealing
- Structure answers around the pipeline stages, not model choice.
- Know the two “suspiciously good accuracy” traps (leakage, imbalance) cold.
- Mentioning monitoring/versioning/pipeline ops beats reciting architectures.