Machine Learning System Design Interview

Paolo Perrone · System Design Newsletter issue #146 · plain-English field guide to ML system design — free portion below covers concepts 1–11; full 38-concept version is subscriber-only · read

What it is

Issue #146 of systemdesign.one (“Shipping Production AI”), guest-written by Paolo Perrone — ML engineer with 8+ years of production AI experience, read by 1M+ engineers across LinkedIn/Substack (runs The AI Engineer newsletter).

Premise: most ML-interview prep offers algorithm drills or LeetCode, but what you actually need first is vocabulary + building blocks + how they fit together — no TensorFlow or gradient-descent math required. It traces how real systems work throughout: Netflix recommendations, Uber ETA prediction, Spotify Discover Weekly.

Core frame: an ML system is just a pipeline — “data flows in, patterns get extracted, predictions flow out, and the system keeps learning from its own results.” Every concept maps to a stage of that pipeline.

Free concepts (1–11)

Raw material: how data becomes intelligence

  1. Features — single measurable inputs the model predicts from (“columns in a spreadsheet”); quality matters more than model sophistication. Spotify doesn’t “hear” music — it reads features like top genres, time of day, skip rate.
  2. Feature engineering — reshaping raw data into signal: Uber’s raw timestamp 2024-12-25 08:47:32 UTC becomes day-of-week/morning-rush/holiday/rainy. Best features are often aggregates that don’t exist in raw data (“avg rides per week”).
  3. Labels & ground truth — label = what you train on, ground truth = what actually happened; the latter is surprisingly ambiguous (“did the user enjoy this movie?” — watched to end? rated? similar next?). 💡 Your label definition shapes everything: pick the wrong definition of success and the model optimizes for the wrong thing.
  4. Training/validation/test sets — typical splits 70/15/15 or 80/10/10; non-negotiable rule: test set must never influence development. Netflix-style time-based splits (train Jan–Oct, validate Nov, test Dec) matter because real-world data drifts.
  5. Class imbalance & data leakage — fraud where 1-in-10,000 cases means “predict not fraud always” scores 99.99% and is useless; leakage = answer hiding inside training data (a “customer complaint filed” column that exists only because delivery was late). 💡 Rule of thumb: >95% accuracy on first try → check leakage; >99% → check imbalance too.
  6. Data pipelines & ETL — extract/transform/load plumbing; breaks are silent (dropped row, null, late run — no errors, just worse predictions). Uber runs thousands daily with monitoring and quality checks. Mentioning this in interviews signals you think about practice, not papers.
  7. Data versioning — Git-for-data so you can answer “did performance drop because of the model or the data?” Also required for reproducibility in finance/healthcare audits. Tools: DVC, MLflow.

Pattern machines: how models learn

  1. Model training & loss functions — loss measures how wrong predictions are; the model nudges internal settings to shrink it millions of times. MSE penalizes big misses (regression), cross-entropy punishes confident wrongness (classification); choice of loss shapes behavior downstream.
  2. Parameters vs hyperparameters — parameters are learned (model memory; GPT-4-scale hundreds of billions), hyperparameters you set before training (learning rate/batch size/epochs/architecture — the study plan vs knowledge metaphor).
  3. Embeddings — represent anything (word/user/product) as vectors where meaning = proximity on a map (“comedy” and “rom-com” neighbors, “horror” distant same galaxy, “accounting” another universe). Spotify song embeddings, Netflix movie×user embeddings power recommendations.
  4. Overfitting & regularization — (section begins, then paywall) the danger of a model learning its training data too well.

What’s behind the paywall

Concepts 11–38 continue into serving/production territory per the pitch (feature stores, real-world architectures, scale/reliability/performance) as subscriber-only “golden member” content. The visible page also carries a partner ad read (AgentField — open-source harness orchestration to call coding agents programmatically from Python/TS/Go) — sponsored insert, not part of the guide.

Interview takeaways worth stealing

  • Structure answers around the pipeline stages, not model choice.
  • Know the two “suspiciously good accuracy” traps (leakage, imbalance) cold.
  • Mentioning monitoring/versioning/pipeline ops beats reciting architectures.