Lesson 6 of 6

An ML system design checklist

Dev's meal-kit cancellation model is finished. It was built in a reproducible pipeline, versioned, wrapped in a serving endpoint, and wired to a drift monitor. On held-out data it looks good. So Dev's boss asks the only question left: is it safe to turn on?

A model that scores well in a notebook is not the same thing as a system that is safe to run. Between the two sits a short list of design questions, the same six every time, that decide whether a model helps or quietly does harm once it touches real subscribers and real money. This lesson turns everything you built across this course into that checklist.

By the end of this lesson you will be able to:

  • Reframe a model as a weekly decision, and hold it to the baseline it must beat
  • Set the operating threshold from what a wrong call actually costs, not the default 0.5
  • Catch a feature that leaks the future or goes missing at serving time
  • Specify how the model serves, what happens when it cannot, and what you watch after launch

Prerequisites: the rest of this course, you can [fit a model and read predict](Your-First-End-to-End-Model-in-R.html), you know batch vs real-time serving and monitoring with PSI, and you can read a confusion matrix.

Why a checklist

A validated model is not yet a system

In Lesson 5 you watched a launched model decay as the world drifted, with no error to warn you. That is one of many ways a model that passed every offline check still fails in production. The others are just as quiet: a metric that does not match the decision, a feature that is missing when you actually score someone, a serving path with no plan for when the model is down.

None of these show up in a notebook, because a notebook has the whole labelled dataset in memory and no real users. Production has neither. So before you ship, you answer a fixed set of questions, the same six every time, each one a way a model has burned somebody before.

Key Insight
Shipping a model is not "is the accuracy high enough?" It is "have we answered every question on the checklist?" A high score with an unanswered question on the list is how good models cause bad outcomes.

The six questions, and where each one draws on this course:

  1. Decide - what action does a score trigger, and what baseline must it beat?
  2. Cost - what does a wrong call cost, and what threshold minimizes it?
  3. Data - where does each feature come from at serving time (Lessons 1-2)?
  4. Serve - batch or real time, how fast, and what is the fallback (Lessons 3-4)?
  5. Monitor - what do you log and watch, and what trips an alarm (Lesson 5)?
  6. Document - the one page that records every answer above.

We will walk them in order, on Dev's model, and end with the page that records the answers.