What it is
Model drift is the gradual loss of accuracy in an AI model as the real-world data it scores against shifts over time. Market conditions change, candidate pools grow or shrink, companies behave differently than they did a year ago, and a model trained or tuned on an earlier period keeps applying patterns that no longer hold. Nothing crashes and no error appears. The score still looks exactly like a normal number. It is simply, quietly, more wrong than it used to be.
That is what separates drift from an AI hallucination. A hallucination is one output, a single invented fact or signal a model generates with unearned confidence. Drift is systemic: a slow accuracy decline across many outputs over weeks or months, caused by a shifting world rather than a single bad guess. It is a governance and monitoring concept, not a one-off error you catch and move past.
Model drift never announces itself. The score still looks normal. It is just increasingly wrong.
Why it matters
Lead scoring, ICP-fit and candidate-matching models are tuned on patterns from a given period: which signals used to convert, which accounts used to close, which candidates used to get placed. Markets move. A sector that was hiring aggressively cools off, a candidate pool thins out or floods, a company's buying behaviour changes after a funding round or a leadership change. If the model behind the score is never revisited, its rankings start reflecting a market that no longer exists.
The real risk is that nothing flags it. The number keeps arriving in the same format, a consultant keeps trusting it out of habit, strong accounts quietly slide down the list, and effort keeps going toward patterns that used to convert and have stopped. Nobody notices until reply rate or win rate has already slipped for weeks, by which point the cost was not one bad call but a slow bleed across every call the model made.
How boilr handles it
boilr treats model drift as something to monitor continuously, not a problem fixed once on a retraining schedule and forgotten. Every scoring and matching model is checked against real recruiter outcomes, placements, replies and the corrections a consultant actually makes, flowing back through the Company Brain. When a model's calls stop lining up with what is actually converting, that gap becomes visible instead of staying hidden inside a score that still looks normal.
Confidence scores and human-in-the-loop review double as an early tripwire: a rise in low-confidence detections or corrections on a particular type of account signals that the model's assumptions no longer match the market, which triggers retraining before the gap costs a desk its pipeline. Because that history sits in the Company Brain rather than one consultant's head, the correction holds even when the desk changes hands.