- by x32x01 ||
The model says: P(Churn) = 0.87
Does that mean the customer will leave? ❌
Not necessarily.
This is where real-world Machine Learning gets more interesting. A model can produce a probability, but turning that number into a reliable decision requires much more than checking Accuracy.
Imagine we are building a customer churn model for a telecom company.
We have:
And the model predicts: Customer #1024
A beginner might immediately ask:
"Is the model accurate?"
A Data Scientist asks a much longer list of questions.
Why does this matter?
Because you might discover a heavily imbalanced target:
A model that predicts
For example:
But if your actual goal is to identify customers who are likely to leave, that model may have very little practical value.
📌 Accuracy alone is not enough.
You need to understand what the model is actually getting right and wrong.
Suppose your dataset contains:
You might create your features like this:
But there is a problem.
That means the model is effectively getting a glimpse of the future.
This is Data Leakage.
The result can look impressive during training or testing, while the model performs much worse in real-world use.
📌 The rule is simple:
Any information that would not have been available at the time of the decision should not be used as a feature for that decision.
For example:
Conceptually, the process becomes:
This is often closer to the real-world scenario:
You train using information from the past, then ask the model to make predictions about data it has not seen yet.
That is much more useful than accidentally allowing future information to influence the past.
For a real dataset, categorical and other non-numeric features should be encoded appropriately before they are passed to
Notice the use of:
This helps preserve approximately the same class distribution in the training and test sets.
For example:
Why use
Because these two methods answer different questions.
While:
For example:
And now we are back to the number from the beginning: 0.87
But what does it actually mean?
It does not automatically mean:
"This customer will leave with an 87% certainty."
And it does not guarantee that 87 out of every 100 customers receiving this prediction will actually churn.
That is why Probability Calibration matters.
You can examine calibration with tools such as:
Suppose a group of predictions is around:
If the model is well calibrated, we would expect the actual churn rate for comparable groups to be reasonably close to 80% over repeated observations.
This is the basic idea behind Probability Calibration.
A model can rank customers reasonably well while still producing probabilities that are poorly calibrated.
as if it were a universal rule.
It is not.
A classification threshold determines when a probability becomes a positive prediction.
For example:
Why would you change it?
Because the cost of different mistakes may not be the same.
Consider Customer Churn.
False Negative
The model predicts:
But the customer actually leaves.
The company may lose the customer.
False Positive
The model predicts:
But the customer stays.
The company may spend money on a retention offer that was not necessary.
So the threshold should be connected to the Business Cost of these errors, not selected only because 0.5 is a common default.
start asking better questions:
Now the question changes.
Instead of asking:
"What's the best Accuracy?"
you start asking:
"Which decision best meets the business objective while considering the cost of errors?"
That is a much more useful Machine Learning question.
and evaluated using:
Now imagine that several months later, the following things change:
This can lead to problems such as: Data Drift and Concept Drift.
That is why Machine Learning does not end at:
A real ML system may look more like:
The model is part of a larger system that needs to be monitored after deployment.
When a model tells you: P(Churn) = 0.87
don't immediately ask: "Is the model correct?"
Ask:
It is a complete process:
And that is the point where you move from:
"I know Machine Learning algorithms."
to:
"I can build and evaluate a Machine Learning system for a real-world problem."
Does that mean the customer will leave? ❌
Not necessarily.
This is where real-world Machine Learning gets more interesting. A model can produce a probability, but turning that number into a reliable decision requires much more than checking Accuracy.
Imagine we are building a customer churn model for a telecom company.
We have:
X = customer_featuresy = churnAnd the model predicts: Customer #1024
P(Churn) = 0.87A beginner might immediately ask:
"Is the model accurate?"
A Data Scientist asks a much longer list of questions.
🔎 1. Is the Data Clean?
Before training any model, inspect the data. Python:
df.isna().sum()
df.duplicated().sum()
df.dtypes
df["Churn"].value_counts(normalize=True) Because you might discover a heavily imbalanced target:
Code:
Churn = No 95%
Churn = Yes 5% No for almost everyone could achieve: 95% AccuracyFor example:
Code:
No
No
No
No
No
... 📌 Accuracy alone is not enough.
You need to understand what the model is actually getting right and wrong.
🚨 2. Is There Data Leakage?
Data Leakage is one of the most dangerous problems in Machine Learning.Suppose your dataset contains:
Code:
monthly_charge
contract_type
tenure
support_calls
churn_date Python:
X = df.drop(columns=["Churn"]) churn_date may contain information that only became available after the customer actually churned.That means the model is effectively getting a glimpse of the future.
This is Data Leakage.
The result can look impressive during training or testing, while the model performs much worse in real-world use.
📌 The rule is simple:
Any information that would not have been available at the time of the decision should not be used as a feature for that decision.
📅 3. Don't Train Before Thinking About Time
If time matters in your problem, your train and test strategy should reflect how the model will be used in Production.For example:
Python:
train = df[df["month"] < "2026-07"]
test = df[df["month"] >= "2026-07"] Code:
PAST DATA
↓
TRAIN
↓
MODEL
↓
FUTURE DATA
↓
TEST You train using information from the past, then ask the model to make predictions about data it has not seen yet.
That is much more useful than accidentally allowing future information to influence the past.
🤖 4. Build a Logistic Regression Model
Now we can build a simple classification model using Scikit-learn.For a real dataset, categorical and other non-numeric features should be encoded appropriately before they are passed to
LogisticRegression. Assuming X contains suitable numeric features, a basic example looks like this: Python:
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import Pipeline
X = df.drop(columns=["Churn"])
y = df["Churn"]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
stratify=y,
random_state=42
)
model = Pipeline([
("scaler", StandardScaler()),
("classifier", LogisticRegression())
])
model.fit(X_train, y_train) stratify=yThis helps preserve approximately the same class distribution in the training and test sets.
📊 5. Don't Ask Only: "What's the Accuracy?"
A useful evaluation should look at more than one metric.For example:
Python:
from sklearn.metrics import (
classification_report,
confusion_matrix,
roc_auc_score
)
y_pred = model.predict(X_test)
y_prob = model.predict_proba(X_test)[:, 1]
print(confusion_matrix(y_test, y_pred))
print(classification_report(
y_test,
y_pred
))
print(
"ROC-AUC:",
roc_auc_score(y_test, y_prob)
) predict_proba()?Because these two methods answer different questions.
model.predict(X_test) gives you a class prediction, such as: Code:
0 or 1 model.predict_proba(X_test) gives you probabilities.For example:
Code:
P(No Churn) = 0.13
P(Churn) = 0.87 And now we are back to the number from the beginning: 0.87
But what does it actually mean?
🎯 6. What Does P(Churn) = 0.87 Actually Mean?
A probability such as0.87 represents the model's estimated probability for the positive class, assuming the model and probability estimation are appropriate for the problem.It does not automatically mean:
"This customer will leave with an 87% certainty."
And it does not guarantee that 87 out of every 100 customers receiving this prediction will actually churn.
That is why Probability Calibration matters.
You can examine calibration with tools such as:
Python:
from sklearn.calibration import calibration_curve
prob_true, prob_pred = calibration_curve(
y_test,
y_prob,
n_bins=10
) P(Churn) ≈ 0.80If the model is well calibrated, we would expect the actual churn rate for comparable groups to be reasonably close to 80% over repeated observations.
This is the basic idea behind Probability Calibration.
A model can rank customers reasonably well while still producing probabilities that are poorly calibrated.
⚖️ 7. The 0.5 Threshold Is Not a Law
Many beginners treat: 0.5as if it were a universal rule.
It is not.
A classification threshold determines when a probability becomes a positive prediction.
For example:
Python:
threshold = 0.30
y_pred_custom = (
y_prob >= threshold
).astype(int) Because the cost of different mistakes may not be the same.
Consider Customer Churn.
False Negative
The model predicts:
No ChurnBut the customer actually leaves.
The company may lose the customer.
False Positive
The model predicts:
ChurnBut the customer stays.
The company may spend money on a retention offer that was not necessary.
So the threshold should be connected to the Business Cost of these errors, not selected only because 0.5 is a common default.
💡 8. This Is Where Data Science Becomes Real
Instead of stopping at: Model Accuracy = 94%start asking better questions:
- How many customers did we identify correctly?
- How many customers did we miss?
- What is the cost of a False Positive?
- What is the cost of a False Negative?
- Does the model work on new customers?
- Are the predicted probabilities calibrated?
- Does performance change over time?
- Does the chosen threshold match the business objective?
Python:
from sklearn.metrics import precision_score, recall_score
for threshold in [0.2, 0.3, 0.4, 0.5, 0.6, 0.7]:
pred = (y_prob >= threshold).astype(int)
<span>precision = precision_score(<br> y_test,<br> pred<br>)<br><br>recall = recall_score(<br> y_test,<br> pred<br>)<br><br>print(<br> threshold,<br> precision,<br> recall<br>)</span> Instead of asking:
"What's the best Accuracy?"
you start asking:
"Which decision best meets the business objective while considering the cost of errors?"
That is a much more useful Machine Learning question.
📈 9. The Model Can Change Over Time
Suppose your model was trained using data from: Code:
2025 Code:
2026 Now imagine that several months later, the following things change:
- Customer Behavior
- Pricing
- Products
- Competitors
- Economic Conditions
This can lead to problems such as: Data Drift and Concept Drift.
That is why Machine Learning does not end at:
model.fit(...)A real ML system may look more like:
Code:
Data
↓
Validation
↓
Training
↓
Evaluation
↓
Deployment
↓
Monitoring
↓
Drift Detection
↓
Retraining 💙 The Main Lesson
When a model tells you: P(Churn) = 0.87
don't immediately ask: "Is the model correct?"
Ask:
- Is the data correct?
- Is there Data Leakage?
- Were the features actually available when the decision was made?
- Did we choose the right evaluation metrics?
- Are the probabilities calibrated?
- Is the threshold appropriate for the business cost?
- Does the model generalize to new data?
- Does its performance remain stable over time?
It is a complete process:
Code:
Data
→
Features
→
Model
→
Probability
→
Evaluation
→
Decision
→
Monitoring "I know Machine Learning algorithms."
to:
"I can build and evaluate a Machine Learning system for a real-world problem."
❓ Frequently Asked Questions
----------------------Does P(Churn) = 0.87 mean the customer will definitely leave?
No. It means the model assigns an estimated probability of 0.87 to the positive class. Whether that probability is well calibrated depends on the model, data, and evaluation.Is 95% Accuracy a good result for a churn model?
Not necessarily. If 95% of customers do not churn, a model that predicts "No Churn" for everyone could achieve 95% Accuracy while failing to identify customers who actually leave.Why is Data Leakage dangerous?
Because it allows information that would not have been available at prediction time to influence the model. This can produce unrealistically strong evaluation results and poor performance in real-world use.Should I always use a 0.5 classification threshold?
No. The appropriate threshold depends on the problem, the model's probabilities, and the relative costs of False Positives and False Negatives.Why should a Machine Learning model be monitored after deployment?
Because real-world data and behavior can change. Data Drift and Concept Drift can affect model performance over time, making monitoring and possible retraining necessary. Last edited: