A spam classifier is one of the best first machine-learning projects: the data is easy to find, the model trains in seconds on a laptop, and the results are easy to understand. This tutorial builds one in Python with scikit-learn and a Naive Bayes model, measures it properly, saves it, and turns it into a small web API you can deploy.
Load a labelled set of spam and normal messages, split it into training and test sets, and put a TfidfVectorizer and a MultinomialNB classifier in one scikit-learn Pipeline. Tune the alpha setting with cross-validation, judge the model on precision and recall rather than accuracy, and save the whole pipeline with joblib. Train on your own computer; serve the saved model from a small Flask API on a platform built for apps.
1. What you need
- Python 3.12 or newer on your computer, in a virtual environment.
- Packages: install them with
pip install scikit-learn pandas joblib flask gunicorn. - A labelled dataset. Two well-known free ones: the SMS Spam Collection from the UCI Machine Learning Repository (about 5,500 text messages, each marked
spamorham) and the SpamAssassin public corpus (real emails in folders). This tutorial uses the SMS set because it is a single file.
The SMS file has one message per line: the label, a tab, then the text.
2. Load and split the data
import csv
import pandas as pd
from sklearn.model_selection import train_test_split
df = pd.read_csv(
"SMSSpamCollection",
sep="\t",
header=None,
names=["label", "text"],
quoting=csv.QUOTE_NONE,
)
X = df["text"]
y = (df["label"] == "spam").astype(int) # 1 = spam, 0 = ham
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42
)
print(y.value_counts())Two details matter here. quoting=csv.QUOTE_NONE stops pandas from treating quote marks inside messages as field delimiters. stratify=y keeps the same share of spam in the training and test sets; only about 13% of this dataset is spam, and a random split can skew that.
3. Turn text into numbers
A model can't read words, so each message becomes a vector of numbers. TF-IDF (term frequency, inverse document frequency) gives each word a weight: high when it appears often in this message but rarely across all messages. Words like "free", "prize" and "claim" end up carrying a lot of signal; "the" and "you" carry almost none.
TfidfVectorizer does the cleaning for you: it lowercases the text, strips punctuation and splits it into words. Useful settings:
ngram_range=(1, 2)also counts two-word phrases such as "call now";min_df=2ignores words that appear in only one message;sublinear_tf=Truestops a word repeated ten times from counting ten times as much.
4. Build and train the model
Multinomial Naive Bayes works well on word counts and weights. It applies Bayes' theorem with the "naive" assumption that words appear independently of each other. That assumption is wrong, but the model still classifies text well and trains almost instantly.
Put the vectorizer and the model in one Pipeline, so the exact same text processing is used in training and in production:
from sklearn.pipeline import Pipeline
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.model_selection import GridSearchCV
pipe = Pipeline([
("tfidf", TfidfVectorizer(ngram_range=(1, 2), min_df=2, sublinear_tf=True)),
("nb", MultinomialNB()),
])
search = GridSearchCV(
pipe,
{"nb__alpha": [0.01, 0.05, 0.1, 0.5, 1.0]},
scoring="f1",
cv=5,
)
search.fit(X_train, y_train)
model = search.best_estimator_
print("Best alpha:", search.best_params_["nb__alpha"])alpha is the smoothing setting: it controls how much weight the model gives to words it saw rarely. GridSearchCV tries each value with five-fold cross-validation on the training data only, so the test set stays unseen.
5. Measure it properly
from sklearn.metrics import classification_report, confusion_matrix
pred = model.predict(X_test)
print(classification_report(y_test, pred, target_names=["ham", "spam"]))
print(confusion_matrix(y_test, pred))| Metric | What it answers | Why it matters for spam |
|---|---|---|
| Accuracy | What share of all messages was right? | Misleading: marking everything "ham" already scores about 87% here |
| Precision (spam) | Of messages marked spam, how many really were? | Low precision means real messages are lost |
| Recall (spam) | Of all real spam, how much was caught? | Low recall means spam reaches the inbox |
| F1 | Balance of precision and recall | A single number for comparing models |
The confusion matrix shows the four counts behind these numbers. The number to watch is false positives, real messages marked as spam, because a customer's lost enquiry costs more than one extra spam message.
predict uses a 50% cut-off. Use model.predict_proba(texts)[:, 1] to get a spam score between 0 and 1 and flag only messages above, say, 0.8. A higher threshold raises precision and lowers recall; pick the balance your use case needs.
6. Improve it
- Try
ComplementNB. A Naive Bayes variant designed for imbalanced text data; swap it into the pipeline and compare F1. - Try a linear model.
LogisticRegressionorLinearSVCon the same TF-IDF features often beat Naive Bayes by a small margin. - Add features. Message length, the share of capital letters, or whether the text contains a link or a phone number.
- Retrain regularly. Spammers change their wording, so a model trained once slowly gets worse. Add newly labelled messages and retrain.
7. Save the model and serve it as an API
Save the whole pipeline, not just the classifier, so the vectorizer travels with it:
import joblib
joblib.dump(model, "spam_model.joblib")joblib and pickle files can run code when they are loaded. Only load model files you built yourself or received from a source you fully trust, and never accept model uploads from users.
A minimal Flask API, in app.py:
import joblib
from flask import Flask, jsonify, request
app = Flask(__name__)
model = joblib.load("spam_model.joblib")
@app.post("/classify")
def classify():
data = request.get_json(silent=True) or {}
text = data.get("text")
if not isinstance(text, str) or not text.strip():
return jsonify(error="text is required"), 400
score = float(model.predict_proba([text[:20000]])[0][1])
return jsonify(spam=score >= 0.8, score=round(score, 3))The API checks its input, caps the text length and returns a score, so the caller can choose what to do. In production, run it with Gunicorn rather than Flask's development server, and pin your package versions in requirements.txt: a model saved with one scikit-learn version may not load in another.
8. Running this on Domain India
Train the model on your own computer; training uses more memory and CPU than a hosting account should. Then deploy only the saved model and the API.
- App Platform (recommended for an API). Python apps run from a
Dockerfilein your project, such aspython:3.12-slim, installingrequirements.txtand startinggunicorn app:app --bind 0.0.0.0:$PORT. Your app must listen on thePORTvariable. Deploy from GitHub with Deploy Now, or with a deploy token from your terminal or CI. Each app gets 512 MB of RAM, enough for a small Naive Bayes model. See Getting started with the App Platform. - cPanel or DirectAdmin shared hosting. You can run a small Flask app with the Python app tool; see How to deploy a Python app on shared hosting. scikit-learn and NumPy load into your account's memory limit, so test on your plan before you rely on it.
- VPS. For larger models, scheduled retraining or background workers, a VPS gives you root access. It is self-managed.
- 512 MB RAM per app
- 1 vCPU
- 5 GB NVMe SSD
- PostgreSQL Database
If you simply want less spam in your own mailbox, you don't need to build a model: use the spam filters in your control panel. See Configuring spam filters.
Why use Naive Bayes for spam detection?
It trains in seconds, needs little data, and works well on word counts and TF-IDF weights. It is a strong baseline; try logistic regression or a linear SVM afterwards and compare their F1 scores.
Why is accuracy a poor measure for a spam classifier?
Spam datasets are imbalanced. If 87% of messages are normal, a model that never predicts spam scores 87% accuracy while catching nothing. Precision, recall and F1 on the spam class show the real performance.
What does the alpha parameter do in MultinomialNB?
It is additive smoothing. It stops words that appeared rarely, or never, in one class from pushing a prediction to zero. Tune it with cross-validation on the training data.
Should I save the vectorizer separately from the model?
No. Put both in a scikit-learn Pipeline and save the pipeline as one file, so the same text processing is used in training and in production.
Is it safe to load a joblib or pickle model file?
Only if you created it or fully trust its source. Loading these files can run code, so never load model files uploaded by users or downloaded from unknown sites.
Can I run a scikit-learn API on Domain India hosting?
Yes. The App Platform runs Python apps from a Dockerfile with 512 MB of RAM per app. A small Flask app can also run on cPanel or DirectAdmin shared hosting through the Python app tool, within your account's memory limit.
Ready to put your model online? Read Getting started with the App Platform, compare App Platform plans, or open a support ticket if you are unsure which option fits.
Bring a Dockerfile, deploy from GitHub or with a deploy token, and get PostgreSQL and SSL with every plan.
See App Platform plans