Email Inspector
Ready for Analysis
Paste email text or select one of the quick samples above, then click Classify Email.
Confusion Matrix (2 × 2)
| Actual Class | Predicted Class | |
|---|---|---|
| Predicted Ham (0) | Predicted Spam (1) | |
| Actual Ham (0) | 416 True Negative | 135 False Positive |
| Actual Spam (1) | 14 False Negative | 355 True Positive |
Classification Report
| Class | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Ham (0) | 0.97 | 0.75 | 0.85 | 551 |
| Spam (1) | 0.72 | 0.96 | 0.83 | 369 |
Why Gaussian Naive Bayes achieves ~84% & How to Reach 98%+
Gaussian Naive Bayes assumes features follow a continuous Gaussian bell curve and are conditionally independent. Real-world email text has sparse word counts and non-linear patterns. Here are the top ways to dramatically boost classification accuracy:
Multinomial Naive Bayes + TF-IDF
Switching from GaussianNB to MultinomialNB with full vocabulary TF-IDF vectorization matches discrete word occurrences and reduces false positives on legitimate emails.
N-Gram Phrasing Features
Spammers often use multi-word phrases like "urgent wire transfer", "click the link below", or "verify password". Extracting 2-gram and 3-gram tokens captures these exact cues.
Header & Authentication Checks
Incorporate email metadata checks: SPF, DKIM, DMARC pass/fail status, sender domain age, and mismatched Reply-To addresses.
URL & Domain Reputation Analysis
Detect hidden hyperlinks, IP address URLs, known shorteners (bit.ly, tinyurl), and cross-reference domains against real-time anti-phishing blacklists.
Ensemble or Gradient Boosting
Training a Random Forest or Logistic Regression alongside Naive Bayes resolves feature correlation errors, commonly lifting accuracy from 84% to 95%+.
Transformer / ONNX Embeddings
Use Node.js ONNX Runtime with a lightweight model (e.g. all-MiniLM-L6-v2) for semantic deep learning that understands contextual intent rather than just word frequencies.
REST API Integration
You can classify emails programmatically by sending HTTP requests to the backend.
/api/classifyAccepts email text and returns predictions, probabilities, and 57-feature breakdown.
curl -X POST http://localhost:3000/api/classify \
-H "Content-Type: application/json" \
-d '{ "text": "Win $1,000,000 free cash prize now! Reply with credit info." }'
{
"success": true,
"prediction": "Spam",
"isSpam": true,
"confidence": 99.82,
"probabilities": {
"spam": 99.82,
"ham": 0.18
},
"riskLevel": "High",
"processingTimeMs": 0.42,
"featureAnalysis": {
"totalWords": 10,
"detectedKeywords": [
{ "word": "free", "frequency": 10.0, "isSpamIndicator": true },
{ "word": "credit", "frequency": 10.0, "isSpamIndicator": true }
],
"capitalStats": {
"average": 3.0,
"longest": 3,
"total": 3
}
}
}