Email Inspector

Quick Samples:

Ready for Analysis

Paste email text or select one of the quick samples above, then click Classify Email.

83.80%
Test Set Accuracy
Evaluated on 920 holdout test emails
3,681
Training Samples
80% stratified split
57
Extracted Features
48 words + 6 symbols + 3 capital runs

Confusion Matrix (2 × 2)

Actual Class Predicted Class
Predicted Ham (0) Predicted Spam (1)
Actual Ham (0) 416 True Negative 135 False Positive
Actual Spam (1) 14 False Negative 355 True Positive

Classification Report

Class Precision Recall F1-Score Support
Ham (0) 0.97 0.75 0.85 551
Spam (1) 0.72 0.96 0.83 369
Architecture Analysis

Why Gaussian Naive Bayes achieves ~84% & How to Reach 98%+

Gaussian Naive Bayes assumes features follow a continuous Gaussian bell curve and are conditionally independent. Real-world email text has sparse word counts and non-linear patterns. Here are the top ways to dramatically boost classification accuracy:

01

Multinomial Naive Bayes + TF-IDF

Switching from GaussianNB to MultinomialNB with full vocabulary TF-IDF vectorization matches discrete word occurrences and reduces false positives on legitimate emails.

02

N-Gram Phrasing Features

Spammers often use multi-word phrases like "urgent wire transfer", "click the link below", or "verify password". Extracting 2-gram and 3-gram tokens captures these exact cues.

03

Header & Authentication Checks

Incorporate email metadata checks: SPF, DKIM, DMARC pass/fail status, sender domain age, and mismatched Reply-To addresses.

04

URL & Domain Reputation Analysis

Detect hidden hyperlinks, IP address URLs, known shorteners (bit.ly, tinyurl), and cross-reference domains against real-time anti-phishing blacklists.

05

Ensemble or Gradient Boosting

Training a Random Forest or Logistic Regression alongside Naive Bayes resolves feature correlation errors, commonly lifting accuracy from 84% to 95%+.

06

Transformer / ONNX Embeddings

Use Node.js ONNX Runtime with a lightweight model (e.g. all-MiniLM-L6-v2) for semantic deep learning that understands contextual intent rather than just word frequencies.

REST API Integration

You can classify emails programmatically by sending HTTP requests to the backend.

POST /api/classify

Accepts email text and returns predictions, probabilities, and 57-feature breakdown.

cURL Request Example
curl -X POST http://localhost:3000/api/classify \
  -H "Content-Type: application/json" \
  -d '{ "text": "Win $1,000,000 free cash prize now! Reply with credit info." }'
JSON Response Example
{
  "success": true,
  "prediction": "Spam",
  "isSpam": true,
  "confidence": 99.82,
  "probabilities": {
    "spam": 99.82,
    "ham": 0.18
  },
  "riskLevel": "High",
  "processingTimeMs": 0.42,
  "featureAnalysis": {
    "totalWords": 10,
    "detectedKeywords": [
      { "word": "free", "frequency": 10.0, "isSpamIndicator": true },
      { "word": "credit", "frequency": 10.0, "isSpamIndicator": true }
    ],
    "capitalStats": {
      "average": 3.0,
      "longest": 3,
      "total": 3
    }
  }
}