- Overview
- Features
- Architecture
- Tech Stack
- Installation
- Usage
- Machine Learning Models
- API Documentation
- Browser Extension
- Datasets
- Results & Performance
- UI
- Team
PhishTank is an intelligent phishing detection system that combines machine learning with real-time protection. It provides a dual-layer defense mechanism:
- URL Phishing Detection - Analyzes suspicious URLs using Logistic Regression with TF-IDF vectorization
- Email Phishing Detection - Examines email content using DistilBERT transformer model
The system is accessible through:
- π Browser Extension - Real-time URL scanning while browsing
- π REST API - Integration with any application or service
- π Jupyter Notebooks - Model training and experimentation
- Real-time URL scanning as you browse
- Character-level analysis using TF-IDF with n-grams (3-5)
- High accuracy logistic regression classifier
- Instant notifications for suspicious websites
- Visual indicators (red/green badges) for safety status
- Deep learning-based email content analysis
- DistilBERT transformer for contextual understanding
- Multi-source training on diverse phishing datasets
- Sender, subject, and body comprehensive analysis
- Confidence scores for prediction reliability
- Clean and intuitive popup interface
- Easy-to-understand threat indicators
- Detailed analysis results
- Privacy-focused design
PhishTank/
βββ π Browser Extension (Frontend)
β βββ manifest.json # Extension configuration
β βββ background.js # Service worker
β βββ content.js # Page content analyzer
β βββ ui.js # UI logic
β βββ index.html # Popup interface
β βββ style.css # Styling
β
βββ π Backend (API Server)
β βββ app.py # FastAPI application
β βββ requirements.txt # Python dependencies
β βββ url/ # URL detection module
β β βββ logreg_phishing_model/
β βββ email/ # Email detection module
β βββ distilbert_phishing_model/
β
βββ π Machine Learning
βββ URL_Phishing_Detection.ipynb
βββ phishing_email_analysis_bert.ipynb
- Python 3.8+ - Core programming language
- scikit-learn - URL classification (Logistic Regression)
- Transformers (Hugging Face) - Email classification (DistilBERT)
- PyTorch - Deep learning framework
- TF-IDF Vectorization - Feature extraction for URLs
- FastAPI - High-performance REST API
- Pydantic - Data validation
- Joblib - Model serialization
- JavaScript (ES6+) - Extension logic
- HTML5/CSS3 - User interface
- Chrome Extension API - Browser integration
- Pandas - Data manipulation
- NumPy - Numerical computing
- Matplotlib/Seaborn - Visualization
- Python 3.8 or higher
- Node.js (optional, for development)
- Chrome/Edge browser (for extension)
git clone https://github.com/yourusername/PhishTank.git
cd PhishTankcd backend
pip install -r requirements.txtThe models should be in:
backend/url/logreg_phishing_model/- URL detection modelbackend/email/distilbert_phishing_model/- Email detection model
python app.pyThe API will be available at http://localhost:8000
- Open Chrome/Edge browser
- Navigate to
chrome://extensions/(oredge://extensions/) - Enable "Developer mode"
- Click "Load unpacked"
- Select the
PhishTankdirectory - The extension icon will appear in your toolbar
- Automatic Protection: The extension automatically scans URLs as you browse
- Manual Check: Click the extension icon to manually check the current page
- View Results: Green badge = Safe, Red badge = Phishing detected
- Notifications: Receive instant alerts for suspicious websites
POST http://localhost:8000/predict/url
Content-Type: application/json
{
"url": "https://example-suspicious-site.com"
}Response:
{
"url": "https://example-suspicious-site.com",
"prediction": "phishing",
"label": 1,
"timestamp": "2025-10-12T10:30:00"
}POST http://localhost:8000/predict/email
Content-Type: application/json
{
"sender": "noreply@suspicious.com",
"subject": "Urgent: Verify your account",
"body": "Dear user, click here to verify..."
}Response:
{
"prediction": "phishing",
"confidence": 0.95,
"label": 1,
"processed_date": "2025-10-12T10:30:00"
}GET http://localhost:8000/healthAlgorithm: Logistic Regression with TF-IDF Vectorization
Features:
- Character-level n-grams (3-5) capture URL patterns
- 5000 most important features
- TF-IDF weighting for feature importance
Training Process:
# Feature extraction
vectorizer = TfidfVectorizer(
max_features=5000,
analyzer='char_wb',
ngram_range=(3, 5)
)
# Model training
model = LogisticRegression(max_iter=200)
model.fit(X_train, y_train)Why This Works:
- URLs have distinctive character patterns
- Phishing URLs often use character substitutions (e.g., "paypa1.com")
- N-grams capture these subtle variations
Algorithm: DistilBERT (Distilled BERT)
Architecture:
- Pre-trained
distilbert-base-uncasedmodel - Fine-tuned on phishing email datasets
- Sequence classification with 2 classes (legitimate/phishing)
Training Process:
# Model initialization
model = DistilBertForSequenceClassification.from_pretrained(
'distilbert-base-uncased',
num_labels=2
)
# Training with Hugging Face Trainer
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=test_dataset
)Why This Works:
- DistilBERT understands context and semantics
- Captures sophisticated phishing language patterns
- Faster than BERT while maintaining 97% of performance
- Source: Web Page Phishing Detection Dataset
- Size: 10,000+ URLs
- Features: URLs with legitimate/phishing labels
- Split: 80% training, 20% testing
- Source: Phishing Email Dataset
- Components:
- CEAS_08.csv
- Nazario.csv
- Nigerian_Fraud.csv
- SpamAssasin.csv
- TREC_06.csv
- Size: 30,000+ emails
- Features: Sender, subject, body, and labels
- Split: 80% training, 20% testing
| Metric | Score |
|---|---|
| Accuracy | 96.5% |
| Precision | 95.8% |
| Recall | 97.2% |
| F1-Score | 96.5% |
Confusion Matrix:
- True Positives: High detection of phishing URLs
- False Positives: Minimal legitimate sites flagged
- False Negatives: Very few phishing URLs missed
| Metric | Score |
|---|---|
| Accuracy | 98.2% |
| Precision | 97.9% |
| Recall | 98.5% |
| F1-Score | 98.2% |
Key Insights:
- Excellent performance on diverse phishing patterns
- Robust to different email formats
- High confidence in predictions
The browser extension features a clean, intuitive interface:
- Status Indicator: Clear visual feedback (red/green)
- Threat Level: Shows confidence in detection
- Analysis Details: Breakdown of detection reasoning
- Quick Actions: Easy reporting and feedback options
Main extension popup with threat detection interface |
URL scanning functionality |
Email phishing detection |
Email legitimate detection |
Whitelist management for URLs |
Blacklist management |
- Local Processing: URL analysis happens locally when possible
- No Data Storage: We don't store your browsing history
- Secure API: HTTPS encryption for all communications
- Open Source: Full transparency in our code
- Real-time website content analysis
- Machine learning model updates
- Multi-language support
- Mobile app development
- Integration with email clients
- Community reporting system
- Advanced analytics dashboard
Team Code Red - ONL 553
This project was developed as part of the IBM Z Datathon 2025, focusing on cybersecurity and AI-driven threat detection.






