Haydra Prov2.1.2
PREDATOR · BadDomains · DomainProfiler · Triplet Embeddings · DNS Traffic · DGA Ensemble · Live SURBL
Authorized use only · Trust & Safety
Settings ⓘ Auto-saved to browser
No saved settings
🔍 Checker
📊 Dashboard
🏢 Registrar Intel
🔎 Pattern Recognition
🧠 Models
🌐 SURBL TLDs
🔑 Keywords
📈 Trending Keywords
⛏ Keyword Mining
🏷 Brands
Input 0
Drop CSV/TXT or click to browse
SURBL not loaded
Results

Paste or upload domains, then click Analyse domains.

Run an analysis first to populate the dashboard.

Run an analysis first to populate registrar intelligence.

Corpus upload Independent from Checker — accepts any domain list (SURBL / suspended / active / mixed)
Columns: Domain · Registrar · Transaction Date · Detection Date · Status · Source  (all beyond Domain optional)

Upload a corpus above, then click Analyse corpus to detect clusters and campaigns.

Research model ensemble — signal weights

Each slider controls how much that research model's signal contributes to the final composite abuse score. Higher = more influence.

Campaign patterns (APWG) 0
Model descriptions & research citations
PREDATOR — Time-of-Registration (Hao et al., ACM CCS 2016)

Signal: Burst registration score, SLD textual similarity to batch-mates, structural name features, registrar reputation.
Key finding: 70% detection rate, 0.35% FPR, predicts abuse days/weeks before DNS blacklists. Miscreants register many domains in bursts to ensure profitability and attack agility.
Implementation here: Burst proxy via SLD numeric suffix patterns (brand+digits, sequential naming), structural name diversity score, registrar risk heuristic.

BadDomains — Phishing Registration Signals (CERT Polska / PMC 2025)

Signal: Impersonation target presence in SLD, suspicious registrant contact patterns, TF-IDF on character n-grams, domain lifecycle anomalies.
Key finding: LightGBM model using trademark presence + registrant reuse history achieved highest F1 among evaluated systems on .pl phishing registry data. Data drift noted across 2023–2025 deployment period.
Implementation here: Brand trademark match scoring (direct + fuzzy), registrant anomaly proxy via privacy-indicator patterns, character n-gram rarity (TF-IDF approximation).

DomainProfiler — Temporal Variation Patterns (Chiba et al., IJIS 2018)

Signal: How and when a domain appears/disappears from legitimate vs malicious lists. Survival analysis for domain lifespan.
Key finding: Predicted malicious domains 220 days beforehand with TPR 0.985. Domains that appear on Alexa and malicious lists simultaneously show highest abuse probability.
Implementation here: TVP proxy via domain age + list-status heuristic, suspicious activation pattern scoring (new TLD + high-risk SLD = fast-abuse profile).

Multimodal Triplet Embeddings (Casino et al., RAID 2021)

Signal: Character-level, n-gram, and structural embeddings mapped into a triplet space where malicious domains cluster near known-bad anchors.
Key finding: Triplet embedding approach achieves better separation of DGA families and targeted phishing infrastructure than single-modality classifiers. Embedding distance to known-bad anchor <0.3 is a strong malicious signal.
Implementation here: Simulated via multi-dimensional distance scoring — n-gram distance to known-bad patterns + structural similarity to documented attack templates + brand proximity in embedding-like combined space.

DNS Traffic Classification (Bilge et al. EXPOSURE, NDSS 2011 + Sajid et al. 2021)

Signal: Passive DNS query patterns — NX ratio, TTL variance, IP diversity, query volume spikes.
Key finding: EXPOSURE showed malicious domains have distinct DNS behavior: short TTLs, high NX ratios, and rapid IP churn. Batch + online ML models on DNS logs achieve >95% classification accuracy.
Implementation here: DNS behavior proxy via TLD fast-flux likelihood, IP churn proxy via hosting infrastructure signals, TTL-proxy via domain age heuristic.

DGA Multi-Model Ensemble (Woodbridge 2016 + Harishkumar 2024 + Zhao 2023 + Vranken 2022 + Sun & Liu 2023)

Five-signal DGA ensemble: Shannon entropy (char randomness), N-gram rarity (bigram/trigram vs English corpus), vowel-consonant ratio (phonics deviation), TF-IDF lexical score (legitimate SLD corpus distance), consecutive consonant/digit run detection. 2024 consensus: GRU+attention BiLSTM achieves 99% accuracy. Implemented as weighted composite.

SURBL TLD Abuse Rates (ICANN DAAR 2023 + SURBL live 2025)

SURBL publishes monthly TLD abuse counts at surbl.org/static/tld-abuse-complete-rankings.txt. Current top abusers: .top (657K), .info (427K), .cn (126K), .cc (123K), .vip (100K). Scores normalised 0–100 against highest-ranked TLD and fed directly into TLD signal weight.

SURBL TLD abuse rankings

Source: https://www.surbl.org/static/tld-abuse-complete-rankings.txt — updated monthly.

Not loaded
CSAM indicators restricted — matches trigger immediate block + mandatory NCMEC CyberTipline referral.
Malware0
Malvertising0
CSAM — restricted0
Restricted. Matches: block + NCMEC referral.
Phishing / fraud0
Suspicious / general0
Keyword Mining Engine
Fetches live domain threat feeds → tokenises SLDs → scores tokens via TF-IDF against Tranco baseline → surfaces emerging abuse vocabulary by category
Sources: URLhaus · PhishTank · ThreatFox · OpenPhish · abuse.ch
Feed sources & extraction settings
URLhaus Malware
abuse.ch bulk CSV — 500K+ malware-distributing URLs tagged by malware family (LummaC2, AsyncRAT, Remcos, FormBook…)
api.urlhaus.abuse.ch/v1/urls/recent/
ThreatFox Malware IOCs
abuse.ch IOC feed — domains/URLs tagged by malware family, confidence score, threat type (botnet_cc, payload_delivery…)
threatfox-api.abuse.ch/api/v1/
PhishTank Phishing
Community-verified phishing URLs with brand target labels — ideal for brand+action keyword co-occurrence extraction.
data.phishtank.com/data/online-valid.csv.gz
OpenPhish Phishing
Community phishing feed — plain text URL list updated every ~4 hours. No brand labels but high freshness.
openphish.com/feed.txt
PhishStats Phishing
Phishing URL database with score, IP, country, ASN. Provides good geographic + structural signal for SLD vocabulary.
phishstats.info/phishstats.csv
Feodo Tracker Botnet C2
abuse.ch botnet C2 tracker — domains/IPs for Emotet, Dridex, TrickBot, Qakbot, IcedID. Botnet infrastructure vocabulary.
feodotracker.abuse.ch/downloads/domainblocklist.txt
Ignore tokens appearing fewer times
Skip very long tokens (URLs, hashes)
Skip very short noise tokens
Keywords shown per abuse category
⭐ Upload your own suspended domain list (highest-signal source)
Upload a CSV or TXT of domains your registry has suspended for abuse. Tokens extracted from confirmed-abuse domains carry stronger signal than any external feed — this is your registry's own abuse fingerprint. Supports: one domain per line (TXT), or CSV with configurable columns.
Drop suspended domain CSV/TXT here or click to browse
💡 Your suspended domain list is the highest-quality source — every domain is confirmed abuse. External live feed fetches may be CORS-blocked in browser; use Run demo to preview the pipeline.
Add brand
Two columns: Name, Domain — one brand per row — header row optional
Brands 0