Paste or upload domains, then click Analyse domains.
Paste or upload domains, then click Analyse domains.
Run an analysis first to populate the dashboard.
Run an analysis first to populate registrar intelligence.
Upload a corpus above, then click Analyse corpus to detect clusters and campaigns.
Each slider controls how much that research model's signal contributes to the final composite abuse score. Higher = more influence.
Signal: Burst registration score, SLD textual similarity to batch-mates, structural name features, registrar reputation.
Key finding: 70% detection rate, 0.35% FPR, predicts abuse days/weeks before DNS blacklists. Miscreants register many domains in bursts to ensure profitability and attack agility.
Implementation here: Burst proxy via SLD numeric suffix patterns (brand+digits, sequential naming), structural name diversity score, registrar risk heuristic.
Signal: Impersonation target presence in SLD, suspicious registrant contact patterns, TF-IDF on character n-grams, domain lifecycle anomalies.
Key finding: LightGBM model using trademark presence + registrant reuse history achieved highest F1 among evaluated systems on .pl phishing registry data. Data drift noted across 2023–2025 deployment period.
Implementation here: Brand trademark match scoring (direct + fuzzy), registrant anomaly proxy via privacy-indicator patterns, character n-gram rarity (TF-IDF approximation).
Signal: How and when a domain appears/disappears from legitimate vs malicious lists. Survival analysis for domain lifespan.
Key finding: Predicted malicious domains 220 days beforehand with TPR 0.985. Domains that appear on Alexa and malicious lists simultaneously show highest abuse probability.
Implementation here: TVP proxy via domain age + list-status heuristic, suspicious activation pattern scoring (new TLD + high-risk SLD = fast-abuse profile).
Signal: Character-level, n-gram, and structural embeddings mapped into a triplet space where malicious domains cluster near known-bad anchors.
Key finding: Triplet embedding approach achieves better separation of DGA families and targeted phishing infrastructure than single-modality classifiers. Embedding distance to known-bad anchor <0.3 is a strong malicious signal.
Implementation here: Simulated via multi-dimensional distance scoring — n-gram distance to known-bad patterns + structural similarity to documented attack templates + brand proximity in embedding-like combined space.
Signal: Passive DNS query patterns — NX ratio, TTL variance, IP diversity, query volume spikes.
Key finding: EXPOSURE showed malicious domains have distinct DNS behavior: short TTLs, high NX ratios, and rapid IP churn. Batch + online ML models on DNS logs achieve >95% classification accuracy.
Implementation here: DNS behavior proxy via TLD fast-flux likelihood, IP churn proxy via hosting infrastructure signals, TTL-proxy via domain age heuristic.
Five-signal DGA ensemble: Shannon entropy (char randomness), N-gram rarity (bigram/trigram vs English corpus), vowel-consonant ratio (phonics deviation), TF-IDF lexical score (legitimate SLD corpus distance), consecutive consonant/digit run detection. 2024 consensus: GRU+attention BiLSTM achieves 99% accuracy. Implemented as weighted composite.
SURBL publishes monthly TLD abuse counts at surbl.org/static/tld-abuse-complete-rankings.txt. Current top abusers: .top (657K), .info (427K), .cn (126K), .cc (123K), .vip (100K). Scores normalised 0–100 against highest-ranked TLD and fed directly into TLD signal weight.
Source: https://www.surbl.org/static/tld-abuse-complete-rankings.txt — updated monthly.