Phishing remains one of the most stubborn threats on the internet, and the tools built to stop it have long forced defenders into an uncomfortable choice. Classical machine learning classifiers can explain which features of a suspicious link triggered an alert, but they struggle to keep pace with rapidly evolving attack patterns. Deep learning models, by contrast, achieve impressive accuracy on raw URL strings yet behave as opaque black boxes, handing security analysts nothing more than a label and a confidence score. A newly published study in Discover Informatics proposes a way out of this dilemma, combining a hybrid deep neural network, dual explainable AI techniques, and a carefully sandboxed large language model into a single framework that both detects phishing URLs with near-perfect accuracy and explains, in analyst-style language, exactly why each verdict was reached.
The research team, led by Hoc Minh Le of Konya Technical University together with Van Thuong Nguyen of Eskişehir Technical University, built their system around a multi-branch neural architecture. The first branch ingests raw URL strings at the character level, embedding each token into a 64-dimensional vector space before passing the sequence through a one-dimensional convolutional neural network with 128 filters and then a bidirectional LSTM with 64 hidden units per direction. The convolutional layers act as pattern detectors for suspicious local character combinations, such as brand names or special symbols, regardless of where they appear in the string, while the bidirectional recurrence captures long-range dependencies that span both the domain prefix and the path suffix. The second branch processes twenty handcrafted lexical and domain-based indicators, including URL length, digit and special-character ratios, Shannon entropy of the character distribution, and structural properties of the registered domain, through a small dense network stabilized by batch normalization.
The two branches are fused by simple concatenation into a 160-dimensional tensor, a deliberate design choice the authors justify on interpretability grounds: because the branches operate on fundamentally different input modalities, attention-based fusion would introduce additional learned parameters whose own behavior would require separate explanation, working against the framework’s transparency goals. The final model contains only 168,737 trainable parameters and occupies 0.65 megabytes, making it dramatically lighter than transformer-based alternatives. Classification latency is roughly 2.6 milliseconds per URL on a Tesla T4 GPU, rising to about 399 milliseconds when local explanation is included, and the language model reasoning step, which adds between 500 and 2,000 milliseconds, is invoked selectively only on flagged or borderline instances.
Explainability is delivered through two complementary techniques applied to a Random Forest surrogate model trained on the same twenty-dimensional handcrafted feature space. LIME generates local, instance-level attributions by fitting weighted linear models to perturbed samples around each prediction, while SHAP provides globally consistent feature importance rankings grounded in Shapley value theory. Using a shared surrogate resolves what would otherwise be an attribution normalization problem across heterogeneous branches, since both explainers operate on an identical feature space. The approach proved capable of subtle reasoning: in one highlighted case, the model correctly flagged a phishing URL that contained no suspicious keywords, digits, or hyphens, with the explanation instead driven by the domain’s structural randomness and gibberish character composition, a signature of algorithmically generated phishing domains. Globally, the most influential features included the count of suspicious keywords, hyphen frequency, digit counts, dot counts, special-character ratios, and Shannon entropy, aligning closely with established expert heuristics.
The framework’s most novel component is its evidence-grounded language model reasoning module, instantiated with Google’s gemini-2.5-flash at a fixed temperature of 0.2. Crucially, the LLM never acts as a classifier. The neural network makes the phishing decision; the language model receives a structured JSON evidence packet containing the sanitized URL, the prediction label, the confidence score, and the top three LIME and SHAP feature attributions, and is instructed to synthesize an analyst-style report consisting of exactly four fields: an explanation, a categorical risk level, a recommended action, and an uncertainty flag. The system prompt explicitly prohibits the model from introducing outside knowledge, ensuring the narrative remains anchored to model-derived evidence rather than the LLM’s potentially outdated parametric memory.
Quantitative evaluation of this reasoning layer produced striking numbers. Across a benchmark of 100 instances spanning verified phishing, legitimate, borderline, and misclassified cases, the module achieved a mean semantic faithfulness score of 0.833, meaning the vast majority of top-ranked evidence features were explicitly referenced in the generated explanations. Repeated generation of the same evidence packet yielded 100 percent agreement on risk-severity labels, with an average pairwise text similarity of 0.597, indicating stable categorization alongside lexically varied but semantically coherent rationales. Every single response across all runs passed the strict JSON schema validation, recording zero formatting failures.
Security hardening extends throughout the inference pipeline, which the authors describe as a zero-trust-inspired interface at the LLM boundary rather than a claim of end-to-end system security. Inputs are treated as untrusted by default: regex-based sanitization strips null, empty, and malformed entries while preserving phishing-relevant structural elements, structured prompting constrains the model’s role, and programmatic validation checks every output against the required schema, retrying up to twice before surfacing an explicit error rather than forwarding unvalidated content. Red-team testing across twelve adversarial vectors, including direct instruction overrides, percent-encoded payloads, role-play jailbreaks, Base64 encoding, homoglyph substitution, and nested parameter spoofing, revealed an honest limitation: the regex sanitizer alone caught only 25 percent of injection attempts and none of the advanced obfuscations. Yet the downstream defenses held firm, with zero schema violations across all twelve cases, as manual inspection confirmed the model consistently synthesized feature attributions rather than executing embedded malicious instructions.
Detection performance was validated under demanding cross-dataset conditions. The model was trained on a primary corpus of 428,616 URLs and evaluated on an independent test set of 142,859 URLs pairing confirmed PhishTank 2026 phishing entries with a strictly disjoint pool of legitimate addresses never seen during training. On this shifted distribution, the framework achieved an accuracy of 0.9933, an F1-score of 0.9926, and a ROC-AUC of 0.9999, with a false-negative rate of roughly 0.17 percent and false positives bounded at 1.13 percent. The contrast with classical baselines was revealing: Random Forest, despite strong in-distribution accuracy, collapsed catastrophically to a recall of just 0.0663 on the unseen data, correctly identifying fewer than 7 percent of phishing URLs, while logistic regression performed respectably at an F1 of 0.9894. Multi-seed testing across five runs confirmed the hybrid model’s advantage was statistically significant, with a paired t-test yielding p equals 0.0022 and a McNemar’s test on individual predictions yielding p below 0.00001. An ablation study further showed that each component contributed: the handcrafted branch alone reached an F1 of 0.9658, the combined CNN-BiLSTM sequence encoder reached 0.9827, and full fusion pushed performance to its peak.
The authors are candid about limitations. The surrogate-based explanations cover the structured feature space but not the full recurrent architecture, the current feature set ignores dynamic signals such as WHOIS age, SSL certificates, and page content, and domain-level overlap of up to 48 percent across data splits suggests that domain-aware grouped splitting deserves future investigation. Sending raw enterprise URLs to cloud-hosted LLM endpoints also raises data residency concerns, prompting the authors to recommend air-gapped, self-hosted small language models such as quantized LLaMA-3-8B or Mistral-7B for production security operations centers. Looking ahead, they envision agentic security pipelines in which language models evolve from passive explainers into active cyber-assistants capable of orchestrating multi-step investigations and generating automated incident response playbooks. For now, the framework stands as one of the first empirically evaluated systems to tackle accuracy, interpretability, and LLM security simultaneously, offering a template for how artificial intelligence might finally earn the trust of the analysts who depend on it.
Subject of Research: Explainable deep learning and secure LLM-based reasoning for phishing URL detection
Article Title: A secure and explainable phishing detection framework using deep learning and LLM-based reasoning
Article References: A secure and explainable phishing detection framework using deep learning and LLM-based reasoning. (n.d.). https://doi.org/10.1007/s44564-026-00020-3
Image Credits: AI Generated
DOI: 10.1007/s44564-026-00020-3
Keywords: phishing detection, deep learning, large language models, explainable AI, LIME, SHAP, cybersecurity, prompt injection, CNN-BiLSTM, cross-dataset evaluation, security operations, machine learning
Cite Scienmag News
Blake Davidson. (October 1, 2026). AI Framework Catches Phishing Links and Explains Its Reasoning to Analysts. Scienmag. https://scienmag.com/ai-framework-catches-phishing-links-and-explains-its-reasoning-to-analysts/
Blake Davidson. "AI Framework Catches Phishing Links and Explains Its Reasoning to Analysts." Scienmag, 1 October 2026, https://scienmag.com/ai-framework-catches-phishing-links-and-explains-its-reasoning-to-analysts/. Accessed 1 October 2026.
Blake Davidson. "AI Framework Catches Phishing Links and Explains Its Reasoning to Analysts." Scienmag. October 1, 2026. https://scienmag.com/ai-framework-catches-phishing-links-and-explains-its-reasoning-to-analysts/

