E-PhishGEN : Unlocking Novel Research in Phishing Email Detection

dc.contributor.authorPajola, Luca
dc.contributor.authorCaripoti, Eugenio
dc.contributor.authorBanzer, Stefan
dc.contributor.authorPizzi, Simeone
dc.contributor.authorConti, Mauro
dc.contributor.authorApruzzese, Giovanni
dc.contributor.departmentDepartment of Computer Science
dc.date.accessioned2026-09-24T13:42:01Z
dc.date.available2026-09-24T13:42:01Z
dc.date.issued2025-12-30
dc.descriptionPublisher Copyright: © 2025 Copyright is held by the owner/author(s). Publication rights licensed to ACM.en
dc.description.abstractEvery day, our inboxes are flooded with unsolicited emails, ranging between annoying spam to more subtle phishing scams. Unfortunately, despite abundant prior efforts proposing solutions achieving near-perfect accuracy, the reality is that countering malicious emails still remains an unsolved dilemma. This "open problem"paper carries out a critical assessment of scientific works in the context of phishing email detection. First, we focus on the benchmark datasets that have been used to assess the methods proposed in research. We find that most prior work relied on datasets containing emails that - we argue - are not representative of current trends, and mostly encompass the English language. Based on this finding, we then re-implement and re-assess a variety of detection methods reliant on machine learning (ML), including large-language models (LLM), and release all of our codebase - an (unfortunately) uncommon practice in related research. We show that most such methods achieve near-perfect performance when trained and tested on the same dataset - a result which intrinsically hinders development (how can future research outperform methods that are already near perfect?). To foster the creation of "more challenging benchmarks"that reflect current phishing trends, we propose E-PhishGEN, an LLM-based (and privacy-savvy) framework to generate novel phishing-email datasets. We use our E-PhishGEN to create E-PhishLLM, a novel phishing-email detection dataset containing 16616 emails in three languages. We use E-PhishLLM to test the detectors we considered, showing a much lower performance than that achieved on existing benchmarks - indicating a larger room for improvement. We also validate the quality of E-PhishLLM with a user study (n=30). To sum up, we show that phishing email detection is still an open problem - and provide the means to tackle such a problem by future research.en
dc.description.versionPeer revieweden
dc.format.extent13
dc.format.extent1087985
dc.format.extent64-76
dc.format.extent
dc.identifier.citationPajola, L, Caripoti, E, Banzer, S, Pizzi, S, Conti, M & Apruzzese, G 2025, E-PhishGEN : Unlocking Novel Research in Phishing Email Detection. in Proceedings of the 18th ACM Workshop on Artificial Intelligence and Security, AISec 2025. Proceedings of the 18th ACM Workshop on Artificial Intelligence and Security, AISec 2025, Association for Computing Machinery, Inc, pp. 64-76, 18th ACM Workshop on Artificial Intelligence and Security, AISec 2025, Taipei, Taiwan, Province of China, 13/10/25. https://doi.org/10.1145/3733799.3762967en
dc.identifier.citationconferenceen
dc.identifier.doi10.1145/3733799.3762967
dc.identifier.isbn9798400718953
dc.identifier.other250864006
dc.identifier.othere81127bf-e808-4ef0-ae6b-b98f0a8ce1ff
dc.identifier.other105027095261
dc.identifier.urihttps://hdl.handle.net/20.500.11815/8348
dc.language.isoen
dc.publisherAssociation for Computing Machinery, Inc
dc.relation.ispartofseriesProceedings of the 18th ACM Workshop on Artificial Intelligence and Security, AISec 2025; ()en
dc.relation.ispartofseriesProceedings of the 18th ACM Workshop on Artificial Intelligence and Security, AISec 2025; ()en
dc.relation.urlhttps://www.scopus.com/pages/publications/105027095261en
dc.rightsinfo:eu-repo/semantics/openAccessen
dc.subjectbenchmarken
dc.subjectdataseten
dc.subjectdetectionen
dc.subjectemailen
dc.subjectlarge language modelsen
dc.subjectspamen
dc.subjectArtificial Intelligenceen
dc.subjectComputer Networks and Communicationsen
dc.subjectSoftwareen
dc.titleE-PhishGEN : Unlocking Novel Research in Phishing Email Detectionen
dc.type/dk/atira/pure/researchoutput/researchoutputtypes/contributiontobookanthology/conferenceen

Skrár

Original bundle

Niðurstöður 1 - 1 af 1
Nafn:
3733799.3762967.pdf
Stærð:
1.04 MB
Snið:
Adobe Portable Document Format