Skip to content
#small language model Open access

A Staged Anti-Phishing Browser Extension for One-Day Scam Pages

Oct 2026 · Zenodo (CERN European Organization for Nuclear Research)
Spam and Phishing Detection

Abstract

Legion is an extension for Chrome, Edge, and Brave. It looks at the page a person already has open and warns them when that page asks for a password, a card, an identity document, or a one-time code and behaves like a one-day phishing page. Such pages collect data in the first hours, while their addresses are still absent from shared blocklists. The warning appears after the page has been examined. A login form by itself is not enough. A strong, corroborated risk covers the screen: the person can leave or continue anyway. A weaker signal shows a yellow caution and leaves the page open. An ordinary site with no ground for concern passes quietly. A small analysis service runs beside the extension. In the present version of the project it runs on the same computer, at 127.0.0.1, port 8787. While Legion remains a student and research prototype, that local service is the intended mode. Full hosting belongs to the later moment when the product is published for general use. Moving the service does not change the order of checks: the settings store a different address. The problem Lists of dangerous addresses fire after someone has already noticed a page and written it down. Until that entry exists, the page can look like a bank, a government portal, a shop, or a game service and can ask for secrets. In 2025 the Anti-Phishing Working Group recorded 3.8 million phishing attacks. In the same year the FBI Internet Crime Complaint Center received 191,561 phishing and spoofing complaints, the most common complaint type of the year, with 215.8 million dollars in reported losses. In Verizon’s 2025 breach data, phishing remains the path of initial access in about 15 percent of cases. Legion occupies the interval between the opening of such a page and the appearance of its address on a shared list. The product is built for an ordinary browser user and stands beside antivirus software and beside the browser’s own check. How a page is checked The decision is staged. Each stage either returns an answer at once or passes the page on. Network calls run after the local lists and after the memory of earlier answers, so a familiar site and a repeat view of the same page do not wait for the service. 1. A trusted address is passed at once. Trust covers the site and its subdomains. Three hosts that publish other people’s content — wordpress.com, accounts.wordpress.com, and livejournal.com — are trusted only as exact names. A foreign subdomain on them receives no trust. 2. An address on the local threat list, or a hit from Google Safe Browsing, produces a full-screen red warning. What leaves the machine is the scheme, the host, and the path. The query string is removed first, because it can carry access tokens and personal identifiers. 3. If the same page, with the same content, has already been scored, the stored answer is reused. 4. A page with no fields for personal or secret data is treated as calm. The fields themselves matter: a password, a card, a document, a one-time code. The word “password” in help text does not count as such a field. 5. Whatever remains is scored from the text, from a known brand mentioned on a foreign address, and from a look-alike host name. Likeness uses Damerau–Levenshtein distance and a brand token standing as its own word inside the host. 6. A red screen for a brand requires a second signal: a look-alike host, a form posted to another site, plain HTTP, or text markers that are already strong enough. One mention of a brand is not enough. A high score, and a form that sends a password or a card to another address, also produce a red screen. 7. A yellow caution has two cases. Tier A: an unknown site asks for a password, a card, a document, or a code. Tier B: the site asks for an email address, a phone number, or a name, and one further clue is present — a look-alike host, a brand on a foreign domain, a form pointed at another address, or plain HTTP. 8. With no supporting signal, the page stays free of a warning. A language model is optional and is called only inside a narrow band of uncertain scores. Without a key, the decision comes from the local rules and, when a key is configured, from Safe Browsing. What the user sees Badge Meaning Gray ellipsis A check is in progress Green check A trusted site, or no ground for concern Yellow question mark Caution. The screen stays open, and the reasons are shown Red exclamation mark Corroborated risk. A full-screen warning Gray dot The person continued anyway, or the page could not be read The red screen offers two actions: leave the page, or continue. If the person continues, the gray dot stays on the icon so that choice remains visible afterwards. What the prototype contains The client is built on Manifest V3. A page script reads ordinary http and https sites, watches for forms as they appear, including forms inside dialogs and shadow trees, and sends the background script a short description: the address, the field types, whether the form posts elsewhere, the title, and small text snippets. The background script runs the protocol and sets the badge. A settings page can add a personal trusted address and can set the service URL. The service is written in Python with FastAPI and Uvicorn. It performs the Safe Browsing check, repeats the rules, may call a language model, and keeps a small local record of verdicts. That record does not change the decision. Keys live in a local environment file and are not packaged in the extension. Requests to the service are accepted only from the extension itself. Version 0.5.4 ships with: a local whitelist of 899 trusted names and 3 hosts trusted only as exact names; a short teaching blacklist of four names, used to demonstrate that stage; a list of twelve brands often impersonated for a Russian-speaking user: Gosuslugi, Sber, T-Bank, VTB, Alfa-Bank, VK, Yandex, Steam, Discord, Mos.ru, Wildberries, and Ozon; demonstration pages for the red, yellow, and calm cases; a set of sixteen checks that record the verdict the protocol is expected to return, including earlier false warnings the program is required not to repeat. With the service switched off, the extension continues on the whitelist, the teaching blacklist, and the local rules. What has been shown On the demonstration phishing pages assembled while the project was being built, 97 percent were identified. Identified means the page received a red warning or a yellow caution. The other 3 percent were not flagged. This is the author’s result on pages prepared for the project. It shows that the rules recognize the lures they were written for. It is a different quantity from the public complaint statistics: those figures count people who reported an incident, while 97 percent counts pages from the project itself. The sixteen automatic cases check the protocol with the service switched off. Five pages in the demo directory show the same rules inside a browser. A pressure lure that asks for a Gosuslugi password, and a lure that appears only inside a dialog, are expected to produce a red screen. An unknown login is expected to show the yellow badge. A page that asks only for an email address, with no further clue, is expected to stay green. How the project runs today The service listens only on the local address 127.0.0.1:8787. For the project as it stands, for the contest, and for an open deposit of the materials, that is the arrangement: one computer, one browser, one local process. Without the service, the extension still applies the lists and the rules that live in the client. When the product is prepared for general use, the same service can be placed on ordinary hosting. The extension already reads the service address from its settings, so the experience for the person and the order of the stages stay as they are. Hosting belongs to that later publication. It is outside the prototype this text describes. Boundaries Legion is an additional layer. Antivirus software, enterprise protection, and the browser’s built-in check stay in place. The extension reads ordinary web pages. The browser platform does not hand it internal browser pages or the extension store. The project needs no user account, and it sends no click history anywhere. The teaching blacklist of four names demonstrates one stage and is not, by itself, a threat intelligence feed: known threats are left to Safe Browsing when a key is configured. The twelve brands were chosen for impersonations a Russian-speaking user is likely to meet. The language-model path is present in the code and stays quiet until a key is supplied. In one sentence Legion is a browser extension that, on one-day pages, separates an ordinary login from a corroborated phishing risk, warns only after that examination, and in version 0.5.4 keeps its analysis service on the same computer as the browser.

View source

Similar papers

#small language model Dataset Open access Oct 2026

Socratic guiding questions in synthetic arithmetic data: matched LoRA runs (revision v2)

Supporting data, adapters, predictions and code for the article *Low-Cost LoRA Fine-Tuning of Small Language Models for Multi-Step Arithmetic Reasoning* by Jake O'Grady, Asena Isik Gürhan, Chee Fong Ting and Effirul Ramlan (University of Galway). We generated 20,000 GSM8K-derived arithmetic problems with step-by-step s...

O'Grady, Jake, Gürhan, Asena Isik, Chee, Fong Ting et al. · 465 citations
#computer vision Open access Jun 2016

Software Development in Startup Companies: The Greenfield Startup Model

The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.

Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al. · 178 citations · ⚡14
#computer vision Open access Oct 2016

Software Startups - A Research Agenda

Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.

M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al. · 157 citations · ⚡17
#machine learning Review Open access Oct 2016

“Failures” to be celebrated: an analysis of major pivots of software startups

This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.

Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al. · 127 citations · ⚡15
#computer vision Review Open access May 2015

A survey study on major technical barriers affecting the decision to adopt cloud services

The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.

Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al. · 111 citations · ⚡8
#computer vision Open access Feb 2018

Lean Internal Startups for Software Product Innovation in Large Companies: Enablers and Inhibitors

This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.

Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al. · 78 citations · ⚡6

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.