Building a Small AI Model for XSS Detection

x32x01
  • by x32x01 ||
I’ve been experimenting with a question that kept coming up while working with AI models for programming and cybersecurity:
Why should I use a huge Mixture of Experts model when I only need it for a specific security task?
Models such as Qwen, GLM, and others can be extremely capable, but they also contain a lot of knowledge and capabilities that may not be necessary for a focused use case like Application Security or a specific vulnerability class.
That led me to a different idea: build a Small Language Model (SLM) specialized for one security problem.
In this case, I focused specifically on Cross-Site Scripting (XSS).



The Idea: A Small, Specialized Security Model​

Instead of trying to make a small model know everything, the idea is to take a larger model such as Claude or another capable model and use it to help build a much smaller model focused only on the knowledge I actually need.
After some research and experimentation, I ended up with a setup based on:
  • Qwen3-8B as the main model.
  • RAG to provide specialized security knowledge.
  • LoRA/QLoRA for model specialization.
  • Kev-4B as a decision model that evaluates whether new knowledge should actually become part of the learning process.
The goal is not to create another general-purpose AI assistant.
The goal is much narrower:
Can a very small model become highly effective at identifying one specific class of vulnerability?



The Problem With AI-Based Vulnerability Detection​

There is an important problem with using an AI model as the final authority.
An AI model can say:
“This looks like XSS.”
But what if it is wrong?
A model can potentially interpret a suspicious input or response as a vulnerability even when there is no real exploitability.
So I needed a way to separate:
  • What the AI thinks is happening.
  • What the browser can actually prove is happening.
That distinction became one of the most important parts of the project.



Live XSS Assessment System​

I built a Live XSS Assessment System around the model.
The system works by:
  1. Crawling the target.
  2. Discovering inputs and their contexts.
  3. Sending harmless markers.
  4. Performing context-aware probes.
  5. Running the assessment inside Headless Chromium.
  6. Collecting browser-based evidence.
The important part is that the AI model is not allowed to make the final decision by itself.



The Browser Has the Final Say​

I added a strict rule to the system:
No matter how confident the AI model is, it cannot mark a vulnerability as CONFIRMED.
Only the browser-based validation layer can do that.
For example, if the system believes it has found Reflected XSS, the model's confidence is not enough.
The system needs actual evidence that the payload resulted in real JavaScript execution.
That means the assessment needs to observe something that can prove whether the suspected vulnerability is actually exploitable.
This creates a clear separation:
LayerResponsibility
AI ModelAnalyze the target and identify potential vulnerabilities
Decision ModelEvaluate whether new knowledge or findings should be accepted
BrowserValidate actual behavior
EvidenceSupport the final vulnerability decision
This approach helps reduce the risk of treating an AI-generated assumption as a confirmed vulnerability.



Current Results​

So far, the system has reached approximately 93% success in my experiments.
Another interesting part is the model size.
The complete model is under 50 MB.
That is the direction I wanted to explore from the beginning:
A small, specialized model that focuses on one security problem instead of a massive general-purpose model.



What I Want to Try Next​

If the idea continues to produce strong results, the next step I have in mind is a Multi-Agent security system.
Instead of one large model trying to understand every vulnerability class, each agent could specialize in a specific vulnerability.
For example:
  • One agent for XSS.
  • Another agent for SQL Injection.
  • Another agent for SSRF.
  • Another agent for authentication issues.
  • Another agent for a different Application Security problem.
Each agent would have a narrow responsibility and would focus on determining whether a vulnerability actually exists.
The idea is also different from a traditional Generative AI or AI Chat system.
The agent does not need to have a long conversation or generate a huge amount of text.
Its primary job could simply be:
“Is there a real vulnerability here or not?”
The actual evidence and validation layer would then support that decision.



Open Source Project​

The project, including the model, is open source.
If you have ideas for improving the architecture, better ways to validate XSS, or experiments that could help test the approach, I’d be very interested in hearing them.
🔗 https://github.com/SecFathy/xss-specialist

The main idea is simple:
Instead of making AI bigger, what if we make it more specialized?
And instead of allowing the AI to decide that a vulnerability is real, what if we make the real-world execution evidence responsible for confirming it?
00.webp
 
Similar threads
x32x01
Replies
0
Views
117
x32x01
x32x01
x32x01
Replies
0
Views
57
x32x01
x32x01
x32x01
Replies
0
Views
133
x32x01
x32x01
x32x01
Replies
0
Views
109
x32x01
x32x01
x32x01
Replies
0
Views
6
x32x01
x32x01
Forum Statistics
Threads
1,097
Messages
1,103
Members
16
Latest Member
b_a_s_m_a_l_a7
Back
Top