How it was built
The engineering decisions behind vektor-guard: architecture, data, training, and evaluation.
Build phases#
- Phase 1complete
Data collection, cleaning, deduplication, and train/val/test splits.
- Phase 2complete
Fine-tuned
Training runModernBERT-largeas a binary classification baseline, published asvektor-guard-v1. - Phase 3complete
5-class classification plus the synthetic data pipeline, published as
Training runvektor-guard-v2. - Phase 4complete
VektorGuard SDK and FastAPI guard service.
- Phase 5complete
Re-ran the synthetic pipeline with v2 as its own validator.
- v3in progress
Expanded training corpus and real-world evaluation.
Why ModernBERT-large#
The base model is answerdotai/ModernBERT-large, released in December 2024.
- Trained on 2 trillion tokens of recent data, including code and technical content.
- Uses rotary positional embeddings for better length generalization.
- Has an 8,192-token native context window, 16× DeBERTa's 512. The long native context leaves room to fine-tune on longer inputs, where indirect injection hides in retrieved documents.
- v2 is fine-tuned at a maximum of 2,048 tokens; longer inputs are truncated.
| Model | Params | Context | Train VRAM | Inference VRAM |
|---|---|---|---|---|
| ModernBERT-large (chosen) | 395M | 8,192 | ~16GB (bf16) | ~3GB (fp16) |
| DeBERTa-v3-large | 400M | 512 | ~16GB (bf16) | ~3GB (fp16) |
| DeBERTa-v3-base | 184M | 512 | ~8GB (bf16) | ~1.5GB (fp16) |
| RoBERTa-large | 355M | 512 | ~12GB | ~2.5GB |
| DistilBERT | 66M | 512 | ~2GB | ~500MB |
Hardware
- Local RTX 4070 Super (12GB): development and inference validation.
- A100 80GB on Colab: full training runs.
- H100 80GB: hyperparameter sweeps.
Training data and the synthetic pipeline#
Public prompt-injection datasets are binary only, so per-category examples come from a synthetic pipeline.
Public datasets provide real-world volume and a strong clean-versus-injection signal; the provenance-controlled synthetic set is what teaches the four attack categories, since all three public sources are binary.
| Source | Labels | Train split |
|---|---|---|
deepset/prompt-injections Early public prompt-injection set: hand-collected injection prompts and benign controls. | Binary | 546 |
jackhhao/jailbreak-classification Jailbreak prompts (DAN-style and role-play) alongside benign prompts. | Binary | 1,032 |
hendzh/PromptShield Benchmark from the PromptShield paper (Jacob et al., 2025), curated from open-source datasets and published attack strategies. | Binary | 18,904 |
Synthetic Generated 50/50 by GPT-4.1 and Claude Sonnet and validated in two layers; the only source with 5-class labels. | 5-class | 1,514 |
| Total | 21,996 |
Phase 5 later expanded the synthetic set to a balanced 2,274 examples, which feeds v3 (see the validator flywheel below).
Generation is split 50/50 between GPT-4.1 and Claude Sonnet to reduce single-model bias. Each generated example then passes through two validation layers.
- Layer 1, confidence gate. Every generated example runs through the current vektor-guard model. Attack examples must score as an attack above a 0.85 confidence threshold; anything below is flagged.
- Layer 2, category verification. A second LLM independently classifies each example that passed Layer 1. Disagreement with the intended label flags the example.
Flagged examples go to per-category review files with confidence scores attached.
From seven classes to five#
The original plan had 7 categories. Empirical validation collapsed it to 5.
direct_injectionandinstruction_overridedescribed the same attack from different angles. Forcing a split would teach the model noise, not signal.stored_injectionisindirect_injectionwith persistence: the same mechanism with different delivery timing.
Fixing class imbalance#
The problem
About 16,400 Phase 2 examples were binary and mapped only to clean and instruction_override; 1,514 synthetic examples were spread across five classes. After epoch 1, per-class F1 for indirect_injection, jailbreak, and tool_call_hijacking was 0.00%. The model had collapsed to two classes.
The fix
A WeightedTrainer subclass overrides get_train_dataloader() to use WeightedRandomSampler with inverse-frequency class weights. Rare classes are sampled proportionally more often, without discarding any data.
Validation fix
The Phase 2 validation set had only binary labels, so minority classes scored zero even while learning. 15% of the synthetic data was carved into validation so all five classes are represented.
The validator flywheel#
Phase 5 re-ran the synthetic pipeline using v2 as the Layer 1 validator instead of v1.
v1 as validator
7.5%
tool_call_hijacking Layer 1 pass rate
v2 as validator
94%
tool_call_hijacking Layer 1 pass rate
The synthetic set grew from 1,514 to a balanced 2,274.
| Category | Examples |
|---|---|
| instruction_override | 465 |
| indirect_injection | 448 |
| jailbreak | 440 |
| tool_call_hijacking | 469 |
| clean | 452 |
The idea: each model generation becomes a better data gate for the next, which trains a better model.
How it's evaluated#
The primary metrics are macro F1, which weights every class equally so rare attack types count, and false negative rate: attacks labeled clean, the costly error for a guard.
v1: public data only
v1 was a binary classifier trained and evaluated entirely on the public datasets above: no synthetic data.
| Metric | v1 result |
|---|---|
| Held-out test examples | 2,049 |
| Accuracy | 99.8% |
| F1 | 99.8% |
| Recall | 99.71% |
| False negative rate | 0.29% |
v1 training run on Weights & Biases
v2: five classes
| Metric | Target | v2 result |
|---|---|---|
| Macro F1 | ≥ 90% | 99.81% |
| False negative rate | ≤ 5% | 0.47% |
| Per-class F1 | ≥ 80% each | Lowest 99.51% (instruction_override) |
v2 training run on Weights & Biases
Several public injection datasets recycle the same upstream sources, so a "held-out" real-world set can silently overlap training data. The v3 evaluation set is being built from provenance-checked sources outside the training data's lineage.
See the benchmarks on the home page.
What's next#
v3 adds real-world training data alongside synthetic data, and evaluates on a real-world, per-category holdout.
v2 is the recommended model until v3 ships: vektor-guard-v2 on HuggingFace.