vektor-guard

How it was built

The engineering decisions behind vektor-guard: architecture, data, training, and evaluation.

Build phases#

  1. Phase 1complete

    Data collection, cleaning, deduplication, and train/val/test splits.

  2. Phase 2complete

    Fine-tuned ModernBERT-large as a binary classification baseline, published as vektor-guard-v1.

    Training run
  3. Phase 3complete

    5-class classification plus the synthetic data pipeline, published as vektor-guard-v2.

    Training run
  4. Phase 4complete

    VektorGuard SDK and FastAPI guard service.

  5. Phase 5complete

    Re-ran the synthetic pipeline with v2 as its own validator.

  6. v3in progress

    Expanded training corpus and real-world evaluation.

Why ModernBERT-large#

The base model is answerdotai/ModernBERT-large, released in December 2024.

  • Trained on 2 trillion tokens of recent data, including code and technical content.
  • Uses rotary positional embeddings for better length generalization.
  • Has an 8,192-token native context window, 16× DeBERTa's 512. The long native context leaves room to fine-tune on longer inputs, where indirect injection hides in retrieved documents.
  • v2 is fine-tuned at a maximum of 2,048 tokens; longer inputs are truncated.
Candidate base models compared
ModelParamsContextTrain VRAMInference VRAM
ModernBERT-large (chosen)395M8,192~16GB (bf16)~3GB (fp16)
DeBERTa-v3-large400M512~16GB (bf16)~3GB (fp16)
DeBERTa-v3-base184M512~8GB (bf16)~1.5GB (fp16)
RoBERTa-large355M512~12GB~2.5GB
DistilBERT66M512~2GB~500MB

Hardware

  • Local RTX 4070 Super (12GB): development and inference validation.
  • A100 80GB on Colab: full training runs.
  • H100 80GB: hyperparameter sweeps.

Training data and the synthetic pipeline#

Public prompt-injection datasets are binary only, so per-category examples come from a synthetic pipeline.

Public datasets provide real-world volume and a strong clean-versus-injection signal; the provenance-controlled synthetic set is what teaches the four attack categories, since all three public sources are binary.

Training data sources
SourceLabelsTrain split
deepset/prompt-injections

Early public prompt-injection set: hand-collected injection prompts and benign controls.

Binary546
jackhhao/jailbreak-classification

Jailbreak prompts (DAN-style and role-play) alongside benign prompts.

Binary1,032
hendzh/PromptShield

Benchmark from the PromptShield paper (Jacob et al., 2025), curated from open-source datasets and published attack strategies.

Binary18,904
Synthetic

Generated 50/50 by GPT-4.1 and Claude Sonnet and validated in two layers; the only source with 5-class labels.

5-class1,514
Total21,996

Phase 5 later expanded the synthetic set to a balanced 2,274 examples, which feeds v3 (see the validator flywheel below).

Generation is split 50/50 between GPT-4.1 and Claude Sonnet to reduce single-model bias. Each generated example then passes through two validation layers.

  • Layer 1, confidence gate. Every generated example runs through the current vektor-guard model. Attack examples must score as an attack above a 0.85 confidence threshold; anything below is flagged.
  • Layer 2, category verification. A second LLM independently classifies each example that passed Layer 1. Disagreement with the intended label flags the example.

Flagged examples go to per-category review files with confidence scores attached.

From seven classes to five#

The original plan had 7 categories. Empirical validation collapsed it to 5.

  • direct_injection and instruction_override described the same attack from different angles. Forcing a split would teach the model noise, not signal.
  • stored_injection is indirect_injection with persistence: the same mechanism with different delivery timing.

See the final taxonomy

Fixing class imbalance#

The problem

About 16,400 Phase 2 examples were binary and mapped only to clean and instruction_override; 1,514 synthetic examples were spread across five classes. After epoch 1, per-class F1 for indirect_injection, jailbreak, and tool_call_hijacking was 0.00%. The model had collapsed to two classes.

The fix

A WeightedTrainer subclass overrides get_train_dataloader() to use WeightedRandomSampler with inverse-frequency class weights. Rare classes are sampled proportionally more often, without discarding any data.

Validation fix

The Phase 2 validation set had only binary labels, so minority classes scored zero even while learning. 15% of the synthetic data was carved into validation so all five classes are represented.

The validator flywheel#

Phase 5 re-ran the synthetic pipeline using v2 as the Layer 1 validator instead of v1.

v1 as validator

7.5%

tool_call_hijacking Layer 1 pass rate

v2 as validator

94%

tool_call_hijacking Layer 1 pass rate

The synthetic set grew from 1,514 to a balanced 2,274.

Synthetic examples per category
CategoryExamples
instruction_override465
indirect_injection448
jailbreak440
tool_call_hijacking469
clean452

The idea: each model generation becomes a better data gate for the next, which trains a better model.

How it's evaluated#

The primary metrics are macro F1, which weights every class equally so rare attack types count, and false negative rate: attacks labeled clean, the costly error for a guard.

v1: public data only

v1 was a binary classifier trained and evaluated entirely on the public datasets above: no synthetic data.

v1 results on its held-out test split
Metricv1 result
Held-out test examples2,049
Accuracy99.8%
F199.8%
Recall99.71%
False negative rate0.29%

v1 training run on Weights & Biases

v2: five classes

Targets compared with v2 results on the held-out test split
MetricTargetv2 result
Macro F1≥ 90%99.81%
False negative rate≤ 5%0.47%
Per-class F1≥ 80% eachLowest 99.51% (instruction_override)

v2 training run on Weights & Biases

Several public injection datasets recycle the same upstream sources, so a "held-out" real-world set can silently overlap training data. The v3 evaluation set is being built from provenance-checked sources outside the training data's lineage.

See the benchmarks on the home page.

What's next#

v3 adds real-world training data alongside synthetic data, and evaluates on a real-world, per-category holdout.

v2 is the recommended model until v3 ships: vektor-guard-v2 on HuggingFace.