RESEARCH
The science of evolving trust in multi-agent systems, in the open
Vijil’s products are built on research we publish and tools we release. Everything the platform measures or enforces traces back to a method you can read and a model you can download.
TOPICS
What we work on
Adaptive red-teaming
Probes and attack strategies that surface hallucination, injection, jailbreak, and leakage failures — before adversaries do.
Trust measurement
Metrics and methodology for scoring reliability, security, and safety — with detectors calibrated against human-labeled ground truth.
Governance to guardrails
Small, fast open models for runtime defense — detecting prompt injections at inference time without adding latency.
Agent adaptation
Methods for attributing production failures to root causes and evolving agents under explicit constraints.
PUBLICATIONS & ARTIFACTS
Published work
PAPERS
VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
Fang, Yuan, Kong, Rudner (Vijil) · arXiv:2607.19575 · 2026
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
Lamb, Ivanova, Torr, Rudner (Vijil) · arXiv:2604.07172 · 2026
Open Problems in Frontier AI Risk Management
Ziosi, Plueckebaum, Casper, et al., incl. Rudner (Vijil) · Oxford Martin AI Governance Initiative · 2026
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
Bhardwaj, Kempe, Rudner (Vijil) · arXiv:2510.21891 · 2025
Red Teaming AI Red Teaming
Majumdar (Vijil), Pendleton, Gupta · arXiv:2507.05538 · 2025
Localized LoRA: A Structured Low-Rank Approximation for Efficient Fine-Tuning
Barazandeh, Majumdar (Vijil), Rajyaguru, Michailidis · arXiv:2506.00236 · 2025
Consistency in Language Models: Current Landscape, Challenges, and Future Directions
Novikova, Anderson, Blili-Hamelin, Majumdar (Vijil) · arXiv:2505.00268 · 2025
Improving Consistency in Large Language Models through Chain of Guidance
Raj, Gupta, Rosati, Majumdar (Vijil) · arXiv:2502.15924 · 2025
Stop Treating ‘AGI’ as the North-Star Goal of AI Research
Blili-Hamelin, Graziul, Hancox-Li (Vijil), et al. · ICML 2025 · arXiv:2502.03689
Embedding-Based Classifiers Can Detect Prompt Injection Attacks
Ayub, Majumdar · arXiv:2410.22284 · 2024
Is ETHICS About Ethics? Evaluating the ETHICS Benchmark
Hancox-Li (Vijil), Blili-Hamelin · arXiv:2410.13009 · 2024
Evaluating Defences Against Unsafe Feedback in RLHF
Rosati, Edkins, Raj, Atanasov, Majumdar (Vijil), et al. · arXiv:2409.12914 · 2024
garak: A Framework for Security Probing Large Language Models
Derczynski, Galinkin, Martin, Majumdar (Vijil), Inie · arXiv:2406.11036 · 2024
Representation Noising: A Defence Mechanism Against Harmful Finetuning
Rosati, Wehner, Williams, et al., incl. Majumdar (Vijil) · NeurIPS 2024 · arXiv:2405.14577
Unsocial Intelligence: An Investigation of the Assumptions of AGI Discourse
Blili-Hamelin, Hancox-Li (Vijil), Smart · arXiv:2401.13142 · 2024
Showing all 15 papers · sorted by recency
OPEN MODELS
Prompt-injection detection models
DeBERTa- and ModernBERT-based classifiers · Hugging Face
Lightweight open-source detectors that Dome uses for runtime defense — released for anyone to use, benchmark, or fine-tune.
Download from Hugging Face →
METHODOLOGY
The Vijil Trust Score
200,000+ probes · detectors · dimensions · one score
How probe results roll up through calibrated detectors into dimension scores for reliability, security, and safety — weighted by your policy into a single auditable number.
Read the docs →
More research notes on the Vijil blog.
Collaborate with us
We work with academic and industry partners on agent evaluation, red-teaming, and runtime defense.