Platform/Diamond
VIJIL DIAMOND · TEST ENGINE

Test your agents before you trust your agents

Diamond verifies the trustworthiness of an agent in minutes — against adaptive attacks and your own policies.

Start for free Learn more Download datasheet →
$ vijil evaluate my-agent --baseline
Diamond — create evaluation console
VIJIL TRUST REPORT my-agent · run 14
87 / 100
Reliability92
Security84
Safety88
3 findings · 2 fixes proposed ILLUSTRATIVE
SPEC DISCOVER DIAMOND DEPLOY DOME DARWIN
Where Diamond fits
THE EVIDENCE PROBLEM

A demo that worked is not evidence

Every test you ran was run on your terms, on inputs you thought of. The person who has to sign off answers to a hostile reader, and needs to know which attacks were tried, which ones landed, and what the failure actually looked like. A single leaderboard number answers neither question.

One number, many failures

Reliable-but-leaky and safe-but-brittle can score the same. An aggregate hides the dimension that will hurt you.

Not reproducible

A benchmark you cannot re-run against your agent, your tools and your data is a claim about somebody else’s system.

Tested on your terms

You tested the inputs you thought of. The adversary’s entire job is the ones you did not.

HOW IT WORKS

Any agent in, evidence out

YOUR AGENT SYSTEM UNDER TEST LLM · LLM app custom agent · multi-agent system TEST SUITE built-in + bespoke policies metrics probes detectors Diamond test engine 87 TRUST SCORE adaptive · multi-turn · stateful TRUST REPORT YOUR AGENT SYSTEM UNDER TEST LLM · LLM app custom agent · multi-agent system TEST SUITE built-in + bespoke policies metrics probes detectors Diamond test engine 87 TRUST SCORE TRUST REPORT adaptive · multi-turn · stateful
Every finding in the report drills down to the probe that produced it — evidence, not a verdict you have to take on faith.
Read the Diamond docs →
METHODOLOGY

How the score is computed

1
Probes
Attacks and tasks from OWASP LLM Top 10, MITRE ATLAS, garak, and internal red-team.
2
Detectors
Graders label responses pass / fail — calibrated against human-labeled ground truth.
3
Metrics
Pass rates roll up into 12 dimensions, each scored 0–100 with confidence intervals.
4
Score
Weighted by your policy into one Trust Score — traceable to every probe.
REPRODUCIBLEversioned probes + seeds· TRANSPARENTevery score drills to its evidence· CALIBRATEDbenchmarked vs human labels· CURRENTtracks new attacks
Methodology in depth — download the datasheet →·Adaptive offensive AI security testing (PDF) →
COMPLIANCE

Your policies become the test suite

Government regulations EU AI Act · GDPR · NYC LL 144 Industry standards NIST AI RMF · ISO 42001 · OWASP Organization policies code of conduct · privacy · terms of use Agent-specific policies ✓ permitted ✗ prohibited Diamond compiles them into bespoke tests COMPLIANCE REPORT Government regulations EU AI Act · GDPR · NYC LL 144 Industry standards NIST AI RMF · ISO 42001 · OWASP Organization policies code of conduct · privacy · terms of use Agent-specific policies ✓ permitted ✗ prohibited Diamond compiles them into bespoke tests COMPLIANCE REPORT
Dome enforces the same compiled policies at runtime — what Diamond certifies, Dome defends.
Compliance mapping for your auditors →
INSTALL & RUN

Point it at an agent, get evidence back

Diamond takes any agent behind an endpoint — no code change — runs the harness your policy implies, and returns the failures with the transcript attached. Run it by hand once, then wire the same command into CI so a regression fails the build instead of reaching production.

verify
$ pip install vijil-sdk# CLI + SDK
$ vijil verify my-agent --spec spec.yaml# run the harness
$ vijil evaluate my-agent# gate the pipeline
$ vijil scores history my-agent# track drift over time
ANY AGENT
an endpoint is enough — no code change
POLICY-DRIVEN
your spec becomes the test suite
REPRODUCIBLE
every score carries the run id that produced it
PROOF
SmartRecruiters
Our enterprise customers demand trust verification before deploying AI in hiring workflows. Vijil helps us ship AI agents in six weeks instead of six months while dramatically lowering compliance costs.
Michal Nowak
Senior Vice President, Engineering · SmartRecruiters
6 wks
time-to-trust, down from 6 months
lower cost of compliance — EU AI Act, GDPR, NYC LL 144
Read the SmartRecruiters story →
WHERE IT FITS
spec YOUR INPUT discover DISCOVER verify DIAMOND deploy YOUR CI/CD defend DOME evolve DARWIN what production teaches amends the spec Dashed steps are yours, not ours
Diamond is the one test engine in the Trusted Agent Lifecycle, run for three purposes: the developer evaluates, the risk owner verifies, the business owner validates. Dome enforces what Diamond certifies; Darwin evolves what Diamond scores.

Score your agent today

Free tier, no sales call. Point Diamond at your agent and get a complete trust report in minutes.

Start for free Download datasheet →