Safety Index
Preview benchmark snapshot for AI safety systems
Preview snapshot: 2026-06-26 | Version 1.0.0
Illustrative preview only. Benchmark materials available on request.
Latency
Category Breakdown
Jailbreak Prevention
PII Detection
Content Moderation
Injection Defense
Constitutional Compliance
Competitor Comparison
Benchmark data coming soon. We are committed to transparent, verifiable comparisons against leading AI safety solutions.
Request early access to benchmark materialsMethodology
Test Datasets
- Jailbreak: adversarial prompts from public corpora, internal red team, and generated variations
- PII: synthetic cases using Faker across common PII types
- Content Moderation: cases covering hate, violence, and misinformation
- Injection: prompt injection patterns including encoding attacks
- Constitutional: rule compliance tests per active creed
Evaluation Criteria
- Accuracy: (TP + TN) / Total tests
- False Positive Rate: FP / (FP + TN), legitimate content blocked
- False Negative Rate: FN / (FN + TP), harmful content allowed
- Latency: p50/p95/p99 processing time in milliseconds
Reproducibility
Benchmark code and datasets are available for procurement and audit review. Full methodology documentation is available on request. Run the benchmark in a configured workspace:
make safety-index Quarterly automated runs ensure consistent tracking. In the production benchmark pipeline, results are cryptographically signed. The numbers shown here are a static preview snapshot.