Methodology

Automation without artificial certainty.

Repeatability is valuable only when the product also exposes confidence, coverage, limitations, and the boundary of machine judgment.

01

Authorization before testing

Ownership, boundary, exclusions, and stop conditions come before any test run.

02

Version everything

Test suite, configuration, grading method, scoring rule, and report version remain traceable.

03

Inconclusive means inconclusive

Low-confidence machine judgments do not become confirmed vulnerabilities.

04

Minimum necessary data

Use a test environment, synthetic data, and only the access needed for the approved assessment wherever possible.

05

Coverage made visible

Reports distinguish tested, passed, failed, inconclusive, not applicable, and not tested.

06

Expert escalation is explicit

Manual judgment is optional, separately scoped, and attributed to the specialist providing it.

Layer 1

Controlled automated evaluation

Each assessment selects the behaviors and customer-defined policies relevant to the approved scope. Evil AI runs non-destructive checks, captures the evidence needed to understand each result, and prepares supported findings for reporting.

Testing begins only after access and scope are confirmed. Submitting an inquiry does not authorize or initiate an assessment.

Layer 2

Quality and uncertainty controls

Findings retain their test-suite version, grading method, evidence, confidence, and limitations. Duplicate manifestations should not inflate risk, and historical reports do not silently change with new scoring rules.

Layer 3

Optional expert assistance

Complex architectures, consequential findings, disputed results, and remediation decisions may warrant specialist review. That work is separately scoped and subject to finding a qualified independent specialist.

Alignment

Informed by recognized guidance.

Alignment does not imply endorsement, accreditation, compliance, or certification.

  • OWASP Top 10 for LLM Applications 2026 and Top 10 for Agentic Applications 2026
  • NIST AI RMF 1.0 (AI 100-1, 2023) and Generative AI Profile (AI 600-1, July 2024)
  • ISO/IEC 42001:2023 management-system evidence support and ISO/IEC 42005:2025 impact-assessment input

View dated framework versions, check-by-check coverage, and exclusions →

Private beta · authorized applications only

Find the failure before your users do.

Request beta access