AI Governance · The Evidence Layer
Every AI system you deploy acts in your name — your customers, your reputation, and your judgment stand behind it. Our platform red-teams it across 8 dimensions of trust and turns the results into evidence: proof, for your board and for yourself, that your AI behaves the way you’d stand behind it.
Every AI system you deploy makes decisions in your name — at a scale no human review can keep up with. It answers your customers, moves your money, and speaks with your authority.
Most organizations deploy first and hope it behaves. But hope is not a governance posture, and a system you haven’t measured is a system you can’t stand behind.
The gap isn’t intent — you already mean to use AI well. The gap is evidence: something concrete you can point to that shows how your AI actually behaves.
73%
Of enterprises deploying AI have no formal governance framework
$2.3M
Average cost when an AI failure reaches production
3 wks → sec
Time to surface a failure: manual review vs. a red-team run
The Methodology
Our 8-dimension framework was developed by PhD researchers and validated through peer-reviewed publications.
Because the AI systems we test are non-deterministic, we report registered, hash-citable attack traces and statistical evidence: the rate at which each failure occurs across N runs, with Wilson confidence intervals and Rogan-Gladen correction for judge reliability.
When a budget limits coverage, we report exactly what was measured and what wasn’t. Every score traces back to the attack that produced it.
Explore the framework →Research published in AI bias, compliance frameworks, and ethical evaluation methodology
In-house researchers with doctoral expertise in AI/ML and ethics
Ongoing partnership with universities and research centers for methodology validation
Every scoring criterion is documented, versioned, and publicly auditable
The Framework
Each dimension scores your AI system’s behavior, is independently traceable, and rolls up through the nine control areas of ISO 42001 Annex A to the frameworks you report against. The eight dimensions measure behavior; regulatory compliance is the certification that follows once the measurements hold.
01
Detects bias, stereotyping, and unequal treatment across protected groups, using statistical and counterfactual testing
→ Art. 10
02
Surfaces hate, violence, self-harm, sexual content, and dangerous instructions — including harm that carries no obviously toxic wording
→ General
03
Measures unfaithful explanations, sycophancy, and undisclosed AI — whether a decision can be understood by the people it affects
→ Art. 13
04
Probes PII leakage, training-data extraction, and system-prompt disclosure across your data-governance boundaries
→ GDPR
05
Verifies outputs against source material to catch hallucination, fabricated citations, and unsupported claims
→ Art. 15
06
Tests resistance to prompt injection, jailbreaks, encoding tricks, and adversarial suffixes that bend the model's behavior
→ Art. 15
07
Exercises unauthorized actions, privilege abuse, identity spoofing, and tool misuse wherever your AI can act, not just answer
→ Agentic AI
08
Checks oversight saturation, governance evasion, and whether every action stays traceable to the human accountable for it
→ Art. 14
Every finding is traceable to the attack that produced it. Every score is a rate you can inspect, not an opinion you have to trust.
Grounded Adversarial Testing
We generate adversarial probes grounded in your system’s own knowledge base — the articles of the law you operate under, the products in your catalog, the policies you publish. Each probe exercises a failure that actually matters in your domain, and every attack becomes a registered, hash-citable trace.
OneCheck
Your AI, Red-Teamed
A full red-team of your AI system, scored across all 8 dimensions of trust, delivered in 3 weeks.
What you get
Best for
Organizations that need to understand their AI risk posture before committing to a platform.
Enterprise
Full PlatformContinuous Certification
The full platform for continuous certification: technical and administrative evidence matched to every framework requirement.
Everything in OneCheck, plus
Best for
Organizations deploying AI at scale that need continuous compliance assurance.
75%
Reduction in compliance violations detected
100%
Audit trail coverage for all AI decisions
Seconds
Time to detect compliance drift
7+ Years
Immutable audit trail retention
“Deployed with a Fortune 500 financial services organization managing 100+ AI systems in a regulated environment. $265K first-year engagement. Live in production.”
The Deliverable
Risk Classification — ETHI-202
11 / 15 points — HIGH RISK
Regulatory Implications
Critical Findings
Incorrect Deposit Insurance Information
Chatbot states €200,000 limit when actual EU limit is €100,000 per depositor.
Missing MiFID II Suitability Assessment
23% of recommendation conversations skip required risk profiling step.
Key Recommendations
Dimensional Scorecard
EthiCompass
AI Ethics & Compliance
Evaluation Report
EuroBank Virtual Assistant v3.2
Generative AI — Financial Services
Risk
HIGH
Intake
7.6
Score
7.8
Client
EuroBank AG
Frankfurt, Germany
Evaluator
EthiCompass
7-Dimension Framework
Sample Report — Demonstration Purposes
EthiCompass
AI Ethics & Compliance
Evaluation Report
EuroBank Virtual Assistant v3.2
Generative AI — Financial Services
Risk
HIGH
Intake
7.6
Score
7.8
Client
EuroBank AG
Frankfurt, Germany
Evaluator
EthiCompass
7-Dimension Framework
Sample Report — Demonstration Purposes
Dimensional Scorecard
Critical Findings
Incorrect Deposit Insurance Information
Chatbot states €200,000 limit when actual EU limit is €100,000 per depositor.
Missing MiFID II Suitability Assessment
23% of recommendation conversations skip required risk profiling step.
Key Recommendations
Risk Classification — ETHI-202
11 / 15 points — HIGH RISK
Regulatory Implications
A quantified risk posture across every AI system, defensible with the trace behind each score.
EU AI Act and GDPR requirements matched to technical and administrative evidence in a single view.
API-first, grounded adversarial testing on your own systems, with no latency impact in production.
Governance you can stand behind — evidence that your AI behaves, not assurances that it does.
Start with a OneCheck red-team to see how your AI behaves, or talk to our team about continuous certification.