EthiCompass

Protection · Evidence · Trust

Protect your AI.
Protect what it can affect.

Every AI system you deploy makes decisions in your name — and you carry the consequences. We show you how your AI can fail, what those failures can affect, and give you verifiable evidence to act before they become consequences.

Or explore Check →
8 Dimensions of Observable Behavior·ISO 42001·Immutable Audit Trail
Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Ethics & Compliance
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Risk

HIGH

Onboarding

7.6

Score

7.8

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Dimensional Scorecard

Fairness & Non-Discrimination
7.2COND
Safety & Harmful Content
9.4PASS
Transparency & Explainab.
6.1ACTION
Privacy & Data Protection
8.5PASS
Factuality & Accuracy
7.8COND
Robustness & Adv. Resil.
8.1COND
Security & Access Control
8.3PASS
Accountability & Oversight
7.4COND
Composite Score
7.8/10CONDITIONAL
Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Chatbot states €200,000 limit when actual EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

23% of recommendation conversations skip required risk profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Risk Classification — ETHI-202

MINIMAL
LIMITED
HIGH
UNACC.

11 / 15 points — HIGH RISK

FactorPtsMax
Vulnerable Groups Affected33
Sector in EU AI Act Annex III33
Decision Type13
Reversibility12
Population Scale (2.3M)33
TOTAL1115

Regulatory Implications

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
Explore the Full 18-Page Report

Your AI Acts in Your Name.
You Carry the Consequences.

Your AI answers your customers, moves your money, reaches your data, operates your tools and speaks with your authority — at a scale no human review can keep up with.

So there are two things to protect. The system itself, against prompt injection, jailbreaks, poisoned context and tools, and actions it was never authorized to take. And everything it can reach: the people, the data, the decisions, the assets, the processes, the reputation that stands behind it.

The gap isn’t intent — you already mean to use AI well. The gap is evidence: something concrete you can point to that shows how your AI actually behaves, and what that behavior can reach.

8

Dimensions of behavior, each scored and independently traceable

5

Framework axes the same evidence projects into — threat, security, risk, governance, regulatory

N runs

Every risk is an occurrence rate you can inspect, never a single anecdote

The Methodology

Evidence You Can Act On.

Our 8-dimension framework was developed by PhD researchers and validated through peer-reviewed publications.

Because the AI systems we test are non-deterministic, we report registered, hash-citable attack traces and statistical evidence: the rate at which each failure occurs across N runs, with Wilson confidence intervals and Rogan-Gladen correction for judge reliability.

When a budget limits coverage, we report exactly what was measured and what wasn’t. Every score traces back to the attack that produced it — which is also what lets an auditor verify it later, if you need it to.

Explore the framework →

PEER-REVIEWED PUBLICATIONS

Research published in AI bias, compliance frameworks, and ethical evaluation methodology

PhD RESEARCH TEAM

In-house researchers with doctoral expertise in AI/ML and ethics

ACADEMIC COLLABORATION

Ongoing partnership with universities and research centers for methodology validation

TRANSPARENT METHODOLOGY

Every scoring criterion is documented, versioned, and publicly auditable

The Framework

Eight Dimensions of
Observable Behavior.

Each dimension scores your AI system’s behavior and stays independently traceable to the run that produced it. The dimensions classify the evidence — they are not a framework. That same evidence then projects into the threat, security, risk, governance and regulatory frameworks you report against, none of which decides what we are able to see.

01

FAIRNESS & NON-DISCRIMINATION

Detects bias, stereotyping, and unequal treatment across protected groups, using statistical and counterfactual testing

Disparate impact

02

SAFETY & HARMFUL CONTENT

Surfaces hate, violence, self-harm, sexual content, and dangerous instructions — including harm that carries no obviously toxic wording

Harmful output

03

TRANSPARENCY & EXPLAINABILITY

Measures unfaithful explanations, sycophancy, and undisclosed AI — whether a decision can be understood by the people it affects

Unfaithful explanation

04

PRIVACY & DATA PROTECTION

Probes PII leakage, training-data extraction, and system-prompt disclosure across your data-governance boundaries

Data exposure

05

FACTUALITY & ACCURACY

Verifies outputs against source material to catch hallucination, fabricated citations, and unsupported claims

Fabrication

06

ROBUSTNESS & ADVERSARIAL RESILIENCE

Tests resistance to prompt injection, jailbreaks, encoding tricks, and adversarial suffixes that bend the model's behavior

Adversarial input

07

SECURITY & ACCESS CONTROL

Exercises unauthorized actions, privilege abuse, identity spoofing, and tool misuse wherever your AI can act, not just answer

Unauthorized action

08

ACCOUNTABILITY & HUMAN OVERSIGHT

Checks oversight saturation, governance evasion, and whether every action stays traceable to the human accountable for it

Oversight failure

Every finding is traceable to the attack that produced it. Every score is a rate you can inspect, not an opinion you have to trust.

Grounded Adversarial Testing

We attack your AI
with your own world.

We generate adversarial probes grounded in your system’s own knowledge base — the articles of the law you operate under, the products in your catalog, the policies you publish. Each probe exercises a failure that actually matters in your domain, and every attack becomes a registered, hash-citable trace.

The Evidence Architecture

One observation.
Five independent readings.

Everything we observe converges into one preserved object: what happened, how often across N runs, and the trace that proves it. From there the same evidence reads five different ways — as a threat, as a control failure, as risk, as a governance requirement, as a legal obligation. Each reading is a projection. None of them is the structure underneath.

What we observe

Grounded probes from your own world

Multi-turn adversarial attacks

Tool invocations and agent actions

Model outputs and refusals

Administrative attestations

Classified as

Security & Access ControlAccountability & Human Oversight

Illustrative example

The agent invoked the payment tool without valid authorization.

Occurred in 7 of 40 runs · coverage 82%, 9 probes not run · trace 4f2a…c19

Change a framework and only the reading below changes. The evidence above it does not move.

Threat

MITRE ATLAS

at a pinned snapshot

the adversarial technique used

Security & control

OWASP Agentic

at a pinned edition

agent authorization and control failure

Risk

NIST AI RMF

version 1.0

unauthorized financial action and its impact

Governance

ISO/IEC 42001

2023 edition

the control requirements this speaks to

Regulatory

EU AI Act

at the applicable text

the obligation that may apply in context

What you change, what you govern, and when your AI deserves your trust.

What We Build

Two Products.
One Standard of Evidence.

Proof protects the AI that acts in your name. Check protects what you put your name behind. Different objects, one discipline: observe, verify, preserve the evidence, understand the exposure, support the human decision. Every finding is a registered, hash-citable trace and a founded opinion — never a verdict.

Proof

Red-team your AI

Adversarial red-teaming of your AI system. We generate multi-turn attacks, then type and score what breaks — by dimension, with a coverage report and an occurrence rate over N runs.

What you get

  • Multi-turn adversarial attacks across all 8 dimensions of behavior
  • A Score Card: per-dimension risk as an occurrence rate over N runs
  • Registered, hash-citable attack traces behind every finding
  • Available as Proof · Cloud, pay per analysis, or on-premise with multi-framework readiness

Ideal for

Teams putting AI in front of customers or regulators.

Explore Proof →

Check

Verify your content

Regulatory verification of your content and documents, human- or AI-authored. We hold each one against the requirements of the regulation it must answer to, and ground every objection in a citable source.

What you get

  • Requirement-driven verification against the regulation your content must answer to
  • Every objection grounded in a citable source — a demonstrable gap, not an opinion on truth
  • Two exports: evidence-only for the auditor, remediation for whoever revises the document
  • Content that is human- or AI-authored, across regulated domains

Ideal for

Anyone who submits content to a regulator, court, or board.

Explore Check →

Evidence That Holds Up
Line by Line.

Every run resolves into one Score Card: what we observed, what it can affect, what deserves attention, and how strong the evidence behind each finding is — recorded in an immutable audit trail that you, or an auditor, can re-verify independently.

SOC 2 ControlsEU AI Act MappedGDPR-ReadyEncrypted End-to-End

The Deliverable

A Score Card for every
AI system you deploy.

Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Ethics & Compliance
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Risk

HIGH

Onboarding

7.6

Score

7.8

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Dimensional Scorecard

Fairness & Non-Discrimination
7.2COND
Safety & Harmful Content
9.4PASS
Transparency & Explainab.
6.1ACTION
Privacy & Data Protection
8.5PASS
Factuality & Accuracy
7.8COND
Robustness & Adv. Resil.
8.1COND
Security & Access Control
8.3PASS
Accountability & Oversight
7.4COND
Composite Score
7.8/10CONDITIONAL
Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Chatbot states €200,000 limit when actual EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

23% of recommendation conversations skip required risk profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Risk Classification — ETHI-202

MINIMAL
LIMITED
HIGH
UNACC.

11 / 15 points — HIGH RISK

FactorPtsMax
Vulnerable Groups Affected33
Sector in EU AI Act Annex III33
Decision Type13
Reversibility12
Population Scale (2.3M)33
TOTAL1115

Regulatory Implications

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
Explore the Full 18-Page Report

Built for the People Who Own AI Risk.

FOR THE CRO

A quantified risk posture across every AI system, defensible with the trace behind each score.

FOR THE DPO

EU AI Act and GDPR requirements matched to technical and administrative evidence in a single view.

FOR THE CISO

API-first, grounded adversarial testing on your own systems, with no latency impact in production.

FOR THE BOARD

The evidence to decide when your AI deserves your trust — and to say why, on the record.

Your AI Is Already Deployed.
Your Evidence Should Be Too.

Start with a Proof red-team of the AI you run, or a Check review of the content you stand behind. We won’t tell you to trust your AI — you leave with the evidence to decide when you should.