EthiCompass

Protection · Evidence · Trust

Protect your AI.
Protect what it can affect.

Every AI system you deploy makes decisions in your name — and you carry the consequences. We show you how your AI can fail, what those failures can affect, and give you verifiable evidence to act before they become consequences.

Why independence →
8 Dimensions of Observable Behavior·ISO 42001·Immutable Audit Trail
Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Behaviour & Exposure
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Runs

40

Coverage

82%

Findings

14

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Occurrence Rate by Dimension

Fairness & Non-Discrimination
7/40OBSERVED
Safety & Harmful Content
0/40NOT OBSERVED
Transparency & Explainab.
14/40OBSERVED
Privacy & Data Protection
2/40OBSERVED
Factuality & Accuracy
9/40OBSERVED
Robustness & Adv. Resil.
5/40OBSERVED
Security & Access Control
3/40OBSERVED
Accountability & Oversight
—NOT COVERED
Coverage
82%

9 of 50 probes not run — reported, not scored as zero

Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Stated a €200,000 limit in 7 of 40 runs; the EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

9 of 40 recommendation conversations skipped the required risk-profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

EU AI Act Classification

Minimal
Limited
High
Unacceptable

Annex III — credit scoring of natural persons

Triggering Criteria

Vulnerable groups affected
Sector listed in Annex III
Influences access to an essential service
Decision is not readily reversible
Population in scope: 2.3M

Obligations That May Apply

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
Explore the Full 18-Page Report→

Your AI Acts in Your Name.
You Carry the Consequences.

Your AI answers your customers, moves your money, reaches your data, operates your tools and speaks with your authority — at a scale no human review can keep up with.

So there are two things to protect. The system itself, against prompt injection, jailbreaks, poisoned context and tools, and actions it was never authorized to take. And everything it can reach: the people, the data, the decisions, the assets, the processes, the reputation that stands behind it.

The gap isn’t intent — you already mean to use AI well. The gap is evidence: something concrete you can point to that shows how your AI actually behaves, and what that behavior can reach.

8

Dimensions of behavior, each scored and independently traceable

5

Framework axes the same evidence projects into — threat, security, risk, governance, regulatory

N runs

Every risk is an occurrence rate you can inspect, never a single anecdote

The Methodology

Evidence You Can Act On.

Our 8-dimension framework was developed by PhD researchers and validated through peer-reviewed publications.

Because the AI systems we test are non-deterministic, we report registered, hash-citable attack traces and statistical evidence: the rate at which each failure occurs across N runs, with Wilson confidence intervals and Rogan-Gladen correction for judge reliability.

When a budget limits coverage, we report exactly what was measured and what wasn’t. Every score traces back to the attack that produced it — which is also what lets an auditor verify it later, if you need it to.

Explore the framework →

PEER-REVIEWED PUBLICATIONS

Research published in AI bias, compliance frameworks, and ethical evaluation methodology

PhD RESEARCH TEAM

In-house researchers with doctoral expertise in AI/ML and ethics

ACADEMIC COLLABORATION

Ongoing partnership with universities and research centers for methodology validation

TRANSPARENT METHODOLOGY

Every scoring criterion is documented, versioned, and disclosed in full with the engagement

The Framework

Eight Dimensions of
Observable Behavior.

Each dimension scores your AI system’s behavior and stays independently traceable to the run that produced it. The dimensions classify the evidence — they are not a framework. That same evidence then projects into the threat, security, risk, governance and regulatory frameworks you report against, none of which decides what we are able to see.

01

FAIRNESS & NON-DISCRIMINATION

Detects bias, stereotyping, and unequal treatment across protected groups, using statistical and counterfactual testing

→ Disparate impact

02

SAFETY & HARMFUL CONTENT

Surfaces hate, violence, self-harm, sexual content, and dangerous instructions — including harm that carries no obviously toxic wording

→ Harmful output

03

TRANSPARENCY & EXPLAINABILITY

Measures unfaithful explanations, sycophancy, and undisclosed AI — whether a decision can be understood by the people it affects

→ Unfaithful explanation

04

PRIVACY & DATA PROTECTION

Probes PII leakage, training-data extraction, and system-prompt disclosure across your data-governance boundaries

→ Data exposure

05

FACTUALITY & ACCURACY

Verifies outputs against source material to catch hallucination, fabricated citations, and unsupported claims

→ Fabrication

06

ROBUSTNESS & ADVERSARIAL RESILIENCE

Tests resistance to prompt injection, jailbreaks, encoding tricks, and adversarial suffixes that bend the model's behavior

→ Adversarial input

07

SECURITY & ACCESS CONTROL

Exercises unauthorized actions, privilege abuse, identity spoofing, and tool misuse wherever your AI can act, not just answer

→ Unauthorized action

08

ACCOUNTABILITY & HUMAN OVERSIGHT

Checks oversight saturation, governance evasion, and whether every action stays traceable to the human accountable for it

→ Oversight failure

Every finding is traceable to the attack that produced it. Every score is a rate you can inspect, not an opinion you have to trust.

Grounded Adversarial Testing

We attack your AI
with your own world.

We generate adversarial probes grounded in your system’s own knowledge base — the articles of the law you operate under, the products in your catalog, the policies you publish. Each probe exercises a failure that actually matters in your domain, and every attack becomes a registered, hash-citable trace.

The Evidence Architecture

One observation.
Five independent readings.

Everything we observe converges into one preserved object: what happened, how often across N runs, and the trace that proves it. From there the same evidence reads five different ways — as a threat, as a control failure, as risk, as a governance requirement, as a legal obligation. Each reading is a projection. None of them is the structure underneath.

A diagram in three bands. Five kinds of observation — grounded probes, multi-turn adversarial attacks, tool invocations and agent actions, model outputs, and administrative attestations — converge into a single preserved evidence object, shown here with an illustrative example. That one object then diverges into five independent framework readings, listed below it: threat, security and control, risk, governance, and regulatory. No reading sits above the others, and none of them is the structure the evidence rests on.

What we observe

Grounded probes from your own world

Multi-turn adversarial attacks

Tool invocations and agent actions

Model outputs and refusals

Administrative attestations

Classified as

Security & Access ControlAccountability & Human Oversight

Illustrative example

The agent invoked the payment tool without valid authorization.

Occurred in 7 of 40 runs · coverage 82%, 9 probes not run · trace 4f2a…c19

Change a framework and only the reading below changes. The evidence above it does not move.

Threat

MITRE ATLAS

at a pinned snapshot

the adversarial technique used

Security & control

OWASP Agentic

at a pinned edition

agent authorization and control failure

Risk

NIST AI RMF

version 1.0

unauthorized financial action and its impact

Governance

ISO/IEC 42001

2023 edition

the control requirements this speaks to

Regulatory

EU AI Act

at the applicable text

the obligation that may apply in context

What you change, what you govern, and when your AI deserves your trust.

Why Us

We Don’t Sell the System
We’re Asked to Examine.

An evaluation is worth what its evaluator can afford to find. We build no model, host no agent and ship no guardrail, so there is nothing in our catalogue that a finding could embarrass. That is the whole reason our evidence is worth handing to someone whose job is to disbelieve it.

01

No stake in the verdict

We do not build the models, host the agents or sell the guardrails we test. A finding costs us nothing, which is precisely what makes it credible to a regulator, a board or a customer who asked for assurance.

02

One view across your clouds

Native tooling is strong inside its own platform and goes quiet at the boundary. Your agents do not all live in one place, so the evidence about them is brought into a single view rather than three dashboards that never agree.

03

Evidence a reviewer can check

Every finding carries the trace that produced it, the method that ran it and the coverage around it, each citable by hash. A reviewer can check the evidence against the record rather than against our summary of it.

Evidence That Holds Up
Line by Line.

Every run resolves into one Score Card: what we observed, what it can affect, what deserves attention, and how strong the evidence behind each finding is — recorded in an immutable audit trail that you, or an auditor, can re-verify independently.

●

Multi-turn adversarial attacks across all 8 dimensions of behavior

●

A Score Card: per-dimension risk as an occurrence rate over N runs

●

Registered, hash-citable attack traces behind every finding

●

Can run entirely inside your own infrastructure, with the scope agreed in writing before anything is pointed at anything

SOC 2 ControlsEU AI Act MappedGDPR-ReadyEncrypted End-to-End

The Deliverable

A Score Card for every
AI system you deploy.

Confidential
Doc Ref: ETHIC-RPT-2026-00147
Version: 1.0 — Final
Eval ID: eval_mock_eurobank
Date: March 15, 2026

EthiCompass

AI Behaviour & Exposure
Evaluation Report

EuroBank Virtual Assistant v3.2

Generative AI — Financial Services

HIGH RISK — EU AI Act Annex III

Runs

40

Coverage

82%

Findings

14

Client

EuroBank AG

Frankfurt, Germany

Evaluator

EthiCompass

8-Dimension Framework

Sample Report — Demonstration Purposes

EthiCompassCONFIDENTIAL

Occurrence Rate by Dimension

Fairness & Non-Discrimination
7/40OBSERVED
Safety & Harmful Content
0/40NOT OBSERVED
Transparency & Explainab.
14/40OBSERVED
Privacy & Data Protection
2/40OBSERVED
Factuality & Accuracy
9/40OBSERVED
Robustness & Adv. Resil.
5/40OBSERVED
Security & Access Control
3/40OBSERVED
Accountability & Oversight
—NOT COVERED
Coverage
82%

9 of 50 probes not run — reported, not scored as zero

Page 6 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

Critical Findings

P07 Day Deadline

Incorrect Deposit Insurance Information

Stated a €200,000 limit in 7 of 40 runs; the EU limit is €100,000 per depositor.

P014 Day Deadline

Missing MiFID II Suitability Assessment

9 of 40 recommendation conversations skipped the required risk-profiling step.

Key Recommendations

PActionRef
P0Fix deposit insurance to €100KDir. 2014/49
P0Add MiFID II suitability gateMiFID II Art.25
P1Add AI disclosure to responsesAI Act Art.52
P1Implement explanation moduleAI Act Art.13
P1Add confidence indicatorsAI Act Art.14
Page 7 of 18eval_mock_eurobank_2026Q1
EthiCompassCONFIDENTIAL

EU AI Act Classification

Minimal
Limited
High
Unacceptable

Annex III — credit scoring of natural persons

Triggering Criteria

Vulnerable groups affected
Sector listed in Annex III
Influences access to an essential service
Decision is not readily reversible
Population in scope: 2.3M

Obligations That May Apply

Conformity assessment (Art. 43)
EU AI database registration (Art. 49)
Fundamental rights assessment (Art. 27)
Quality management system (Art. 17)
Post-market monitoring (Art. 72)
Incident reporting (Art. 73)
Page 4 of 18eval_mock_eurobank_2026Q1
Explore the Full 18-Page Report→

Built for the People Who Own AI Risk.

FOR THE CISO

Adversarial testing against the agents you already run, through the endpoint they already expose. No SDK in your runtime, no change to how the agent is served, and it can run entirely inside your own infrastructure.

FOR THE ENGINEER

The exact turn that broke, reproducible, with the trace behind it. Risk arrives as an occurrence rate over N runs with a coverage report — not a number you have to take on faith.

FOR THE BOARD AND THE AUDITOR

Evidence that survives someone whose job is to disbelieve it: the method, the coverage, and every finding citable by hash.

Your AI Is Already Deployed.
Your Evidence Should Be Too.

We red-team the AI you run and hand you what we found, with the method and the coverage attached. We won’t tell you to trust your AI — you leave with the evidence to decide when you should.