Semantic failure-mode assurance for AI deployments

Your AI passed every test.
It's still wrong.

Scanners test whether it can be broken — we test whether it can be wrong in the way that costs you.

claims-triage · same deployment, same daydemo
generic-scanner --target triage-agent --full prompt injection ······················· PASS jailbreak resistance ··················· PASS refusal bypass ························· PASS data exfiltration ······················ PASS harmful content ························ PASS 47/47 checks passed · 0 findings   autogeny assess --spec claims-triage.spec failure modes elicited from claims supervisor · 14 C-03 severity misclassification ········ REACHABLE └─ plausible fraud pattern tiered as routine payout 1 reachable failure · consequence: financial
Nothing in the output was harmful. The system failed at the thing it was built to do.
01
The problem

Real failures don't look like attacks.

C-01 · Missed detectionA safety monitor misses a disclosure phrased in euphemism.Scanner: pass
C-03 · Severity misclassificationA triage agent mis-tiers a plausible adversarial case.Scanner: pass
C-02 · Trust-boundary crossingA copilot answers helpfully, using a document the user was never entitled to see.Scanner: pass
C-07 · Compositional harmAn agent chains six individually-safe actions into one unauthorised transaction.Scanner: pass

Every scanner on the market passes all four. Nothing in the output is harmful. The system simply failed at the thing it was built to do — and that is the failure that carries legal, clinical and financial consequence.

02
An honest split

Half of this problem is already solved. Free.

Structural testing — commodityJailbreaks, injection strings, refusal bypass.

Open-source harnesses do this well, and the layer is commoditising to zero. Run it free, in CI, on every build. If a vendor charges you for this layer, ask what else they're wrong about. We'll tell you which tools to use.

Semantic testing — the gapWhether the system is wrong at its actual job.

No generic harness can test what it doesn't know your system is for. Finding these failures requires your deployment's purpose, your data boundaries, and your domain's definition of catastrophe — extracted and converted into adversarial test design.

03
The method

Hazard analysis, pointed at an AI system.

Fifty years of safety engineering already solved how to find failure modes in complex systems. We adapted the four canonical techniques to AI deployments.

Your fraud lead, clinical director or claims supervisor already knows what catastrophic failure looks like. Nobody has asked them in a form that produces a test.

FMEARank failure modes by severity, occurrence, detectability
FTAStart at the catastrophic outcome, work backwards
HAZOPGenerate deviations from intent, systematically
STPAFind harm emerging between components, not within them

Findings map to NIST AI RMF · OWASP GenAI LLM Top 10 · MITRE ATLAS — so reports survive your governance chain.

04
What we look for

Eight ways a correct-looking system gets it wrong.

C-01Missed detection in your risk category
C-02Trust-boundary crossing
C-03Severity misclassification
C-04Asserting where it should escalate
C-05Subpopulation blind spot
C-06Drift across a long conversation
C-07Safe actions composing into harm
C-08Injection through retrieved content
Domain-independent — the reason the method transfers between industries.
05
Engagements

A closed system, not a project.

Your model provider updates the model underneath you. The assessment has to repeat on material change — so the engagement is built as a cycle from the start.

Specentry
Design
failure mode
Semantic
assessment
Adversarial
build
Continuous
assurance
repeats on material change
Your model changes underneath you. So does the test.
06
Current offer

A no-fee assessment, in exchange for the findings.

We're running a comparative study: generic harnesses against semantic testing, on real deployments. Results get published either way — including where the free tools win.

You get
  • Full failure-mode elicitation session
  • Adversarial testing built for your deployment
  • Findings report and remediation plan
  • Generic-harness baseline for comparison
We need
  • A real deployment, not a prototype
  • 90 minutes with one domain expert
  • Written authorisation to test
  • Permission to publish anonymised findings

Publication terms are specific: your company is never named, no data or prompts are published, no reproducible attack strings are disclosed, and you get ten business days to review anything describing your case. Scope and authorisation are agreed in writing before anything is tested.

Request an assessment or email hello@autogeny.ai · we reply within two business days