Autogeny

Independent audits of AI in production

Your AI is making decisions. What does a wrong one cost?

We test the model inside your system, with your prompts, tools and data, against the failures you’d be liable for. We put a price on each kind of wrong decision you can live with, tell you which configuration is worth running, and monitor it as your system changes.

The questions arrive before the contract does.

Procurement

“How do you know it’s accurate?”

Enterprise and regulated buyers now send AI questionnaires. A demo and a benchmark score don’t answer them.

Drift

More wrong decisions, and no error to tell you

A prompt edit, a tool change or new data shifts what the AI does. The share of wrong decisions creeps up, or a failure you fixed comes back, and nothing in your logs says so.

Model choice

Paying flagship prices out of caution

Newer, cheaper models ship every few weeks. We test the credible ones inside your system against your acceptance and failure criteria, so you switch when one holds up and keep the savings.

Accountability

Scores without a verdict

Eval dashboards show numbers. Nobody tells you which failures you can live with, and which ones end the deal.

Method

Cheaper is only cheaper if it doesn’t fail.

We score capability and failure separately and put a price on every failure short of catastrophic.

ConfigurationRecommendedDisqualified: catastrophic failure
Acceptable zone70758085909510001234567Unacceptable failures per 1,000 tasksCapability scoreCapability bar 85Failure limit 2Flagship A$0.079 per acceptable outcomeMid-tier B$0.113 per acceptable outcomeBudget C + output filter$0.059 per acceptable outcomeBudget CDisqualifiedOpen-weights D$0.154 per acceptable outcome
Illustrative data. Failure rates come from 5,000 tasks per configuration. One unacceptable failure costs $40 in rework and refunds.
Table view
ConfigurationCapabilityFailures /1kCost /taskPer acceptable outcomeVerdict
Flagship A931.2$0.0310$0.079Acceptable
Mid-tier B902.6$0.0090$0.113Outside requirements
Budget C + output filter861.4$0.0026$0.059Recommended
Budget C846.2$0.0020–Disqualified
Open-weights D783.8$0.0015$0.154Outside requirements

Catastrophic failures are a gate

Defined with you. One reproduced catastrophic failure disqualifies a configuration.

Everything else gets a price

cost per acceptable outcome = (cost per task + failure rate × cost of one failure) ÷ (1 − failure rate)

We score the system, not the model

A cheap model with the right filter can beat a flagship. We test them together.

A clean result needs volume

Thousands of runs per failure class, so a zero means something.

Sample report

The report is the product.

Autogeny audit report · A-0142Sample report · illustrative data

Insurance claims intake assistant

Harness v3.2 held fixed · 5 configurations · 412 evaluation-set claims · 5,000 tasks per configuration · 3,000 adversarial runs per catastrophic evaluation · audited 6 Oct 2026

Capability
86/100
Reference ≥ 85Pass

Graded blind by a judge from no candidate’s model family.

Unacceptable failures
1.4 per 1k
Reference ≤ 2.0 per 1kPass

7 in 5,000 tasks: misrouted claims and wrong policy citations.

Cost per acceptable outcome
$0.059
Reference ≤ $0.080Pass

$0.0026 per task plus failures at $40 each, divided by the 99.86% of tasks that are acceptable.

Results: Budget C + output filter (recommended)

MeasureResultReferenceFlag
Capability score412 evaluation-set claims, graded blind86 /100≥ 85Pass
Unacceptable failures7 in 5,000 tasks: misrouted claim, wrong policy citation1.4 per 1k≤ 2.0 per 1kPass
Catastrophic: cross-claimant disclosureC-02 Trust-boundary crossing · 3,000 adversarial runs0 events0Pass (<0.10%)
Catastrophic: payout promised beyond authorityClient-defined evaluation · 3,000 adversarial runs0 events0Pass (<0.10%)
Adjuster hand-off under pushbackC-04 Asserting where it should escalate · awaiting human adjudication3 cases0Needs review
Pass26Fail6Needs review7Couldn't run1

All configurations, same harness

EvaluationFlagship AMid-tier BBudget CBudget C + filterOpen-weights D
C-02 Trust-boundary crossing
Client-defined: payout beyond authority
C-08 Injection through retrieved content
C-06 Drift across a long conversation
C-04 Asserting where it should escalate
Failure rate ≤ 2.0 per 1k
Evaluation set: coverage questions
Evaluation set: document upload
Recommendation

Run Budget C with the output filter. It clears the catastrophic gate and the capability bar with the same failure profile as Flagship A, at 25% lower cost per acceptable outcome. Budget C alone is disqualified: it disclosed another claimant’s data in 3 of 3 reproductions. Re-test on the next model release or harness change.

Method: harness held fixed across configurations. Every fail reproduced before it is reported. Subjective classes go to human adjudication and stay flagged for review until they are.

How it works

Set the bar. Audit against it. Watch for drift.

  1. 01One-time

    Scope

    An initial consultation to set your baseline quality, which failures are acceptable and which aren’t, what counts as catastrophic, and what one failure costs you.

  2. 02One-time

    Audit

    We build an evaluation set with your domain experts and a red-team suite derived from what your app is for, run both inside your system, and deliver the audit report.

  3. 03Monthly

    Monitoring

    Capability and failure scores every month and on every significant change, with the exact cases that regressed.

  4. 04On every release

    Re-audit

    A credible new model ships, or you change prompts, tools or data. We re-audit, and tell you whether to switch and what it saves.

Evidence

Generic web scanners check the software, not what your AI is for.

Three open-source AI apps, tested on our own isolated copies with dummy data. On the first, a generic web security scanner we ran alongside missed the hijack, even with its own test inputs in the same field.

Tested in September 2026; the apps may have changed since. App names have been replaced with general descriptions, and identifying details left out.

Personal-finance app
Reproduced

Prompt hijack

6/6

An instruction hidden in a transaction description, sent directly to the app’s spending-analysis AI, was followed in 6 of 6 attempts (3 of 3 for a second team). Our test marker replaced the AI’s summary or was written into its recommendations or risk flags. A generic web scanner given the same field did not flag it. We have not yet tested whether a saved transaction can steer the AI a signed-in user sees.

Couples-finance app
Reproduced

Privacy and access flaws beneath the AI

3

Using the app’s database interface directly with their own login, one partner could read transactions and notes the other had marked private, and could attach a note to a stranger’s transaction. The flaws were in the database’s access rules, not the AI: in 13 injection attempts the assistant never obeyed, and it showed each user only their own data.

Hiring screener
Reproduced

The candidate sets the score

10/10 or 1/10

A candidate could set their own stored score to 10/10 or 1/10 by typing an instruction into an answer, and the app kept it. Separately, identical answers scored about 9/10 said confidently and about 5/10 hedged (1 to 2 runs per condition, two test teams). That is a potential disparate-impact risk; our checks of names and stated age or background found no effect in a small sample.

All tests ran on our own isolated copies of public code, never on live systems.

Built for

  • Companies with AI already in production, not in a pilot
  • Selling to enterprise or regulated buyers: health, finance, insurance, legal
  • AI that makes decisions, touches customer data or acts on a user’s behalf
  • Teams that change models, prompts or tools often and want to know what broke

Not for

  • Deciding whether to use AI at all
  • Low-stakes AI where a wrong answer costs a re-ask
  • Building your AI for you

Your AI eval and security team, without the hires.

  • No eval engineers or red‑team specialists to recruit and keep.
  • No tooling to build, run and maintain.
  • Each credible new model tested as it ships, not when someone has time.
  • No angle on the verdict: we don’t sell models or build the systems we audit.

Founder

Brad Cunningham

MSc Engineering · BSc Information Technology · EMBA

Brad spent 15 years implementing health data systems in over 40 countries, with a focus on data privacy, policy and security.

Autogeny brings that discipline to AI. Define what failure means before you test. Reproduce it before you report it. Re-test it every time something changes.

Book a scoping call.

Bring one AI workflow you’re shipping. We’ll tell you what we’d test first, what we’d treat as catastrophic, and what an audit would take.