Engineering case

Accountable assurance for agentic AI.

Codance's work begins with an engineering proposition about the conditions required for a valid assurance outcome. From it, we derive a testable engineering hypothesis that Project Aurora is built to examine in practice.

Origin The thesis draws on founder Jonathan Arklay's 42 years of engineering, systems and process experience to frame a specific engineering proposition that can be examined through a falsifiable programme of work.

The thesis

Agentic AI can be appropriately assured only through an evidential process.

Every material action, decision, delegation and consequence of both the probabilistic agent and the application in which it operates must remain attributable, evidentially reconstructible and open to competent challenge.

Declared boundaries, equivalent evidential rigour applied to the assurance apparatus itself, and legitimate human authority are prerequisites to a valid assurance outcome.

Equivalent evidential rigour includes resistance to influence. The assurance apparatus cannot assume that the system under test will remain passive toward the process observing it. Where material interference is possible, the method should be able to establish authorised state, detect meaningful deviation and preserve that deviation as evidence for competent challenge.

Conditions of valid assurance

A valid assurance outcome requires three distinct conditions.

Aurora is designed to accommodate and preserve all three within a controlled assurance campaign, but it cannot legitimately originate them all. Declared boundaries and human authority must come from competent parties outside the platform, while the evidential rigour of Aurora and the assurance procedure must itself remain open to independent challenge.

Declared boundaries and legitimate human authority enter an Aurora campaign from competent external parties, while Aurora provides evidential machinery that remains open to independent challenge. All three are required conditions for a valid assurance outcome.

Declared boundaries

Competent parties define the system, interfaces, behaviours, risks, evidence and acceptance conditions that are actually in scope. Aurora can record, enforce and preserve that declaration; it does not create the competence or legitimacy behind it.

Evidential rigour

Aurora's principal engineering contribution is to execute assurance methods while retaining attributable evidence of both the system under test and material influence exerted by the assurance procedure. The adequacy of that machinery remains challengeable.

Legitimate human authority

Evidence can support a judgement, but authority comes from competent people acting within an appropriate mandate. Aurora can retain that authority and its relationship to a campaign; it does not confer it.

Independent challenge is part of the method, not an endorsement layer. Codance can build and operate the apparatus. Consistently with the thesis, it cannot treat self-authored boundaries, self-assessed evidential adequacy and self-conferred authority as sufficient.

The engineering hypothesis

Can assurance adapt without losing its evidential integrity?

An adaptive assurance campaign can combine deterministic and probabilistic methods while preserving sufficient external evidence for a competent reviewer to reconstruct and challenge both the material operation of the system under test and every material influence exerted by the assurance procedure itself.

Aurora is built to test this proposition in practice. It can fail. Evidence may prove incomplete; a method may not expose the right interfaces; adaptive steps may exert influence that cannot be sufficiently reconstructed; or a competent reviewer may conclude that the retained record is not adequate to support the claim.

The hypothesis tests Aurora's engineering contribution to the thesis. It does not replace the external boundary-setting or legitimate human authority required for a valid assurance outcome.

Recursive assurance

The test apparatus cannot disappear behind its own report.

If an LLM proposes a probe, selects a guide, interprets an observation, resolves model work or materially shapes a finding, that influence is part of the assurance procedure. Aurora therefore requires material assurance influence to remain attributable, evidentially reconstructible and open to challenge alongside the system under test.

Material operation of the system under test and material influence from the assurance apparatus are both retained as evidence for competent review before human judgement.

System under test

The agent and the application in which it operates must remain externally attributable and reconstructible at material points.

Assurance apparatus

Adaptive or probabilistic components that materially influence the campaign are subject to equivalent evidential discipline.

Human authority

The evidential record informs competent judgement. A machine-generated finding does not acquire legitimate authority merely because a machine produced it.

Claim boundaries

Strong propositions need restrained claims.

The thesis is deliberately bounded. It proposes conditions for a valid assurance process; it does not confer trust, claim access to unobservable model cognition, define universal ethics or imply that probabilistic methods belong everywhere.

Evidence does not certify trust.

Qualified assurance, regulatory, certification or other competent authorities determine what conclusion the evidence justifies within their own remit.

Reconstruction is bounded by observability.

Aurora identifies where probabilistic model work materially influenced an execution and retains the evidence available around it. Where internal telemetry or interpretability evidence is lawfully exposed, Aurora can retain that too; it cannot reconstruct signals the system or model provider does not expose to observation.

Boundary-setting remains contextual.

Applicable criteria, values, laws and acceptable risk depend on the domain and must be declared by competent parties rather than invented by Aurora.

Deterministic work should remain deterministic.

Probabilistic methods are used where their capability adds value. Exact rules, authority conditions and other deterministic controls remain deterministic where appropriate.

Research context

An engineering proposition within a recognised assurance challenge.

UK government guidance describes AI assurance as measuring, evaluating and communicating the trustworthiness of AI systems, while the National Physical Laboratory's Centre for AI Measurement is intended to support scientifically robust technical assurance capabilities and help credible tools progress toward market.

These references establish research and policy context only; no endorsement, validation or participation by the referenced organisations is implied.