The Architecture of Trust: Scaling Inference and High Assurance in the Information Economy

Since the widespread public adoption of AI, much of the conversation has been focussed on model capabilities, bigger models with better benchmarks and new capabilities. However, as these systems begin to obtain true decision-making authority a more nuanced question begins to dominate. It is no longer simply about how smart the model is, but rather about how it performs under stress. Ultimately, this question will be answered at the inference level once the model has been activated and is consistently running under real-world constraints.

Inference is the operational phase of AI, the moment a trained model runs continuously in real environments, making decisions, classifications, and actions. As it evolves, the intelligence it provides will become part of the infrastructure of finance, healthcare, energy and other systemically critical industries. For this to occur, we must be able to ensure that these systems are reliable, efficient, and predictable under real-world pressure.

Compute has increasingly become the engine of the digital economy, shaping how efficiently high-consequence systems run. In healthcare, inference will flag deterioration and tune closed-loop insulin therapies; in energy, it will forecast demand and balance grids in real time; in transport and logistics, it will route fleets, manage congestion, and optimise safety margins; in cybersecurity, it will detect anomalies and coordinate rapid containment. As inference becomes faster and uses less energy, intelligence moves into the operational core of these sectors to shape outcomes in real time. This introduces a new discipline for the industries – energy efficiency will become financial and safety priorities rather than purely technical or economic ones.

The Scaling Inference Lab (part of ARIA’s broader Scaling Compute Programme) exists to specifically catalyse this outcome, with its overall aim to reduce the cost of AI hardware down by >1000x. The most credible way to achieve this efficiency is through a systems-based approach. Creating efficient inference is not simply a chip problem, or a model problem, or a data centre problem, or any other problem in isolation but rather it is a coordination problem across the entire stack. If you optimise one layer without aligning the rest, you simply move the bottleneck. Whereas, a holistic systems-based approach can simultaneously reduce waste, lower energy use, and narrow the range of performance outcomes, making AI behaviour more stable under stress. CommonAI’s collaborative platform has been selected to steward delivery and provide the neutral, real-world testbed needed to prove these gains in practice. This predictability is a necessary condition to unlock one of the primary pillars of Anthemis’ thesis for AI-enabled business models in regulated and systemically-critical industries: “High Assurance.”

We define High Assurance to mean AI systems that are engineered with execution authority designed to operate in high-consequence environments. Not just drafting text or offering suggestions, but approving transactions, dispatching actions and triggering downstream automation. At a high-level, it is AI moving from a co-pilot to a pilot. In such situations, “usually correct” is not sufficient. We need predictable, reliable, explainable and auditable behaviour if we are to trust AI to be unleashed inside finance, healthcare, infrastructure, and other mission-critical industries. At Anthemis, we believe the investment opportunity for such applications is huge. It is where the technology becomes truly enterprise-grade, where defensibility compounds through compliance and integration, and where the value captured can be orders of magnitude larger than the copilots that dominate today’s headlines.

For this to take place, institutions must be able to see how AI decisions are made through clear execution records, showing which model ran, on what data, and under which controls. When infrastructure becomes easier to understand, governance can keep pace with execution, a key requirement for high assurance AI. This is where the true complementarity of these two engineering themes converge. Scaling Inference becomes the operational base, running real-world logic efficiently while pushing energy discipline and price performance, and High Assurance AI becomes the governance layer, adding explainability, verification, and control as native features of the stack.

Once these layers connect, predictability becomes part of the infrastructure itself. Verifiable execution supports stable decision cycles, reliable handoffs across organisational boundaries, and constant monitoring of risk. Software will increasingly operate within critical industries, monitoring and preventing unsafe behaviour within industrial environments, directing energy load across the grid and using inference to assess consent while enabling online transactions.

Across the broader evolution of the information economy, the meeting point of Scaling Inference and High Assurance AI marks a structural shift. Shared execution infrastructure starts to carry more of the operational burden, creating a common base where efficiency, resilience, and trust reinforce each other. Teams build faster on efficient technology, operators gain confidence in automated systems where behaviour is measurable and provable and oversight bodies engage with environments designed for clarity, from traceability and auditable controls that can be inspected end to end.

Scaling Inference raises the floor on efficiency and reliability, while High Assurance AI raises the ceiling on what automation can safely do, whether that’s managing energy flows, coordinating logistics, controlling clinical pathways, or settling value in real time. Trust moves into the infrastructure itself, expressed through the performance, transparency, and stewardship of the systems that increasingly underpin modern life.

share to my network

Story by Sean Park