In #06, we showed why the agent is not the system. In #07, we tested one consequence of that architecture: how much reasoning capability a bounded responsibility actually requires when the surrounding system owns more of the operating context and constraints.

In #08, we take the next step: what happens when the domain changes, the reasoning approaches change, and the responsibility boundary itself becomes part of the experiment? We evaluate these principles in commercial lending underwriting—a domain introducing multi-entity financial structures, financial ratio evaluation, document freshness verification, and qualitative credit policy criteria.

Model requirements are a function of system architecture.
The more responsibility the system assumes for context, evidence, policy, authority, state, execution, and verification, the less responsibility we need to concentrate in the reasoning model.

And its reciprocal statement:

System requirements are a function of where model uncertainty is allowed to matter.
Design the operating model by deciding deliberately where responsibility belongs: what the system knows, what evidence can establish, what models may infer, what humans must judge, and what authority is required before any of it becomes an outcome.


Mixed-Method Evidence: Scorecard & Governance Boundary Progression

To test how model capability interacts with system architecture, we evaluated five model configurations across a frozen corpus of 38 synthetic validation and stress-test cases representing commercial lending underwriting scenarios. The models spanned three distinct reasoning approaches: a hosted frontier-model control (GPT-4o-mini), an open-weight small language model baseline (Qwen3-8B across synthetic adaptation levels), and a specialized typed-decision model (TypeSafe Jev).

We evaluated these models across two system governance configurations:

  1. Simulation A (Incomplete Policy Representation): In this configuration, some execution-critical conditions remained dependent on model judgment because qualitative policy conditions were not independently represented outside the reasoning layer.
  2. Simulation B: In Simulation B, execution-critical responsibilities and human escalation paths were evaluated within an independent simulation harness. The simulation environment evaluated model outputs against deterministic policy expectations without referencing target decision answer keys.

Benchmark Performance Scorecard

Model & Configuration Recommendation Accuracy Parser-Flagged Disclosures Sim A False-Proceed Sim B False-Proceed Sim B False-Block Measured Avg Latency Straight-Through Rate (STP)
Hosted Frontier Control
(GPT-4o-mini) Live Inference
84.21% (32/38) 15.79% (6/38)* 0.1316 (5/38) 0.0000 (0/38) 0.0263 (1/38)** 1,545.68 ms 60.53% (23/38)
Specialized Typed-Decision Model
(TypeSafe Jev) Live Inference
71.05% (27/38) 0.00% (0/38)*** 0.1316 (5/38) 0.0000 (0/38) 0.0263 (1/38)** 105.21 ms (14.7× faster) 60.53% (23/38)
Self-Hosted SLM — Level 0 (Scripted Baseline) Scripted Baseline1 84.21% (32/38) 15.79% (6/38) 0.1579 (6/38) 0.0000 (0/38) 0.0263 (1/38)** N/A (Scripted Delay) 60.53% (23/38)
Self-Hosted SLM — Level 1 (Scripted Baseline) Scripted Baseline1 94.74% (36/38) 5.26% (2/38) 0.0526 (2/38) 0.0000 (0/38) 0.0263 (1/38)** N/A (Scripted Delay) 60.53% (23/38)
Self-Hosted SLM — Level 2 (Scripted Baseline) Scripted Baseline1 97.37% (37/38) 2.63% (1/38) 0.0263 (1/38) 0.0000 (0/38) 0.0263 (1/38)** N/A (Scripted Delay) 60.53% (23/38)

*Control parser-flagged disclosures reflect limitation statements generated on abstain decisions. **Simulation B False-Block Rate (1/38 / 2.63%) reflects a single unnecessary escalation of a compliant case to human review caused by an exploratory policy rule revision. ***The specialized typed-decision model achieved a 0.00% parser flag rate because its native typed choice interface eliminates freeform text generation. 1Scripted baselines are controlled synthetic model-output profiles used to evaluate governance behavior; they are not live Qwen inference results.

Scope, Harness & Coverage Limitation Disclosures:

  1. Evaluation Corpus Scope: This benchmark was conducted on a frozen corpus of 38 synthetic validation and stress-test cases designed to evaluate specific underwriting edge cases, rather than a held-out production evaluation.
  2. Simulation B Harness Mechanics: In Simulation B, execution-critical responsibilities and human escalation paths were evaluated within an independent simulation harness. The simulation environment evaluated model outputs against deterministic policy expectations without referencing target decision answer keys.
  3. Execution Recovery Coverage Limitation: Partial execution recovery mechanics were not exercised in this benchmark pre-execution policy gate (0 RECONCILED outcomes).

The Governance Boundary Finding

In Simulation A, incomplete policy coverage allowed model errors to affect outcomes. When qualitative policy conditions were not independently represented outside the reasoning layer, model reasoning failures leaked across the execution boundary, resulting in false-proceed rates ranging from 2.63% up to 15.79%.

In Simulation B, when execution-critical responsibilities were independently represented or explicitly escalated:

Across the evaluated cases, no prohibited execution crossed the complete governance boundary.

[Observed False-Proceed Rate under Complete Governance (Simulation B)]
Self-Hosted SLM L0 Baseline:    [] 0.0000 (100% Contained)
Specialized Typed Model:       [] 0.0000 (100% Contained)
Self-Hosted SLM L1 Baseline:    [] 0.0000 (100% Contained)
Self-Hosted SLM L2 Baseline:    [] 0.0000 (100% Contained)
Hosted Frontier Control:       [] 0.0000 (100% Contained)

ANNOTATION:
"Better reasoning reduced exposure. Better architecture removed model errors from the execution boundary."
          

This establishes an important outcome-level distinction:

Architecture establishes the execution boundary; model capability affects performance within it.


Grounding & Evidence Provenance: Correct Decision ≠ Verified Reasoning

A key finding from auditing the control model's output highlights a critical distinction for production AI governance:

A model can reach the correct decision while producing supporting statements that the system cannot independently treat as verified evidence.

Across six validation cases, the control model correctly determined that the expected action was ABSTAIN due to incomplete or stale evidence, returning a structured decision of ABSTAIN.

However, detailed audit showed that in those cases, the model generated freeform limitation disclosures in its response text (such as "financial statements for FY2025 are incomplete") that were flagged by the scoring parser.

Parser Limitation Disclosure: The scoring parser flagged all non-empty limitation text as non-conforming disclosures. While these statements reflected genuine document limitations present in the source fixture text, returning freeform unconstrained prose alongside decision choices prevents the system from independently establishing evidence provenance against verified source offsets.

Placing unverified explanatory statements in the evidence record would violate evidence provenance. Plausibility does not establish provenance. Even when an LLM's explanation is plausible and its final decision is correct, an enterprise operating system cannot convert unverified model prose into durable evidence. The reasoning layer can recommend an action; the evidence layer must independently verify the facts.


Typed Decisions & Bounded Choice: Specialized Model Architectures

One of the key findings in Experiment #2 was evaluating specialized typed-decision models alongside general-purpose generative LLMs.

Generative LLMs produce freeform text or JSON strings that must be parsed and validated post-hoc. A specialized typed-decision model operates on a native typed choice interface, outputting structured probability distributions over a closed set of domain decisions.

Operating Characteristics

  1. Ultra-Low Latency: The specialized typed-decision model delivered a measured average latency of 105.21 ms—operating 14.7× faster than the hosted general-purpose control (1,545.68 ms).
  2. No Freeform-Prose Failure Surface: The typed-decision interface did not generate freeform explanatory prose, eliminating that particular failure mode from the evaluated task.
  3. Failure Surface Transformation: Different model architectures exhibit materially different capability, latency, and failure characteristics while operating under the same execution boundary. Generative LLMs fail through schema violations and ungrounded prose; typed choice models fail through misclassification among valid options—a failure mode where, in this experiment, the surrounding system intercepted the consequential misclassifications tested.

This demonstrates that different reasoning architectures can occupy different responsibilities inside the same operating model while holding the surrounding execution contract constant.


Workflow Economics: Human Review vs. Inference Burden

A common mistake in enterprise AI evaluation is focusing narrowly on marginal token costs. Experiment #2 examined workflow economics across the evaluation corpus, revealing why token price optimization is secondary to system architecture.

Structural Escalation Ceiling

We differentiate mandatory structural policy escalations (7 of 38 cases / 18.42% required by domain policy rules whenever unrepresented conditions or evidence insufficiencies exist) from model-induced escalations caused by uncertainty or extraction errors.

Under complete governance, all five model configurations achieved an identical Straight-Through Processing (STP) rate of 60.53% (23 of 38 cases permitted straight-through). The frontier control and the typed decision model resulted in identical straight-through rates because a smarter model does not silently acquire more authority. When policy requires human credit review, no model—regardless of confidence or capability—is permitted to bypass that gate.

Workflow Cost Structure

Evaluating workflow economics across the test corpus demonstrates the primary economic leverage point:

  • Human review dominated inference cost by orders of magnitude in our modeled workflow economics.
  • Across the evaluated workload, marginal model inference cost represented less than 0.01% of total outcome cost.

Modeled Economics Qualification: These economic calculations reflect modeled workflow assumptions (including hardcoded API rates, constant SLM hosting costs, and modeled human review times of 0.1 hours per escalation at $150/hr). The conclusion that model inference represents under 0.01% of total outcome cost is strictly dependent on these modeled operational parameters.

Optimizing token prices by fractions of a cent per request yields negligible savings. By contrast, increasing straight-through processing by safely automating a single structural exception improves operational throughput dramatically. Enterprise optimization must focus on straight-through efficiency and exception containment rather than token prices in isolation.


Architectural Governance: Bounded Responsibilities & Adaptation Boundaries

To operationalize these research findings, Agent Atlas applies two core architectural principles to reasoning models:

1. Responsibility Allocation

Atlas can assign different bounded responsibilities to different reasoning approaches while holding the surrounding execution contract constant.

The experiment compared multiple reasoning architectures under the same operating boundary and found materially different capability, latency, and failure characteristics.

2. Preregistered Adaptation Boundaries

When evaluating smaller open-weight models against synthetic error baselines, adaptation progress must be evaluated against operational usefulness rather than safety. Because system safety under complete governance is maintained by system policy gates, model accuracy governs exception volume rather than execution safety.

In our synthetic harness evaluation, candidate model error baselines met our preregistered operational target criteria, so parameter fine-tuning was not triggered.


Authority Invariants & Conclusion

The commercial lending underwriting experiment reinforces the foundational principles of Agent Atlas. As reasoning models become faster, smaller, more specialized, and highly diverse, keeping authority anchored in the operating system becomes essential.

Core Operating Model Invariants

  1. Reasoning can propose. It cannot self-authorize: High model confidence is an inference signal, but it cannot grant or substitute for execution authority.
  2. Architecture establishes the execution boundary; model capability affects performance within it: Across the evaluated cases, no prohibited execution crossed the complete governance boundary. Model capability governs how much work runs straight-through versus how much escalates to human judgment.
  3. Output architecture dictates failure modes: Matching model output architecture to task responsibility eliminates entire classes of system risk.
  4. The cheapest model is not necessarily the cheapest workflow: Workflow economics are driven by straight-through processing fidelity and exception containment, not marginal token prices.

The reasoning layer will continue to evolve rapidly. Models will get faster, cheaper, and more specialized. The system around them must ensure that no matter which model proposes a decision, state remains durable, policy remains explicit, evidence remains verifiable, and authority remains invariant.

Reasoning can propose. It cannot self-authorize.