Empirical Record · Supplementary
Empirical Session Record
First live four-model Ontinuity session — April 9, 2026
Overview
This is a documented record of the first fully automated four-model Ontinuity session. It constitutes the primary empirical evidence for the claims the companion papers make about the cognitive-ecology architecture, the Teaching Leash safety mechanism, and the behavioral-tendency mapping methodology. It is presented in full to address the distinction between architectural claims and demonstrated results — the session did not produce a published paper; it produced the empirical foundation the published papers rest on.
1Session Configuration
The objective was to develop a framework for evaluating the quality of adversarial challenges — what makes a challenge useful versus spurious, and how to measure that distinction. The four roles ran on different providers: a frontier model as the Researcher (primary work), a different-distribution frontier model as the Challenger (adversarial review), a lightweight model as real-time out-of-band friction scoring, and a local on-device model for session-end Knowtext distillation. The routing client ran on a laptop, checkpointing every three cycles. The session ran seven cycles in roughly twenty-five minutes.
2Session Lifecycle Record
The complete cycle-by-cycle record — each cycle’s status tag, friction signal with the scorer’s stated reason, and the routing action taken. The two override events are the drift catches.
| Cycle | Tag | Signal | Routing action | Friction reason |
|---|---|---|---|---|
| 1 | CONTINUE | 2 — Caution | Cycled automatically | Slight departure — exploring components rather than assessing against challenges |
| 2 | CHECKPOINT | 4 — Override | Override — human input required | Departed from objective — designing test scenarios rather than stress-testing component overlaps |
| 3 | CONTINUE | 1 — Nominal | Cycled automatically | Returned to objective following operator redirection |
| 4 | ALIGNMENT_NEEDED | 2 — Caution | Paused — human judgment required | Genuine framework decision fork — consolidation of two components |
| 5 | CONTINUE | 2 — Caution | Cycled automatically | Some retraction pressure as the session consulted the Challenger before returning |
| 6 | CONTINUE | 4 — Override | Override — human input required | Returned to component relevance rather than redundancy — second drift detection |
| 7 | SESSION_END | 2 — Caution | Terminated — distillation triggered | Clear direction established; adversarial prompt introduces potential drift from goals |
3Key Observations
Friction-signal accuracy. The scorer demonstrated accurate real-time drift detection across all seven cycles. The signal rose from 2 to 4 between cycles 1 and 2 as the Researcher consistently elaborated on component descriptions rather than executing the stress-testing objective. The override at cycle 2 correctly identified genuine drift — designing test scenarios rather than evaluating whether the framework components were the right ones, exactly the describe-versus-execute drift a model reviewing its own output would not catch. After redirection the signal dropped to 1 at cycle 3, then rose to 4 again at cycle 6, catching a recurrence of the same pattern. Two independent overrides, both accurately locating the same behavioral tendency, is meaningful empirical support for the signal’s drift-detection capability.
ALIGNMENT_NEEDED at a genuine fork. At cycle 4 the Researcher issued ALIGNMENT_NEEDED after building a five-point argument for consolidating two components of the framework. This was a genuine architectural decision — accepting it would change the framework being developed — and the Researcher correctly flagged it as requiring human judgment rather than autonomous continuation. The operator routed the consolidation argument to the Challenger for adversarial review before deciding. The tag functioned as designed: the human was engaged precisely when judgment was needed, and not before.
Behavioral tendency observation. Across cycles 1–3 the Researcher-role model exhibited a consistent tendency to describe and elaborate rather than directly stress-test, persisting across two separate operator redirections and recurring at cycle 6. This is recorded as a hypothesis-generating single-session finding — confirmation as a stable tendency requires replication across multiple sessions, operators, and problem types. The methodologically important note: this tendency, a liability in the Researcher role, is predicted to be an asset in the Challenger role, where thorough elaboration of objections is exactly what adversarial review requires. That role-assignment implication was implemented in the session script immediately after, swapping the two models.
4Challenge-Event Record
This session is the first formally documented entry in the behavioral-tendency database the Psychology of AI Data framework is designed to accumulate. The three significant routing events:
| Event | Situation | Cycle | Outcome |
|---|---|---|---|
| Override | Researcher describing components instead of stress-testing them | 2 | Upheld — genuine drift; operator redirection required |
| ALIGNMENT_NEEDED | Consolidation of two components into one | 4 | Routed to Challenger — genuine decision fork requiring human judgment |
| Override | Researcher returned to relevance rather than redundancy evaluation | 6 | Upheld — same drift recurred after two redirections |
5What This Session Demonstrates
The record establishes four claims the companion papers make architecturally. First, the out-of-band friction signal detects genuine drift without contaminating the working dialogue — it rose to 4 twice, both times correctly identifying the same tendency, and the working models were not informed of the value until after their responses were produced. Second, the ALIGNMENT_NEEDED tag functions as a genuine decision-fork detector — the Researcher interrupted autonomous operation not on a timer but because it recognized a substantive decision. Third, the routing client correctly handled all five routing cases across seven cycles. Fourth, a behavioral-tendency observation was generated from live data, and the role-assignment implication it produced became the first entry in the behavioral-tendency database.
The session did not produce a published paper. It produced the empirical foundation for the architectural claims the published papers make. The claims are not speculative — they are grounded in a documented session that ran, produced real events, and generated real data.
One honest detail preserved from the record: at session end, the distillation step was interrupted by an operating-system focus event on the laptop; the session content was preserved in the terminal output and extracted manually. The lifecycle from initialization through SESSION_END and the distillation trigger functioned as specified — the interruption is logged rather than smoothed over, which is the point of an empirical record.
Where this sits in the corpus
This is the receipt — the documented session that the rest of the corpus refers back to when it says “validated in live sessions.” It is supplementary to A New Growth Vector and is the concrete run of the Tetraform protocol. It is deliberately modest about scope: it is one session, the behavioral-tendency finding is explicitly hypothesis-generating and awaits replication, and even the interrupted distillation is logged rather than hidden. The Synthesis carries this same discipline across the whole system — this page is where you can see exactly what “empirically observed” means when the corpus uses the phrase.