The field is racing to make a single AI model trustworthy enough to act on its own. Ontinuity takes the opposite bet: that no model should ever be the thing you trust.
A team of AI seats, held to a contract fixed before the work begins, by a gate that won't let anything ship until a second, independent seat signs off — with every step folded into an auditable record. Reliable output from interchangeable, often inexpensive, models. Because the reliability was never in the models.
The field is racing to make a single AI model capable and aligned enough to be trusted to act on its own — and managing what it can't guarantee with guardrails and human oversight. Ontinuity takes the opposite bet: that no single model should ever be the thing you trust, and that reliability is therefore a property of the structure around the models, not of the models themselves.
Bigger models, better alignment, longer context. When the agent is still unreliable, add a guardrail model or a human in the loop. Every safety mechanism lives inside the same system being checked — so a confidently wrong agent sails straight through.
Assume no model is trustworthy enough to be a decision gate. Build a team of seats that adversarially review each other, hold them to a contract fixed before the work, and forbid any seat from shipping its own output. Trust is produced by the structure, not placed in a part.
Because reliability comes from the harness, the model in the seat can be swapped, downgraded, or made inexpensive without losing it. Reliable output from interchangeable, often open-weight, models — because the reliability was never in the models.
The core mechanism: a contract is fixed before a session runs, and a gate refuses to let the session close until the output matches it — indifferent to which model produced the work. This is measured, not asserted.
The working seat held by five different models — most often an inexpensive open-weight one — with per-cycle measurement and a randomized-signal control that distinguishes a genuine adversarial response from the mere appearance of one. The reliability holds across the model swaps. That is the whole claim, measured.
The seat that deploys is structurally forbidden from being the seat that authored the deployed bytes. A different seat must review and sign off; a self-deploy writes a violation row. No single judgment — however confident — ships unchecked.
Because every action is logged to an append-only ledger, a seat's claim to have done work can be checked against what was actually done. There is a documented case of the gate catching a model asserting a result the record did not support — and refusing to close until it was grounded.
The same verification-first design runs at two layers: a session engine that produces verified deliverables, and an engineering team of seats that builds and runs that engine. Both governed by the same gate, contract, and grounding discipline — which is why the system can be said to be built the way it works.
Peer model instances that build and adversarially review each other's work. The reviewer is always a different seat than the author, structurally forbidden from rubber-stamping. The unit of cognition is a team with a chain of custody — not one agent with a self-critic whose blind spots are correlated with its own.
A contract fixes the criteria before the work begins, so it can't be quietly relaxed to fit whatever the model produced. The gate refuses to let the session close until the output meets it, and forbids the author of the bytes from being the one who deploys them. Separation of powers, applied to AI.
Every action folds into an append-only record keyed on shared identifiers, so any decision can be walked from conversation to commit to receipt and back. Privileged actions run through a small set of named, bounded, logged operations — never a general shell. A fresh seat boots by reading the record and inherits the whole system.
The same idea applied to a single urgent problem: making regulated data safe to work with by ensuring identity never crosses the line.
A regulated home-care agency can use an AI-assisted scheduling tool on its caregiver and client data without any private information ever leaving its own computer. The data is de-identified at the boundary — on the operator's machine, before a single field crosses — and an independent verifier re-scans the output for any residual identifier before it is trusted. The service operates entirely in token-space — nothing identifiable from that dataset is sent, so there is nothing in it for the service to breach. Built and tested; the first shipped instance of this architecture in a regulated domain.
"Some predict that a future in which AI dominates will be a dark one. The disagreement is this: if the foundational architecture is built correctly now — verifiable, bounded, accountable — the future stops being a wildcard. It becomes a leashed and bounded one."
The wager underneath Ontinuity is that the most important thing to get right about autonomous AI is not capability but accountability — and that the window to set that pattern is open now, while the systems are still small enough to shape. A verification-first architecture is a small demonstration of a large claim: that we can have autonomous systems that are trustworthy not because we believe the model, but because the structure around it makes every action checkable, reversible, and on the record.
The theoretical and empirical foundation — the conceptual ground the working system was built from. Read straight through, or start with whichever thesis pulls you.
A tiered reading guide — foundation, implementation, synthesis, evidence.
These papers document the architecture's foundation. The system has since matured into the gate-and-contract design described above — the current research notes track what runs now.
The engine is real and runs. Access today is through direct engagement, not self-serve — the work is delivered as verified output, not a tool you configure. Hosted self-serve sessions are on the roadmap.
Open-source and forkable. The architecture, the corpus, and the operating record are public. The open questions — long-horizon autonomy, coherence ceilings, per-identity authentication — are real engineering problems.
⌥ GitHub RepositoryFor research collaboration, regulated-data use cases, or serious technical engagement — the de-identification work and the verification engine are available through direct conversation.
◈ Get in touch