Verification-first AI · Open source · Live

The field is racing to make a single AI model trustworthy enough to act on its own. Ontinuity takes the opposite bet: that no model should ever be the thing you trust.

Reliability belongs in the architecture, not the model.

A team of AI seats, held to a contract fixed before the work begins, by a gate that won't let anything ship until a second, independent seat signs off — with every step folded into an auditable record. Reliable output from interchangeable, often inexpensive, models. Because the reliability was never in the models.

The Bet

Reliability belongs in the architecture, not the model

The field is racing to make a single AI model capable and aligned enough to be trusted to act on its own — and managing what it can't guarantee with guardrails and human oversight. Ontinuity takes the opposite bet: that no single model should ever be the thing you trust, and that reliability is therefore a property of the structure around the models, not of the models themselves.

THE FIELD
Trust the model

Bigger models, better alignment, longer context. When the agent is still unreliable, add a guardrail model or a human in the loop. Every safety mechanism lives inside the same system being checked — so a confidently wrong agent sails straight through.

ONTINUITY
Structure the boundary

Assume no model is trustworthy enough to be a decision gate. Build a team of seats that adversarially review each other, hold them to a contract fixed before the work, and forbid any seat from shipping its own output. Trust is produced by the structure, not placed in a part.

THE RESULT
Reliable from cheap parts

Because reliability comes from the harness, the model in the seat can be swapped, downgraded, or made inexpensive without losing it. Reliable output from interchangeable, often open-weight, models — because the reliability was never in the models.

Live · running now
The Evidence

A gate that won't let bad work close

The core mechanism: a contract is fixed before a session runs, and a gate refuses to let the session close until the output matches it — indifferent to which model produced the work. This is measured, not asserted.

The Gate in Close-Up
  • 3 refusalsone documented session, a three-criterion contract
  • forced a paraphrase up to a real quotation
  • forced an inference up to a cited mechanism
  • forced a description up to an applicable edit
  • thenaccepted the consolidated deliverable
Separation of Duties

The seat that deploys is structurally forbidden from being the seat that authored the deployed bytes. A different seat must review and sign off; a self-deploy writes a violation row. No single judgment — however confident — ships unchecked.

How It Works

One architecture, instantiated twice

The same verification-first design runs at two layers: a session engine that produces verified deliverables, and an engineering team of seats that builds and runs that engine. Both governed by the same gate, contract, and grounding discipline — which is why the system can be said to be built the way it works.

01
The Seats
A team, not an agent

Peer model instances that build and adversarially review each other's work. The reviewer is always a different seat than the author, structurally forbidden from rubber-stamping. The unit of cognition is a team with a chain of custody — not one agent with a self-critic whose blind spots are correlated with its own.

A model reviewing its own output shares the framing that produced the error. A genuinely separate seat does not.
02
The Gate & Contract
Reliability boundary

A contract fixes the criteria before the work begins, so it can't be quietly relaxed to fit whatever the model produced. The gate refuses to let the session close until the output meets it, and forbids the author of the bytes from being the one who deploys them. Separation of powers, applied to AI.

Trust is not placed in any seat. It is produced by forcing two independent seats to agree before anything changes.
03
The Corpus
Auditable memory

Every action folds into an append-only record keyed on shared identifiers, so any decision can be walked from conversation to commit to receipt and back. Privileged actions run through a small set of named, bounded, logged operations — never a general shell. A fresh seat boots by reading the record and inherits the whole system.

Least privilege and a tamper-evident trail — built as the substrate the system runs on, not bolted on for compliance.
The Principle, Shipped

The Boundary De-Identifier

The same idea applied to a single urgent problem: making regulated data safe to work with by ensuring identity never crosses the line.

Don't trust the component. Structure the boundary so the sensitive thing cannot leak.

A regulated home-care agency can use an AI-assisted scheduling tool on its caregiver and client data without any private information ever leaving its own computer. The data is de-identified at the boundary — on the operator's machine, before a single field crosses — and an independent verifier re-scans the output for any residual identifier before it is trusted. The service operates entirely in token-space — nothing identifiable from that dataset is sent, so there is nothing in it for the service to breach. Built and tested; the first shipped instance of this architecture in a regulated domain.

Where This Goes

A bounded future, built on purpose

"Some predict that a future in which AI dominates will be a dark one. The disagreement is this: if the foundational architecture is built correctly now — verifiable, bounded, accountable — the future stops being a wildcard. It becomes a leashed and bounded one."

The wager underneath Ontinuity is that the most important thing to get right about autonomous AI is not capability but accountability — and that the window to set that pattern is open now, while the systems are still small enough to shape. A verification-first architecture is a small demonstration of a large claim: that we can have autonomous systems that are trustworthy not because we believe the model, but because the structure around it makes every action checkable, reversible, and on the record.

Research Corpus

Nine papers, one argument

The theoretical and empirical foundation — the conceptual ground the working system was built from. Read straight through, or start with whichever thesis pulls you.

Read the corpus →

A tiered reading guide — foundation, implementation, synthesis, evidence.

These papers document the architecture's foundation. The system has since matured into the gate-and-contract design described above — the current research notes track what runs now.

Access

Service now, product later

The engine is real and runs. Access today is through direct engagement, not self-serve — the work is delivered as verified output, not a tool you configure. Hosted self-serve sessions are on the roadmap.

The System

Open-source and forkable. The architecture, the corpus, and the operating record are public. The open questions — long-horizon autonomy, coherence ceilings, per-identity authentication — are real engineering problems.

⌥ GitHub Repository
Engagement

For research collaboration, regulated-data use cases, or serious technical engagement — the de-identification work and the verification engine are available through direct conversation.

◈ Get in touch
Contact

For collaboration or serious technical engagement: contact@ontinuity.org