Limits
The strongest case against this work, stated plainly. A system built around refusing overclaims should refuse its own.
Per-object claims, verification dates and promotion conditions are in the status ledger. Where this page and the ledger disagree, the ledger is newer.
What is qualified, and what is not
Alpha 1, which ended with the released 0.1.0-alpha.6, proved the semantics: a real composed occurrence, its authority, custody and currentness boundaries, recovery, receipts and refusals. Alpha 2, accepted on 2026-10-02, shows that a packaged set of versions can be installed from published artifacts onto fresh machines by someone outside the development environment and run one governed effect with independent postconditions. Beta will mean an operator can actually live in it, with an operator-facing workflow, coherent status and refusal surfaces, and real setup, upgrade, recovery and day-to-day operation; that is in preparation, and nothing here claims it yet. The 2026-09-25 Alpha 2 release candidate was qualified but never accepted; the accepted Alpha 2 showing supersedes it.
- Alpha.6 qualifies one named Linux local-copy composition, not arbitrary executors, providers, or production deployments.
- The qualified release is single-host. The over-the-wire roadmap defines the contracts, failure cases, ownership and evidence required before a network-separated composition is called supported.
- One live Slack delivery was observed on 2026-10-01, with the owner confirming receipt; that is one message, not qualified delivery. Discord and PagerDuty delivery and any human acknowledgment loop remain unverified. Deterministic adapters and local inbox delivery are narrower evidence. See notifications.
- Cross-version restore, migration, and rollback remain unqualified.
- Operational burn-in is a separate, bounded post-release phase; it does not change the immutable alpha.6 result. The first installed observation profile on an operator host ran for about an hour before failing closed; see the record.
- The NQ 0.2.0 package covers one component. It does not qualify any other component or composition.
- ATProto received an exact shared handback. Its application binding, production deployment, migration, notifications, rollback, and legacy-NQ retirement remain ATProto-owned work.
- ECAD/design-flow remains pre-alpha. Operations is the first proving ground, not the permanent boundary of the project.
What has been shown only once
- The alpha.6 governed effect. One exact plan created one small local file, once. The public reproduction verifies retained evidence and refusal behavior without repeating the provider call, authority spend, or effect.
- The Alpha 2 showing (2026-10-02). One fixed governed service start on two fresh virtual machines, once, from a published bundle; the two refused setup attempts and the accepted run are recorded. No stranger has repeated it. See the ledger row.
- The observation profile on one operator host (2026-10-01). One installation, about an hour of real observation, one live notification, one unattended recurrence tick, then fail-closed refusal. See the record.
- The earlier Alpha 2 candidate’s clean install (2026-09-25), history. That candidate bundle, never accepted, has one recorded clean-install run, which passed from packaged artifacts in a fresh virtual machine, and one recorded upgrade-continuity run. In both, a local fixture stood in for the model provider's review, and the test harness, not a person, performed the acceptance. A clean-room newcomer run succeeded. A hostile review found a blocker in the kit pin check; it is fixed, and a later clean install (run-004) refused the review’s attack kits. See the ledger row.
- The four-office governed vertical (2026-07-26). It ran once, against one repository and one effect class. See the archived exhibit.
What is private
- The 2026-09-25 candidate’s packaged artifacts and its clean-install and upgrade-continuity run records are not public, while the accepted Alpha 2 bundle, result record and records archive are public in its release. That candidate’s harness source is public on an unmerged branch; see the ledger row.
- The four-office vertical's identifiers, digests, and verdicts are held in a private campaign record and are not publicly inspectable. What is public is the declaration document, the prior three-office run written up in full, and each office's own statement of its jurisdiction and maturity. A reader can therefore check that the offices agree about their boundaries; a reader cannot independently verify the run. That is the exhibit's stated promotion condition, and until it is met the composition is a dated report, not a reproducible artifact.
- The ABSD operating-system source is not public. The ABSD page reports facts from its own documentation; none of them can be checked from here yet.
What is not production-ready
Everything. Constellation is alpha, solo development, and nothing here is deployed at scale; this page will say so until that changes. Each component states its own maturity on the components page, and none claims production deployment.
The contracts describe how responsibilities are divided; deployment must actually preserve that division. Service identities, filesystem access, executor configuration, database access, operator privileges, and access to a container daemon or its socket can bypass a software workflow if they are too broad. Constellation does not make those deployment privileges safe merely by recording a well-formed plan.
The local trust boundary
On a single host, the governance protects against the worker and the agent, not against whoever controls the operator’s own host account. In Alpha 2 that account can run the packaged executor directly, without an authorization or custody, and it also holds the authorizing key and the working directory; the Alpha 2 bundle generates issuer custody inside each virtual machine, which does not close this. Closing that gap needs component changes (a separate account for effects, or an executor that checks a signed token) and is an open item for a beta. The ledger row records it.
The paid review route also sends one non-generating warm-up request besides the review itself, from the pinned App Server. Whether the provider bills that warm-up is not verified; it is another beta item.
Where evidence does not generalize
- One composition being exercised is evidence about those components and that one effect class. It is not evidence that any other pair of components composes, and there is no universal pipeline behind it.
- A successful recorded run does not qualify another host or authorize a successor. A settled or successful attempt does not establish that the objective behind it was met.
- Specimens run against lab substrate are compatibility evidence, not live testimony about a deployed estate. NQ has never caught a real firewall failing to enforce a real block; it has shown, in a lab, that it would notice.
- Component revisions are pinned per release. Do not infer compatibility between repository heads, between classic and successor components, or between cuts that share a version string.
What this is not
Not an all-or-nothing AI platform. Not a generic agent framework. Not a claim that every artifact composes with every other. Components remain independently useful; supported compositions are named, pinned and qualified one at a time.
What the proofs and receipts don't prove
The Lean theorems prove class boundaries: that a given refusal kind is required by the custody discipline's own definitions, and, in the design-constraints family, that a rounded score, a perturbation distance, a spend limit or a provenance chain does not by itself license the inference a design might want to draw from it. They do not prove that any deployed system is safe, that the classic Python implementation is bug-free, or that a particular production incident was machine-checked. Receipts prove instance facts: what was observed, what was decided, under which policy hash. They do not prove intentions, and they cannot make a wrong observation right — only attributable. Component-level formal work is bounded the same way; see how it works.
What content-addressing does and doesn't fix
A receipt id is a hash over the canonical decision inputs, so the same inputs under the same schema always give the same id. That is a real property and it holds. It is also narrower than it sounds: this site pinned dda5a1e5… in prose for months, and when the receipt schema gained fields the id became 3f8b93c1… while the behaviour it attests did not change at all. Determinism was never violated; the claim was simply about a schema version and was written as though it were about the system. Ids belong in artifacts, not in sentences.
What it costs
Custody moves cost from postmortem reconstruction into runtime. Every gate is latency on the action path. Every receipt is storage. Every typed claim is friction at write time — an agent that used to just do the thing now has to propose it, and something has to verify it. The bet is the same one behind TLS and structured logging: pay steadily at runtime instead of catastrophically at incident time. If your incidents are cheap, this trade is bad for you.
What must be trusted
This architecture relocates trust; it does not eliminate it. The trusted computing base is the receipt store, the sealing path, and the clock witness source — small, boring, and enumerable, but real. Trust with a bill of materials. Trustless is not on offer here.
What happens when witnesses are wrong
A witness can be wrong. The system's answer is not prevention — it is attribution: which witness, attesting what, when, with what coverage. False testimony becomes attributable evidence with a return address instead of an anonymous rumor in a log file. That is a real improvement and a real limit: garbage observed is still garbage; it is merely garbage you can recall by origin.
What approval does not mean
Approval is an explicit operator or gate decision over a specific proposal, at an authority-bearing surface, with scope, time, and evidence recorded. It is not the agent saying it was done. It is not a passing test run. It is not a completed rehearsal. It is not a chat message. It is not a proposal being generated, nor a demo succeeding. Approval is a bounded authority-bearing act — if nothing recorded a scoped operator or gate decision, nothing was approved.
What conventional tools already do well
Authorization checks, RBAC, OPA over inputs you already trust, audit logs for low-stakes flows, rate limits, CI gates. If your inputs are attested by construction or your blast radius is small, the conventional stack is cheaper and you should use it.
If you only need an authorization check, use an authorization check.
Where it's the wrong tool
- Demo or chat applications with no external actions — there is nothing to refuse.
- Hot paths where added microseconds matter more than added evidence.
- Teams that want a turnkey hosted assistant — this is infrastructure you operate, not a service you subscribe to.
- Systems whose failures are cheap to diagnose after the fact — the receipts would be paying for insurance you don't need.
Current maturity
Alpha. Solo development. A research lab with working instruments and proof artifacts, not a product seeking deployments — lab notebook, not an oracle.
Naming is not stability. Several components describe themselves as an “office” of a constellation. That word states jurisdiction — what a component may and may not decide — and says nothing about API stability, packaging, or production-readiness.
The July 2026 exhibits, as then stated
Before the integration releases, two claims sat at the top of the pile, and they were different kinds of claim. The runnable one: a refusal demo that reproduces in about five minutes and whose exit code fails loudly if it passes for the wrong reason — re-verified 2026-07-29 from an editable install, which is not the same as a cold clone. The reported one: four separately-owned offices carrying one bounded change end to end, dated 2026-07-26, whose receipts you cannot see. Neither is a deployment. Two of the four offices stated in their own repositories that they were not production deployable or not operator-ready. Both exhibits are now in the archive.
Live cage execution in that classic lineage is unarmed by construction — the synthetic substrate (ORIGIN_SYNTHETIC) cannot confer operational effect; that refusal is structural, not a configuration flag.