Limits

The strongest case against this work, stated plainly. A system built around refusing overclaims should refuse its own.

Per-object claims, verification dates and promotion conditions are in the status ledger. Where this page and the ledger disagree, the ledger is newer.

What is qualified, and what is not

Alpha 1, which ended with the released 0.1.0-alpha.6, proved the semantics: a real composed occurrence, its authority, custody and currentness boundaries, recovery, receipts and refusals. Alpha 2, accepted on 2026-10-02, shows that a packaged set of versions can be installed from published artifacts onto fresh machines by someone outside the development environment and run one governed effect with independent postconditions. Beta will mean an operator can actually live in it, with an operator-facing workflow, coherent status and refusal surfaces, and real setup, upgrade, recovery and day-to-day operation; that is in preparation, and nothing here claims it yet. The 2026-09-25 Alpha 2 release candidate was qualified but never accepted; the accepted Alpha 2 showing supersedes it.

What has been shown only once

What is private

What is not production-ready

Everything. Constellation is alpha, solo development, and nothing here is deployed at scale; this page will say so until that changes. Each component states its own maturity on the components page, and none claims production deployment.

The contracts describe how responsibilities are divided; deployment must actually preserve that division. Service identities, filesystem access, executor configuration, database access, operator privileges, and access to a container daemon or its socket can bypass a software workflow if they are too broad. Constellation does not make those deployment privileges safe merely by recording a well-formed plan.

The local trust boundary

On a single host, the governance protects against the worker and the agent, not against whoever controls the operator’s own host account. In Alpha 2 that account can run the packaged executor directly, without an authorization or custody, and it also holds the authorizing key and the working directory; the Alpha 2 bundle generates issuer custody inside each virtual machine, which does not close this. Closing that gap needs component changes (a separate account for effects, or an executor that checks a signed token) and is an open item for a beta. The ledger row records it.

The paid review route also sends one non-generating warm-up request besides the review itself, from the pinned App Server. Whether the provider bills that warm-up is not verified; it is another beta item.

Where evidence does not generalize

What this is not

Not an all-or-nothing AI platform. Not a generic agent framework. Not a claim that every artifact composes with every other. Components remain independently useful; supported compositions are named, pinned and qualified one at a time.

What the proofs and receipts don't prove

The Lean theorems prove class boundaries: that a given refusal kind is required by the custody discipline's own definitions, and, in the design-constraints family, that a rounded score, a perturbation distance, a spend limit or a provenance chain does not by itself license the inference a design might want to draw from it. They do not prove that any deployed system is safe, that the classic Python implementation is bug-free, or that a particular production incident was machine-checked. Receipts prove instance facts: what was observed, what was decided, under which policy hash. They do not prove intentions, and they cannot make a wrong observation right — only attributable. Component-level formal work is bounded the same way; see how it works.

What content-addressing does and doesn't fix

A receipt id is a hash over the canonical decision inputs, so the same inputs under the same schema always give the same id. That is a real property and it holds. It is also narrower than it sounds: this site pinned dda5a1e5… in prose for months, and when the receipt schema gained fields the id became 3f8b93c1… while the behaviour it attests did not change at all. Determinism was never violated; the claim was simply about a schema version and was written as though it were about the system. Ids belong in artifacts, not in sentences.

What it costs

Custody moves cost from postmortem reconstruction into runtime. Every gate is latency on the action path. Every receipt is storage. Every typed claim is friction at write time — an agent that used to just do the thing now has to propose it, and something has to verify it. The bet is the same one behind TLS and structured logging: pay steadily at runtime instead of catastrophically at incident time. If your incidents are cheap, this trade is bad for you.

What must be trusted

This architecture relocates trust; it does not eliminate it. The trusted computing base is the receipt store, the sealing path, and the clock witness source — small, boring, and enumerable, but real. Trust with a bill of materials. Trustless is not on offer here.

What happens when witnesses are wrong

A witness can be wrong. The system's answer is not prevention — it is attribution: which witness, attesting what, when, with what coverage. False testimony becomes attributable evidence with a return address instead of an anonymous rumor in a log file. That is a real improvement and a real limit: garbage observed is still garbage; it is merely garbage you can recall by origin.

What approval does not mean

Approval is an explicit operator or gate decision over a specific proposal, at an authority-bearing surface, with scope, time, and evidence recorded. It is not the agent saying it was done. It is not a passing test run. It is not a completed rehearsal. It is not a chat message. It is not a proposal being generated, nor a demo succeeding. Approval is a bounded authority-bearing act — if nothing recorded a scoped operator or gate decision, nothing was approved.

What conventional tools already do well

Authorization checks, RBAC, OPA over inputs you already trust, audit logs for low-stakes flows, rate limits, CI gates. If your inputs are attested by construction or your blast radius is small, the conventional stack is cheaper and you should use it.

If you only need an authorization check, use an authorization check.

Where it's the wrong tool

Current maturity

Alpha. Solo development. A research lab with working instruments and proof artifacts, not a product seeking deployments — lab notebook, not an oracle.

Naming is not stability. Several components describe themselves as an “office” of a constellation. That word states jurisdiction — what a component may and may not decide — and says nothing about API stability, packaging, or production-readiness.

The July 2026 exhibits, as then stated

Before the integration releases, two claims sat at the top of the pile, and they were different kinds of claim. The runnable one: a refusal demo that reproduces in about five minutes and whose exit code fails loudly if it passes for the wrong reason — re-verified 2026-07-29 from an editable install, which is not the same as a cold clone. The reported one: four separately-owned offices carrying one bounded change end to end, dated 2026-07-26, whose receipts you cannot see. Neither is a deployment. Two of the four offices stated in their own repositories that they were not production deployable or not operator-ready. Both exhibits are now in the archive.

Live cage execution in that classic lineage is unarmed by construction — the synthetic substrate (ORIGIN_SYNTHETIC) cannot confer operational effect; that refusal is structural, not a configuration flag.