How it was built and gated¶
The build is review-gated on purpose: the plan at this section's end fixed the subphases and their order before any code, every change lands through a pull request whose checks include the writing rules and the status-truth gates, and AI-USAGE.md keeps the record of what the coding agent got wrong along the way, because that record is the point.
The program-level view across every repository is PIPELINES at the program's home; what follows is this repository's own.
The pipeline, explained¶
Seven workflows run the gates, and the diagram shows where the four that gate merges and releases land their results:
checks runs on every pull request, on the merge to main, and
weekly on a clock. Eight jobs: secrets sweeps the full history with
TruffleHog with verification on, so a found credential is tested
against its provider to learn whether it is live; writing holds
these documents to the writing rules and runs the docs-truth and
digest-parity gates, so a stale status claim or a drifted image pin
blocks the merge; workflows lints and security-audits the workflow
files themselves, because a mistake in the files that gate everything
else is the most expensive kind; links walks every cross-reference
offline, fragments included; floor runs the whole suite on the
oldest supported interpreter, Python 3.11, because the shipped image
runs 3.14 and 3.14 alone forgives annotation patterns older
interpreters refuse (D-053); browser drives the page with a real
browser from a tree pinned apart in requirements-browser.txt, so
the page is proven by use and the browser never enters the image or
the ordinary suite (D-077); application runs the linter, strict
typing, every test under the coverage floor, the mutation check,
the migrations against a real PostgreSQL with drift detection, the
dependency audits, and generates the software bill of materials as
the run's artifact; and container lints the Dockerfile, scans
the pinned base image on the schedule and the built image on every
change, and runs GuardDog over both pinned dependency trees
from its digest-pinned official image, asking the question the
vulnerability audit cannot: whether a package behaves like malware
before any advisory exists (D-052). Every tool the pipeline downloads is fetched from its
canonical release and checksum-verified before it runs, so the
pipeline's own supply chain meets the same bar as the application's.
page runs the page's one lint rule, the one a scanner's rule set
change turned main red on after months of green: a promise nobody
awaits. typescript-eslint reads the plain script through the
TypeScript checker, and the tools install from the lockfile with
integrity hashes and no scripts run. The same rule runs at commit
time, and three more rules of this repository's own, written from
the lessons scanners taught after a push, run with the lint at commit
time and in the application job (D-086); the pipeline's CodeQL
queries also run locally before a push through scripts/scan.sh.
codeql runs deep static analysis over the Python, the page script, and the workflow files, on every pull request, on main, and weekly; its findings land in the repository's code scanning view.
release runs when a version tag is pushed: it rebuilds the artifacts from the tag (the source archive, the sample account at both sizes, the software bill of materials, checksums), attests build provenance for every artifact into the public transparency log, and publishes the release only if the tag's signature verifies.
scorecard runs on main and weekly, rating this repository's own security posture from outside, and publishes the score to the public scorecard service where it can be read without trusting this repository's word; the badge at the top of this document is served live from that service, so the displayed score cannot drift from the published one. The checks scoring zero are structural facts or measurement lag: a single contributor, a repository younger than the window the rater reads, fuzzing not yet adopted, a review requirement newer than most of the history it evaluates, and release signatures the rater looks for as uploaded files rather than in the platform's attestation log where this repository puts them. Its findings deliberately stay out of code scanning: several are recorded accepted risks no pull request can fix, and an alarm that is always red teaches the eye to skip the alarm (D-037).
The weekly clock exists for the scanners whose subject changes while the code does not: a fix shipping for the base image or a new advisory against a pinned dependency is found on schedule instead of waiting to fail whichever pull request comes next (D-043). The base image scan blocks only on that clock and on main; on a pull request it reports, and the image the pull request builds is what blocks, because the build applies Debian's updates and a pull request can fix what it builds but not what upstream has yet to rebuild (D-055).
The eighth job, doctrine, scores this repository against
build-doctrine's six-level
scale with the doctrine's own scorer, checked out at a pinned commit, and
fails when any applicable rule is absent, so the presence baseline, the
pins, and the counted figures are held by the same tool that publishes
the score.
Ten of these checks are required by the branch ruleset, so there is no path to main around them; the ruleset also requires pull requests and plain merge commits and blocks force pushes and deletion. Each tool was vetted at adoption and recorded as a decision, and two of them found real defects here before they were merged.
release also publishes the container image to this repository's
package registry, ghcr.io/manifest-identity/manifest-identity, tagged with the
version, so a consumer can pull instead of build, and attests the
image digest the same way it attests every artifact. A pulled image
verifies with gh attestation verify oci://ghcr.io/manifest-identity/manifest-identity:<tag> -R manifest-identity/manifest-identity.
attest-release is started by hand with a tag name and attests a release that was cut before the release workflow gained its attestation step: it downloads the assets exactly as published, attests those bytes, and attaches the bundle beside them. The provenance says what it is, an attestation of the published files dated the day it ran, not a claim about the original build.
fuzz runs ClusterFuzzLite against the two AWS parsers, the first
that read files from other systems, on every pull request touching them and weekly.
The harnesses under fuzz/ swallow the named refusal each parser
promises for bad input and let anything else escape, so a crash it
finds is an input that reached an exception nobody wrote.
docs publishes this documentation as a site with side navigation
and search at https://manifest-identity.github.io/manifest-identity/, generated at
build time from this README and the root documents by
scripts/build_docs.py, so the site has no source of its own to
drift, and rendered in strict mode so a broken link or anchor fails
the build rather than reaching a reader. A test holds the generator
to the README's section count and resolves every anchor across the
split.
The actions the workflows stand on¶
The workflows themselves run third-party code: twelve published actions,
each pinned to a full commit hash, with the version tag kept as a
comment beside it in the workflow. The hash is what runs; a tag can be
moved to different code, a hash cannot. This table names what runs and
why; the pins live in the workflow files alone, because a pin written
twice is a pin an update tool can only half move (D-061). A gate
(scripts/check_actions_inventory.py) refuses any use that is not
pinned to a full commit hash, and refuses any action or image the table
does not name.
| Action | Where it runs | What it does |
|---|---|---|
actions/setup-node |
page | Installs the pinned Node the page lint runs on; the lint itself installs from the lockfile with integrity hashes |
actions/checkout |
every job of six workflows; the fuzz workflow's actions fetch for themselves | Fetches the repository; credentials are not persisted, so no token outlives the step |
actions/upload-artifact |
checks, the application job | Carries the software bill of materials out of the run |
github/codeql-action/init |
codeql | Sets up the analysis engine for the Python and the workflow files |
github/codeql-action/analyze |
codeql | Runs the queries; findings land in code scanning |
ossf/scorecard-action |
scorecard | Rates the repository's posture and publishes the score off-repository |
actions/attest-build-provenance |
release, attest-release | Attests each artifact's build provenance, and the container image's digest, into the transparency log |
google/clusterfuzzlite/actions/build_fuzzers |
fuzz | Builds the harnesses under fuzz/ with AddressSanitizer from the digest-pinned fuzzing base image |
google/clusterfuzzlite/actions/run_fuzzers |
fuzz | Runs each harness for a bounded time against inputs derived from the change; a crash fails the check |
codecov/codecov-action |
checks, the application job | Publishes the coverage report through the workflow's identity token, no stored secret, so the coverage figure is measured and shown by an outside service |
SonarSource/sonarqube-scan-action |
checks, the application job, when the token is present | Runs SonarCloud's analysis on the same commit the other gates judged, importing the coverage report |
actions/upload-pages-artifact |
docs | Packages the rendered site for Pages |
actions/deploy-pages |
docs | Publishes the packaged site through the workflow's identity token |
One tool runs as a container image rather than an action, and it is held to the same table discipline: the inventory gate requires every image a workflow step runs to be named here.
| Image | Where it runs | What it does |
|---|---|---|
ghcr.io/datadog/guarddog |
container | Scans both pinned dependency trees for malware shapes (D-052); its digest is pinned in the workflow |
Everything else the pipeline runs is downloaded by hand in the workflow steps, fetched from its canonical release and checksum-verified before it executes.
Dependencies are the part of the codebase nobody here wrote, so each one was checked against its canonical source before adoption, and installs are hash-pinned: a substituted artifact fails to install instead of running.
| Package | Canonical source | Role |
|---|---|---|
| fastapi | github.com/fastapi/fastapi | web framework; typed validation as the default path |
| uvicorn | github.com/Kludex/uvicorn | application server |
| SQLAlchemy | sqlalchemy.org | the ORM; parameterization removes injection as a class |
| psycopg | github.com/psycopg/psycopg | PostgreSQL driver |
| alembic | github.com/sqlalchemy/alembic | schema migrations from the first table |
| bcrypt | github.com/pyca/bcrypt | password hashing, used directly, maintained by the Python Cryptographic Authority |
| pydantic-settings | github.com/pydantic/pydantic-settings | fail-fast configuration |
| python-multipart | github.com/Kludex/python-multipart | upload parsing for the two import routes |
| jinja2 | github.com/pallets/jinja | the report engine, escaping by default (D-040) |
The development tree (pytest, Hypothesis, ruff, mypy, pip-audit,
pytest-cov, pip-tools) is verified the same way and isolated in its
own hash-pinned file. Semgrep sits in a third tree of its own,
because it declares a PyJWT line that carries published advisories;
scripts/compile_scan.py compiles that tree and then overrides the
one pin to the fixed release, hashes read from the index, and the
tree installs complete with no resolver, so every audit reads a tree
without a known vulnerability and without an exception (D-086).
How the agent is governed¶
Most of this repository's code was written by an AI coding agent, and that arrangement runs under the same principle as everything else here: mechanisms, not intentions.
- The agent works to written standards. AGENTS.md is the doctrine it is held to from the first commit, and the writing rules, the truth gates, and the secret scan run at commit time on the agent's output exactly as they would on anyone's.
- The agent cannot land anything alone. Main refuses direct pushes; every change travels a branch and a pull request opened under the agent's own identity (D-045), so the author of record and the human who approves are different parties; twelve required checks and a required approving review must pass; and the merge is a human act. Phase and subphase transitions are likewise human declarations, never the agent's.
- The work is attributed, and the attribution states its own limits. Every co-authored commit carries the standard Co-authored-by trailer, naming the assisting system through its shared attribution account. The exact model behind any single commit is not knowable from inside the session, so no trailer claims one; AI-USAGE.md records the incident that taught this. The agent's commits are unsigned; what carries provenance is the reviewed merge, the maintainer's signed tag that starts a release, and the attestation on every artifact (D-087).
- The failures are the record. AI-USAGE.md keeps what the agent got wrong, what caught it, and what each catch changed, because the interesting output of an AI-assisted build is exactly that list; several of this repository's gates exist because an entry there demanded them.
- The human's limits are recorded too. One person reviews this work, and the self-assessment in SECURITY.md states what that costs rather than hiding it.
The software bill of materials, the machine-readable inventory of the
full dependency tree, is regenerated by every pipeline run from the
hash-pinned requirements and published as the sbom artifact on the
latest checks run, rather than committed, because an inventory
committed once and forgotten drifts into a stale claim the moment a
pin moves; the generated one cannot disagree with the tree that was
actually installed. GitHub's dependency graph offers its own export
built from the same pinned file.
Releases carry the same discipline outward (D-050). The version scheme reads from the roadmap: v0.N means the work through phase N is complete. Each release starts from a signed tag, carries a source archive, the sample account at both sizes (the curated set as committed, and a scaled set of a thousand bulk identities per generation for load work), the bill of materials, and checksums, and every artifact has a build provenance attestation verifiable against the platform's transparency log rather than this repository's word; the attestation bundle also ships as a release asset, so the same proof reads offline and by raters that only look at assets:
gh attestation verify sbom-v0.2.0.json -R manifest-identity/manifest-identity
The plan, fixed before code¶
The build was divided into twelve ordered subphases, planned in full in advance and built one at a time. A subphase is built in small commits on its own branch and then stops: a human reads the diff, runs the demo, and reads the tests, and only after that review is the pull request merged with the required checks green, so the merge itself is the public record of the review. While author and reviewer were the same account, no approval was required on the pull request, because a self-approval would have been theater; since the agent gained its own identity, one approving human review is required and is real (D-045), because the author of record and the approver are different actors. There is no testing phase at the end, because every subphase ships its own tests, and no hardening phase in substance, because each control arrives with the thing it protects; the final subphase is proof, not retrofit.
- Foundation. Hash-pinned dependencies checked against canonical sources, the software bill of materials, automated update review, a digest-pinned container image, fail-fast configuration, migrations from the first table, allowlist logging, health.
- Operators. Sign-in with a timing-equal path for unknown names, revocable sessions, the three roles checked per route, the audit spine writing in the same transaction as every action.
- Ingestion one. The credential report parser: bounded, in memory, verified against its own claims, append-only, identities keyed by the provider's immutable identifier, with its property-based fuzz suite.
- Ingestion two. The authorization details parser: roles, trust policies, groups as privilege sources, memberships, policy documents, tags, and recreated-name detection.
- Derivation and credential findings. State from history at read time, and the credential-hygiene findings with their tiers and the minimum observation age.
- Privilege findings. Admin equivalence by capability, escalation paths, external trust exposure, ownership and group findings, membership drift, privilege attributed to its source.
- Sample data. The synthetic generator producing both file formats across three import generations and every archetype the rules need; moved up from eleventh with the reason recorded in D-034, because every subphase since the first parser had needed demo input made by hand, and hand-made input was wrong three times.
- Inventory and frontend. The lists, the detail view with its observation timeline, the dashboard, the as-of banner, and the single page that renders every value as text.
- Governance records. Owner, purpose, flag, and attestation on identities and groups, attributed, audited, clearable.
- Review campaigns. Scoped, deadlined review cycles with per-item dispositions including insufficient evidence, recommendations with their reasons, the change-since-last- certification view, and no bulk certification by design.
- Reports and exports. Escaped CSV and JSON, the self-contained risk report, and the per-campaign evidence export with its population statement.
- Proof, and the stranger drill. Container hardening verified by command, the mutation check with coverage measured to inform it, the external checklist audits, figures verified against the running system, the fresh-clone run on a machine with nothing but Docker, and the documents re-read and shortened.
Those twelve built the observed half and were tagged v0.2.0 with Phase 2. The authorized half then followed the same discipline in sixteen more subphases, numbered 1.1 to 1.16, planned before the first was started and built in batches of two, each batch one pull request with its runtime proof: the scope tree and scoped administration (1.1), the authorization record (1.2), the form and the file door (1.3), authorize from observed (1.4), the delta (1.5), paths and relationships (1.6), role definitions as versioned observations (1.7), campaigns driven by the delta and by expiry (1.8), alerts and their records (1.9), the read API and the change feed (1.10), the table door with the source selector (1.11), GitHub as the second provider (1.12), the page (1.13), and every provider's file (1.14), Active Directory natively through two doors (1.15), and the populated record a fresh clone gets from one script (1.16). Their decisions run from D-071 onward, and the sixteen shipped together as v0.5.0 (D-087).
The order had reasons. Identity before data, because every later route needs the role checks. Parsers before the engine, because reading the data before designing against it is the deepest lesson this project inherits. Credential findings before privilege findings, because the second carries the judgment and gets the hardest review. The frontend in the middle, so every later subphase demonstrates with clicks. Governance before campaigns, because the noun precedes the workflow. Sample data before the frontend, so demonstrations run against realistic data instead of input typed by hand. The plan bound the order, not the learning: a discovery mid-build became a decision, an amendment, or a backlog entry, visibly, so the difference between the plan as written and the build as it happened stays readable in DECISIONS.md.
Phase 2, local Kubernetes. The image orchestrated on kind with Calico, so network policies are enforced rather than silently ignored; role-based access control, pod security standards, admission control.
Phase 3, cloud enclave as code. The AWS environment as code, split into persistent foundation and ephemeral workload. The organization trail lands here, which is also when enrichment deepens: creator attribution and usage beyond the provider's 90 day window become possible only with logs to hold them.
Phase 4, managed Kubernetes. The image promoted by digest into the enclave; the orchestration questions were already answered locally.
Phase 5, security-gated pipeline. The running gates consolidated, the gaps closed, and the set proven by introducing a flaw deliberately and confirming the pipeline stops it.
Phase 6, runtime security and alerting. Detection on the audit events the threat model names, alerts on new high-risk identities, and the first-hour response procedure written and exercised once.
Phase 7, human-triggered remediation. The trust step change: a tightly scoped action credential, deactivate and restore behind step-up authentication, each action shown as a policy diff before it happens and verified against the provider afterward, because clicked is not revoked until the provider says so. Report-only quarantine and review windows arrive here, and this phase requires its own threat model revision before any code, because write access changes what the tool is.
Beyond the phases, in order: expected-profile checks, where a known vendor integration holding exactly its documented permissions is furniture and the same integration holding more is a finding; temporary approved re-elevation, where someone else approves and the clock does the offboarding; and the live pull for each provider, as adapters behind the same append-only ingestion, once the file door has proven the model for it.
Each phase ends in a state that runs and demonstrates on its own, with the diagrams updated, the decisions recorded, and the documents re-read and shortened.