Provenance¶
Provenance is the fingerprint stored in a steering file anchor that represents the last-known-good state of the code. When kedge detects that the current code no longer matches the provenance, it reports drift.
Content-addressed provenance (sig:)¶
The default and recommended format. The provenance value is a 16-hex-character AST fingerprint prefixed with sig::
How it's computed¶
- Parse the source file with tree-sitter
- If a symbol is specified, navigate to that subtree
- Walk the AST, feeding each node's kind and leaf text into a SHA-256 hasher (skipping comment nodes)
- Truncate the hash to 16 hex characters
- Prefix with
sig:
Properties¶
Whitespace/comment immune. Only the code structure (AST node types and identifiers) is hashed. Reformatting, adding comments, or changing indentation doesn't change the fingerprint.
Rebase/amend/squash safe. The fingerprint is derived from the code's structure, not from git history. Rewriting history doesn't invalidate provenance.
Symbol-scoped. When tracking AuthService#validateToken, changes to other methods in AuthService.java don't affect the fingerprint.
No git history needed. Detection compares the stored sig: value directly against the current file's fingerprint. No git show or git diff is required.
64 bits of collision resistance¶
The sig: value is 16 hex characters = 64 bits. The goal is drift detection, not cryptographic security. The probability of an accidental collision is negligible for this use case, and the truncation keeps frontmatter readable.
Legacy SHA-based provenance¶
kedge also supports plain git commit hashes as provenance:
How detection works with SHA provenance¶
- Read the file at the provenance commit:
git show <sha>:<path> - Read the file at HEAD:
git show HEAD:<path> - Compute AST fingerprints for both versions
- If the fingerprints differ, the anchor has drifted
The git history must contain the provenance commit, so shallow clones may not work. The detection output includes a full git diff for SHA-based anchors.
When to use SHA provenance¶
SHA provenance is supported for backward compatibility. Prefer sig: provenance for new steering files. To migrate, run kedge link to replace SHA values with sig: fingerprints.
Stamping provenance¶
kedge link¶
Computes the current fingerprint for each anchor and writes it to the steering file:
Use after: - Creating a new steering file (initial stamp) - Intentionally updating documentation to match current code
kedge sync¶
Same as link: recomputes fingerprints and writes them. The distinction is semantic. Use sync to advance provenance when code changed but the docs are still accurate (e.g., internal refactors).
Automatic advancement during kedge update¶
By default, kedge update advances provenance for anchors triaged as no_update by recomputing the fingerprint and writing it to the steering file. kedge logs these in the provenance_advanced field of the RemediationSummary.
Pass --no-stamp to skip this step. The summary still lists which docs had no_update anchors (with anchors_synced: 0), but kedge does not modify the files. Use --no-stamp in CI when docs live in a separate repo. Run kedge sync after agent MRs merge to advance provenance in a single commit.
Provenance lifecycle¶
Local workflow¶
1. Create steering file with empty provenance
│
▼
2. kedge link → stamps sig:abc123...
│
▼
3. Code changes over time
│
▼
4. kedge check → detects drift (sig doesn't match)
│
├─ kedge update → agent fixes docs, kedge re-stamps provenance
│
└─ kedge sync → manually advance provenance (docs still accurate)
│
▼
5. Back to step 3