Skip to content

Architecture

The chain invariant

For each chain id, the audit_trail table holds rows in id ascending order. Every row carries three signature columns:

  • previous_hash: the SHA-256 hash of the previous row in the same chain (empty string at genesis).
  • hash: SHA-256 of this row's canonicalized payload (including previous_hash), publicly verifiable without any secret.
  • hmac: HMAC-SHA-256(hash, secret), where secret is the operator's signing key resolved per row via the secret_id column.

Two invariants link rows together:

row[n].previous_hash = row[n-1].hash          (linking constraint)
row[n].hash          = SHA-256(CanonicalJson::encode(payload columns))
row[n].hmac          = HMAC-SHA-256(hash, secret(row[n].secret_id))

with row[0].previous_hash = '' (genesis row of each chain).

Chain\CanonicalJson::encode() recursively ksort(SORT_STRING)s the payload and encodes it via json_encode(JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE | JSON_THROW_ON_ERROR). The exact byte output is pinned by tests/src/Unit/CanonicalJsonTest.php so any change to the encoding fails CI before it can ship: otherwise every existing row's stored hash would silently become unverifiable.

The columns the canonical payload is built from:

[
  'channel'                 => '…',  // PSR-3 channel
  'chain'                   => '…',  // chain id
  'severity'                =>  …,    // RFC 5424 numeric severity
  'action'                  => '…',  // event verb (create / update / …)
  'resource'                => '…',  // subject identifier
  'context_permanent'       => '…',  // JSON-encoded permanent bucket
  'context_transient_hash'  => '…',  // SHA-256 of the transient bucket
  'created'                 => '…',  // microsecond Unix timestamp
  'secret_id'               =>  …,    // signing key id
  'previous_hash'           => '…',  // chain link
]

That column set is defined once, in Drupal\audit_trail\Chain\ChainPayload. The verifier at every depth, the archiver writing NDJSON and the entry detail page all read it from there, so none of them can drift from the others and report an intact row as tampered.

Two retention tiers split the row's context across separate columns:

  • context_permanent is signed raw in the canonical payload. Kept forever; never purged by retention.
  • context_transient is signed via context_transient_hash only: the column itself is NULLed at retention by the transient-purge stage, and a covering audit_trail_segment row attests the transition (transient_purged_at != 0). The chain still verifies because only the hash is in canonical.

This shape lets GDPR right-to-erasure operate on the transient column without breaking chain verification, while still distinguishing legitimate purge from an attacker NULLing a column to hide content: a bare NULL on a row whose write-time context_transient_hash was non-empty AND that doesn't fall within any covering segment with transient_purged_at != 0 or archived_at != 0 fails the verifier. (Archive ops legitimize the NULL because the archive NDJSON envelope captured the row's transient state at archive time: see NDJSON envelope shape below, so the row's pre-purge bytes live in the file even after the live column is gone.)

Why publicly-verifiable hash plus operator HMAC

Two layers:

  • The hash column is recomputable by anyone with the row data: no secret required. A third-party auditor can walk the chain, recompute every hash, and confirm the linking constraint. This is the public-verifiability anchor.
  • The hmac column proves the row was signed by an operator who held the matching secret bytes. Without the secret the hmac cannot be forged; with it, an attacker can sign rows but still cannot rewrite the existing chain (their forgery would need to also rewrite every downstream previous_hash and recompute every hash and hmac, and the operator could detect that via a TSA-anchored snapshot).

Verification runs at one of three depths, given to verifyChain() as a VerificationDepth: Strict requires every signature to authenticate and treats an unresolvable secret as a finding, Tolerant checks what the resolvable secrets allow and names the rest, and None reads no signature at all. Whatever a walk accepts without checking is listed on the verdict, and authenticated_ok is NULL whenever anything was skipped rather than checked: it reports what the walk did, so a Tolerant walk that met no missing key says TRUE.

Multi-tamper walk

If the verifier sees a row that doesn't validate, it records the broken range and keeps walking rather than stopping at the first recovered range. The verdict's broken_ranges list contains every contiguous failed range in chain order. The message advertises additional ranges so operators don't ack the first range and assume the chain is otherwise clean.

Why per-chain (not global)

A single global chain has appeal: it gives a total ordering over every event in the system. In practice the cost is too high:

  • Every INSERT must read the global head, serialize behind it, and write. Under load (a busy WebDAV save plus a node update hitting at the same time) the lock acquisitions stack up.
  • A single break: a corrupted row, an accidental schema migration that touched the table: invalidates verification for every row downstream, including unrelated subsystems.
  • Cross-channel cause-and-effect is rarely the forensic question (you usually want "what happened to this acte", not "what happened to the system at 14:23").

The module's default routing maps each PSR-3 channel to a chain of the same id. Operators that want to funnel related channels into a single chain register an audit_trail_chain config entity with the target chain id and a channels[] list naming the source channels. The choice is per-channel on the input side and per-chain everywhere else (storage, verification, retention) so chain composition is a deployment decision, not a code one.

Per-chain write serialization

INSERTs into the same chain need to be serialized: otherwise two concurrent writers might both read the same head hash and emit two rows with the same previous_hash, breaking the chain.

The module acquires a Drupal lock named audit_trail.write:<chain_id> for the duration of SELECT head, compute hash + hmac, INSERT. Locks are per-chain, so unrelated chains do not contend.

Schema-level defense-in-depth: a UNIQUE index on (chain, previous_hash) makes chain forks impossible at the database level. Even if the application-level lock fails to serialize, the second row's INSERT fails with a uniqueness violation rather than silently extending the chain in two places.

If the application lock cannot be acquired within 5 seconds, the log entry is dropped from the chain rather than blocking the request indefinitely: the underlying dblog or syslog modules still record the entry on their side. Every drop bumps audit_trail.dropped_under_contention in State, and a non-zero counter surfaces on /admin/reports/status as a warning so the gap is operator-visible.

The lock cannot reach inside somebody else's transaction

The serialization above holds only while the writer can read a head that reflects committed reality. Inside a caller's open transaction it cannot: the lock is released before the row is visible to anyone else, and the head read is bound to the caller's snapshot, so two callers whose transactions overlap read the same head however the lock is scheduled. The unique key does its job, the chain never forks, and the second entry is dropped instead.

An entry written inside an open transaction is therefore not chained there. It is buffered and chained from a post-transaction callback, once the caller's outermost transaction has resolved. The callback fires once, at that outermost resolution and never on a savepoint release, and it reports success for a committed transaction and for one a DDL statement voided, each of which did make the caller's work durable. The whole batch takes one lock hold and one head read, each subsequent entry linking to the hash the previous one just produced, so buffering makes a multi-entry transaction cheaper here rather than dearer.

Where the entry waits is in_transaction_write_mode: the audit_trail_outbox table by default, which is atomic with the change it describes and survives a crash before the flush.

This applies only to writes that ask the writer for the lock. The module's own lifecycle writes pass acquire_lock: FALSE because they hold the chain lock themselves, across their whole transaction, so they are serialized correctly already, and they read the returned row id to bind segment bookkeeping to it. Neither would survive being buffered, and the acquire_lock test is what keeps them out of it.

Storage layout

A single Drupal-managed table audit_trail holds all chains. Each row carries:

Column Notes
id Serial primary key, monotonically increasing
created Microsecond Unix timestamp (string, 16 chars)
channel Originating PSR-3 channel. ASCII only, the requirement core makes of dblog.type with the same column type; on MySQL the column enforces it
chain Chain group id (defaults to channel)
severity RFC 5424 numeric (0=emergency .. 7=debug)
action Event verb (create / update / delete / …); used by the entries listing
resource Subject identifier (e.g. entity:node/42, webdav:files/contract.docx)
context_permanent JSON-encoded permanent-bucket payload; signed raw in hash
context_transient JSON-encoded transient-bucket payload; NULLed at retention by the transient-purge cron run (the cleared range is attested by a transient-purge segment). When the operator opts out of transient-purge, the raw bytes are preserved into the archive NDJSON envelope at archive time and live there until file-purge; the column itself is gone after live-purge regardless. Signed via context_transient_hash (only the hash is in the canonical, never the raw bytes)
context_transient_hash SHA-256 of context_transient at write time; empty string when the bucket had no content
secret_id Integer id of the audit_trail_secret entity that produced hmac
previous_hash hash of the previous row in the same chain (empty at genesis)
hash SHA-256 of the canonical payload: publicly verifiable
hmac HMAC-SHA-256(hash, secret(secret_id)): operator-verifiable

A compound (chain, id) index covers the hot read path (fetching a chain's head for the next INSERT, walking a chain for verification). A UNIQUE index on (chain, previous_hash) enforces the no-fork invariant at the schema level. Secondary indexes on channel, created, action, resource, severity and correlation_id back the admin entries page and scripted queries alike.

Which of the page's twelve filters an index can actually serve is worth spelling out, because they are not all matched the same way:

  • Chain, action and correlation ID are equality matches, and each has an index that answers them. So is the created range behind the From and To fields.
  • Severity is severity <= n on an indexed column, so whether the planner takes the index comes down to how much of the table the bound admits. A permissive bound reads as a scan.
  • Channel and resource are matched with a leading wildcard (LIKE '%needle%'), which no B-tree can serve. Their indexes are there for callers that match a prefix instead, such as resource LIKE 'entity:node/%' from a script.
  • User, message, client IP and request URI read JSON out of context_transient, which no index covers at all.

Chain, action, correlation ID and a date range are therefore the ones to reach for first on a large table. Combine one of them with whichever of the rest you actually need, and the scan is bounded by the indexed one.

audit_trail_segment.lifecycle_hmac default value

The lifecycle_hmac column defaults to '' (empty string) at the schema level. This is defensive: if a code path ever inserts an archive bookkeeping row without computing the HMAC the column ends up non-NULL but empty, and the verifier treats empty lifecycle_hmac as an invalid signature (the HMAC of any byte string under any secret is 64 hex characters, never the empty string).

In practice every code path that creates an archive record also computes the HMAC inline, so the default is unreachable. It stays in the schema as belt-and-suspenders insurance.

Write-path index trade-off

audit_trail carries seven secondary indexes plus the (chain, previous_hash) UNIQUE key. On high-write workloads (e.g. 1 000 rows/second sustained into a single chain) the bottleneck is index maintenance: each INSERT updates eight B-trees. The compound (chain, id) is one of the seven and the one every write certainly touches, so leaving it out of the count understated the cost this section exists to quantify.

For 1.0 the index set is what every operator wants by default (filter by channel, severity, action, created, chain). Sites that profile a hot write path and confirm specific filters are never used can DROP those indexes via a custom hook_update_N; the chain semantics don't depend on the secondary indexes, only on (chain, previous_hash) (UNIQUE) and (chain, id) (compound). A future settings flag for the secondary index set is on the roadmap if real operators ask for it.

id column width

The id column is declared serial at big size, which on MySQL / MariaDB resolves to BIGINT UNSIGNED: maximum 18,446,744,073,709,551,615 rows. At a million rows a day that is about 50 billion years, so no deployment reaches it.

The width is deliberate rather than defensive. One sequence is shared by every chain, and a value is never reused: auto-archive and live-purge move old rows out of the live table and reclaim its space, but not its numbers. So the ceiling counts every row the site has ever written, not the rows it currently holds. A 32-bit column would have run out after 4,294,967,295 of them, which is about twelve years at a million rows a day, and widening it at that point means rebuilding a table with billions of rows in it.

Every column that stores an audit_trail.id value is as wide as the column it references: audit_trail_checkpoint's last_id, both range bounds on audit_trail_segment and audit_trail_acknowledgment, and each of the four *_event_id attestation references on a segment. A narrower one would fail before the id itself did.

It buys nothing cryptographic: id is not among the columns ChainPayload signs, nor among the values a segment's archive_hmac covers, and an archive records the ids it holds as values. Nothing that has to keep verifying depends on how wide the column is.

What the width costs

Index space, and almost nothing else. InnoDB appends the primary key to every secondary index entry, so the four extra bytes are paid once in the row and again in each of the eight index structures. Measured on MariaDB over 200,000 rows carrying this module's own average payload (533 bytes permanent, 298 transient):

32-bit 64-bit delta
data (primary key leaf) 260.77 MB 260.83 MB +0.0%
all indexes 60.33 MB 70.44 MB +16.8%
whole table 321.10 MB 331.27 MB +3.2%

Four bytes on a row already carrying about 1.3 KB disappears into existing page slack, which is why the data side does not move. A secondary index entry is small enough that the same four bytes is a third of it: on the severity index an entry goes from 13.2 to 18.4 bytes, and that index alone grows 40%.

Two of the eight are load-bearing, and the write-path index trade-off above says which: (chain, previous_hash) and (chain, id). Keeping only those, the same measurement reads +10.7% on indexes and +1.1% on the table, because 69% of the increment sits on the six indexes a site is free to drop when it does not filter on them. correlation_id is the first to consider: it grew 2.00 MB, more than the fork guard did on a structure four times its size, and the schema describes it as observability metadata rather than part of the integrity contract.

No timing cost was measurable in either direction. An index-only range scan over the same data measured 9.2 ms against 9.1 ms per query, and bulk insert throughput was indistinguishable across three alternating rounds.

Retention is the lever, not the schema

Index size is linear in the number of live rows, and the number of live rows is what archive_after and live_purge_after decide. So the way to keep the database at a reasonable size is the one this module already ships: archive, then purge the live rows an archive holds. Over a live table holding ninety days (live_purge_after: P3M):

events/day live rows all indexes required only
1 000 90 000 27 MB to 32 MB 13 MB to 15 MB
100 000 9 000 000 2.7 GB to 3.1 GB 1.3 GB to 1.4 GB
1 000 000 90 000 000 26.5 GB to 31.0 GB 12.8 GB to 14.2 GB

A site that finds its index tight shortens the window: moving live_purge_after from three months to about eleven weeks gives back the whole 16.8%, costs no schema change and no downtime, and loses nothing, because live-purge only removes rows an archive already holds. The stages and their thresholds are described under Retention lifecycle below.

Which is also why the width is worth paying for. Purging bounds the space; it cannot bound the numbers, because a purged id is never reissued. Retention answers how large the database gets, and the column width answers how long the chain can go on being written to. They are different questions, and only one of them has a configuration setting.

A second table, audit_trail_checkpoint, records verification checkpoints: (chain, last_id, last_hash, created, hmac, secret_id) rows minted each time the verifier walks a chain cleanly to its head. The next incremental walk starts after last_id and uses last_hash as its expected_previous, bounding verification cost to events accumulated since the last checkpoint. The checkpoint's own HMAC is recomputed before trust; a forged checkpoint falls back to a full walk from genesis. The checkpoints are an optimization layer for the rows that survive: lose them and a full walk from genesis still proves the integrity of every row still present. What it cannot prove without them is that no rows were removed from the END, because what remains is a consistent shorter chain and the checkpoint is the only record of how far the chain reached. See verification for the operational pattern.

A third table, audit_trail_segment, records the lifecycle state of every contiguous chain id range that has been processed in some way (archived, transient-purged, live- purged, file-purged). One row per range, created lazily on the first lifecycle op that touches the range and surviving the full lifecycle (including file-purge) so the verifier can bridge across long-purged ranges via the row's anchor_before / anchor_after columns. Each lifecycle transition records its timestamp + the audit_trail.id of the chain event that attested it (segment_archived, segment_transient_purged, segment_live_purged, segment_file_purged); the chain event carries the segment id in resource='segment:<N>' and the verifier cross-checks both directions of the mutual reference. Three independent HMACs cover different concerns on the row: hmac (identity, sealed at creation), archive_hmac (archive content, sealed at archive op), lifecycle_hmac (re-signed on every state mutation under the currently-active secret).

A fourth table, audit_trail_acknowledgment, indexes the operator attestations that a specific row range is known to be unverifiable (e.g. signed by a now-deleted secret, or restored from a backup with no chain link). The verifier silently skips acknowledged ranges and reports them in the verdict.

The record is the chain, not the table. Recording, moving and deleting an acknowledgment each append an acknowledgment_* chain event carrying the range, the reason, both anchors and the creation stamp, signed and linked like any other chain row. Replaying those events in id order rebuilds the table exactly, which is what drush audit_trail:reindex-acknowledgments does and what the status report checks on every build.

So the index rows carry no signature of their own. Signing the copy while never reading the original was backwards: the copy was what a single DELETE could revoke without trace, and the tamper-evident original went unread. What is checked now is whether the index still agrees with the chain, which is a question a DELETE, a hand-edited range and a restore from a partial backup all answer loudly.

A live-purge does not evict the index entries covering the range it takes. It used to, because the archive carried a snapshot of each one; with no snapshot, evicting would destroy the only live record of why a range reads broken. What is left is inert while the rows are gone, because its anchors point at rows that are not there, and it covers them again if they are restored.

An acknowledgment eventually outlives its own recording event, which retention archives and purges like any other row. The entry then cannot be rebuilt from the chain, and is not treated as one nobody recorded: its id falls inside a range a signed segment_live_purged event says was taken, and the archive of that range still carries the event for anyone verifying against the file.

The table holds no actor. The acting operator is in the acknowledgment_* event's transient bucket. A uid on the table could not be emptied at all: it would be copied into the archive NDJSON, which no purge stage reaches. The transient bucket is the one tier the retention lifecycle can empty, which is what makes erasing an actor possible at all. It empties on age, over a range of rows, attested by a segment: a retention control rather than a per-subject erasure on request, and one that does nothing until transient_purge_after is set, which it is not by default. The acknowledgments listing links each id to the entries listing filtered on its resource, which is where that history is read.

Retention lifecycle

Five cron stages carry rows through retention, gated by five thresholds clocked from the row's created timestamp, with a sixth stage, coverage, ahead of them all to put rows into segments in the first place:

transient-purge to archive to live-purge to file-purge to compaction
   ↑                  ↑              ↑              ↑             ↑
 transient_purge   archive_after   live_purge    file_purge   compact
 _after (opt-in)   (required)      _after        _after       _after
                                                              (opt-in)

Drupal builds a hook class in order to call it, so a collaborator taken as an ordinary constructor argument is built before the flag that governs it can be read, and the switch stops switching anything off. Every cron hook here therefore takes what it needs to do the work as a closure, and reads its flag first:

  • auto-archive takes the archiver and the chain registry, behind cron_archive.enabled, which ships disabled;
  • auto-verify takes the verifier, behind auto_verify_enabled, which ships enabled;
  • the TSA bridge takes the timestamper, the chain repository and the chain registry, behind audit_trail_tsa.settings' cron_enabled, which ships disabled.

Two of those three ship off, so a tick that builds them is what a site got by installing the module and leaving it alone: 7.2 ms and 3.8 ms of audit graph for the first two, measured on Drupal 11.3, on every tick of a site that wanted neither.

The settings form enforces transient_purge_after < archive_after < live_purge_after < file_purge_after < compact_after at submit time. Transient-purge runs before archive on the same tick because it must: once a row is archived, the NDJSON file has frozen whatever transient state the live row was in. A late transient-purge (after archive) would NULL only the live column while the file still holds the bytes: defeating the operator's data-minimization intent.

Each stage writes a chained event row attesting its transition (segment_transient_purged, segment_archived, segment_live_purged, segment_file_purged). The audit_trail_segment bookkeeping row records the matching *_at timestamp + the chained event's audit_trail.id in the *_event_id column. Both writes happen inside one DB transaction so a process crash between them can't persist a "ghost segment" (state flag set, event id zero).

Two thresholds are optional, and both ship empty. transient_purge_after disables its stage when empty; archives written for that chain carry the raw context_transient bytes into the NDJSON envelope as a sibling of the canonical payload (see below). Operators who DO configure transient-purge can also disable it per chain via the zero-duration sentinel PT0S on the per-chain form. compact_after disables compaction the same way, which costs nothing a verifier can observe: what it folds is the bookkeeping tail behind file-purge.

Segments as the lifecycle unit

Each stage operates on segments: contiguous id ranges recorded in audit_trail_segment, not on individual rows. A segment is just a bookkeeping row pinning [from_id, to_id] plus the lifecycle stamps (archived_at, transient_purged_at, live_purged_at, file_purged_at). Rows themselves stay in audit_trail; the segment table records which ranges have been processed in which way.

A segment is bare when all four stamps are zero (it exists but no lifecycle op has touched the range yet), transient-purged when transient_purged_at != 0, archived when archived_at != 0, and so on. Stamps accrue in the order transient_purged_at to archived_at to live_purged_at to file_purged_at; each lifecycle stage advances segments that match its prerequisite stamp pattern.

Chain writes are tied to segment existence as well as to lifecycle transitions. Creating a segment writes a segment_created event carrying the segment's whole identity envelope, and the audit_trail.id that event lands on becomes the segment's id: the record and the index entry are one number. Each later stage (transient-purge, archive, …) emits its own segment_<verb> event and stamps the matching *_event_id column, and carries the identity envelope too, because segment events age out under retention like any other row and whichever survives longest is the one a rebuild reads.

audit_trail_segment is therefore an index over what the chain attests, the same relationship audit_trail_acknowledgment has. Two consequences follow.

A segment row that goes missing is detectable: its creation event names a row that is not there. Two further signals cover segments minted before creation was recorded, and damage the chain cannot speak about at all: neighboring segments meet exactly, the earlier one's anchor_after being the later one's anchor_before, so SegmentIndex::findAnchorBreaks() reports a segment removed from between two that remain; and coverage mints per run of chain-linked rows, so a hole left by a missing segment is never sealed over by a segment minted straight across it.

And a segment row cannot be removed to suit another operation. Compaction is the one path that removes rows, and it attests what it absorbed in compacted_from; the verifier accepts a reference to an absorbed segment on that evidence. There is no "delete this segment" operation: the retention policy advances segments, and an operator changes the policy rather than a row.

The cron pipeline shares one mental model:

  1. Cover orphan rows (rows past the earliest-needed retention cutoff not yet inside any segment). The coverage stage groups them into closed granularity buckets and mints one bare segment per uncovered gap via ChainArchiver::ensureSegmentCoverage(). Coverage is the SOLE minter of bare segments in the cron pipeline.
  2. Walk existing segments in the prerequisite lifecycle state past each downstream stage's cutoff, and advance them by calling the matching ChainArchiver method. Each downstream stage (transient-purge, archive, live-purge, file-purge) only walks; none mints.

Operator-driven entry points (drush audit_trail:archive, the segments admin Save action, the tugboat seeder) also call ensureSegmentCoverage() directly when they need bares outside the cron cadence.

Restore is a one-way ratchet, not a lifecycle escape

SegmentRestorer::restoreSegment() brings an archived segment's rows back into audit_trail for incident response or operator inspection. On a successful restore, the segment row's live_purged_at is cleared back to 0 and live_purged_event_id is re-pointed at the freshly-minted segment_restored chain event. lifecycle_hmac is re-signed under the restored event's signing secret. Three things follow:

  1. Retention re-applies. The cron live-purge stage selects segments via live_purged_at = 0 AND to_created < cutoff; restored segments match again, so the next eligible tick re-DELETEs the rows and stamps a fresh segment_live_purged event. Restored rows can't dodge the live_purge_after ceiling.
  2. Operators inspecting restored rows must pause cron (or temporarily lift live_purge_after). The unpaused-cron behavior is to re-purge within one tick of the threshold window's cron cadence.
  3. Lifecycle stamps don't progress strictly monotonically anymore. The original transient_purged_at to archived_at to live_purged_at to file_purged_at ordering invariant pins the first time each transition fires; restore + re-purge can produce a chain with multiple segment_live_purged events for the same segment, in id-ascending order, interleaved with a segment_restored event. The segment row carries the LATEST live-purge transition; the chain history carries the full sequence.

The verifier's segment-event cross check is aware of this shape: live_purged_event_id on the segment row is allowed to point at any later same-chain, same-segment event whose action is segment_live_purged or segment_restored. See docs/verification.md for the lookup rule and what tampers it still catches. file_purged_at is left untouched on restore: if the file was already file-purged before restore (only possible via the override_path import path), the cron file-purge stage's own file_purged_at = 0 filter prevents a second file-purge from firing.

Restore ordering: event first, rows second, segment row last

restoreSegment() is a three-step orchestration behind three pre-flight refusals; each step has a specific concurrency contract.

  • Pre-flight: three refusals, all of them before anything is written. The segment already in the finalized-restored state (live_purged_at = 0 AND live_purged_event_id != 0), which is a double-call on a segment that restored cleanly. The segment with a segment_restored event already in flight, which is a previous attempt that did not finalize; asked here as well as under the lock in Step 1, so that a half-finished restore is diagnosed by the event it left rather than by the rows that event put back. And the range whose own chain still has rows in it, which is every archived segment until live-purge takes it.

They exist because of what came after them. The PRIMARY KEY rejects the replay either way, but only in Step 2, and Step 1 has written the segment_restored event by then: SegmentIndex::replaySegmentEvents() reads that event as the live-purge pointer move it would have been, so the chain attests a restore that never happened and compareIndexToChain() reports the segment altered from then on. A refusal that arrives after the attestation is not a refusal.

The range is counted on the segment's own chain, because the range is a span and not the set of ids the replay will use, and the two coincide on that chain alone: every archived id in the range was a row of it. Counting the range across all chains would block a restore that is free to proceed, since interleaved chains leave their own rows at ids inside it, untouched by its purge. What that leaves open is another chain's row on one of the archived ids, which needs the id set rather than the range: #3620412.

The importFromFile-then-restore path is preserved: a freshly-imported segment has no live rows of its own chain and live_purged_event_id = 0.

  1. Step 1: emit the segment_restored chain event under a brief chain-write lock. The lock guards the chain anchor + per-chain id sequence for the single INSERT, then releases. Step 0 (the in-flight-event check) runs inside the lock so two concurrent restoreSegment() calls on the same segment can't both pass the check and both emit; the second caller throws on the in-flight detection (the crashed-mid-restore signature that the pre-flight check cannot detect, because the segment row's lifecycle pointer hasn't moved yet).
  2. Step 2: replay archived rows + acks inside a DB transaction with no chain-write lock held. Logger writes on the same chain proceed unblocked; their AUTO_INCREMENT ids never collide with the explicit archived ids (which are always below the live head). A mid-loop throw rolls the transaction back so partial INSERTs don't persist.
  3. Step 3: reacquire the chain-write lock briefly, reload the segment row, sign lifecycle_hmac against the FRESH lifecycle fields, then UPDATE the four restore-mutated columns (live_purged_at = 0, live_purged_event_id = restored event id, lifecycle_secret_id, lifecycle_hmac). The lock serializes Step 3 against purgeSegmentArchiveFile(), the only sibling lifecycle op whose preconditions are satisfied during the Step 1 -> Step 3 window: a concurrent file-purge would otherwise leave lifecycle_hmac signed against stale file_purged_at = 0 values.

The chain attestation is therefore committed BEFORE rows land in the live table. This is the property that lets Step 2 run without holding the chain-write lock: every intermediate state (Step 1 done / Step 2 partially done / Step 2 fully done / Step 3 done) verifies cleanly against the existing verifier rules, because:

  • The verifier does not require "rows present in [from_id, to_id] while segment.live_purged_at != 0" to be absent. It walks visible rows and validates their hash + previous_hash linkage.
  • segment_restored is a non-cross-checkable lifecycle action; the verifier doesn't strict-equal it against any segment-row column.
  • The existing live-purge supersession rule already accepts segment.live_purged_event_id pointing at a later segment_restored event (post-Step-3 state). Pre-Step-3 (segment row still references the old purge event) is the baseline strict-match case.

If Step 1 succeeds but Step 2 or Step 3 fails, the segment is in a partially-applied state: the chain records the restore intent, but the rows either rolled back (Step 2 failure) or landed without the segment-row finalization (Step 3 failure). A retry of restoreSegment() is refused by Step 0: the operator must verify chain integrity, audit_trail row presence in the archived window, and the segment row state, then finalize directly (re-run the missing step manually under SQL). The restoreSegment() API does not offer an automated resume path; the partial-state error message names the emitted event id and the [from_id, to_id] window so the operator knows exactly what to inspect.

Coverage stage: minting bares around granularity buckets

segment_granularity (hour / day / week / month) controls the bare segment width minted by the coverage stage. hour is intended for staging / testing workflows where waiting a full day for the first segment to materialize is impractical; production installs typically run with day or coarser. The coverage stage groups orphan rows into closed buckets: buckets whose end-of-window has passed the coverage cutoff, and mints one bare segment per uncovered gap inside each closed bucket via ensureSegmentCoverage(). Downstream stages (transient- purge, archive, live-purge, file-purge) operate on segments coverage produced; none of them mint.

                            cutoff (earliest needed)
                              ↓
 time  ──── bucket A ──── │ ── bucket B (still open) ── …
        ┌──────────────┐  │   ┌────────────────────┐
 rows   │ R1 R2 R3 R4  │  │   │ R5 R6 R7           │
        └──────────────┘  │   └────────────────────┘
              ↓ closed    │         ↓ open
        mint one bare     │   skip: a future row may
        segment over      │   still land in bucket B
        [R1.id, R4.id]    │   before its window closes
        ↓
  segment A: bare (no lifecycle stamps yet)
        ↓
  later cron tick past transient_purge_after:
  runTransientPurge() advances segment A,
  stamps transient_purged_at and emits chain event.

The coverage cutoff slides:

  • With transient_purge_after set, coverage uses transient_purge_after_us: rows enter segments just before transient-purge needs them.
  • Without transient_purge_after, coverage falls back to archive_after_us: rows enter segments just before archive needs them.

The "closed bucket" requirement matters because a still-open bucket may receive new rows in the next cron tick, which would either land in a row id below the bare's from_id (impossible: ids are monotonic per chain) or require amending an immutable segment (forbidden by the chain attestation). Closed-bucket minting guarantees the bucket boundaries match what archive will later turn into a single NDJSON file.

The segment spine

The rows of a chain link to each other: every row's hash covers the previous row's, so altering one breaks the next. Segments did not link to each other. Each one was attested only by its own segment_created and segment_archived events, and retention purges those like any other row. From that point the segment row rested on its three HMACs, which is the operator's word and unreadable to anybody without the key.

The spine closes that, by doing to segments what the chain already does to rows. It has a name of its own because "chain" is already the most load-bearing word here, and a sentence mentioning both needs to say which one it means.

The construction

Each archived segment produces a DIGEST over its own fixed facts, and the digests link:

D_k = H(chain, from_id, to_id, anchors, file_hash, row_count)
S_k = H(k, S_k-1, D_k)

S_k is stored on the segment row as spine, with k as spine_height, and written into the segment_archived event's permanent bucket. A consolidated row has no archive event of its own, so the head it adopts is written into its segment_compacted event instead, and a replay reads both.

The digest covers immutable facts only. A segment's lifecycle stamps are re-signed on every transition by design, and a transient-purge months later must not move a value already recorded elsewhere, so the purge stamps, the event ids and verification_failed_at_purge stay outside it.

There is no separate digest over the rows, because anchor_after already is one: a row's hash covers its previous_hash, so the hash of the row at to_id transitively covers every row back to anchor_before.

flowchart LR
  subgraph segments["audit_trail_segment"]
    S1["segment A, h=1"] --> S2["segment B, h=2"]
    S2 --> S3["segment C, h=3"]
  end
  subgraph chain["audit_trail"]
    E1["segment_archived, spine = S1"]
    E2["segment_archived, spine = S2"]
    E3["segment_archived, spine = S3"]
  end
  S1 -. "must agree" .-> E1
  S2 -. "must agree" .-> E2
  S3 -. "must agree" .-> E3

What it buys

A new archive re-attests the old ones. Segment 5's facts are folded into the head on segment 40, whose archive event is recent and still within retention. Evidence for an old segment is carried forward by later activity instead of dying with that segment's own events.

One value covers everything below it. Checking forty segments used to need forty surviving events; a head at height 40 commits to all forty.

A local edit becomes a global rewrite. Altering segment 5 changes every head above it, so every later segment row and every later archive event has to change too, and those are chain rows whose HMACs an attacker without the key cannot produce.

That last one is the whole of it, and its limit is the same as everywhere else in this module: someone holding the signing secret can perform the global rewrite. What they cannot do is a small one.

What compaction leaves

Compaction writes one consolidated row over a run of file-purged segments and deletes every segment in the run, which takes with it the fields their digests were derived from. The consolidated row records the height and head of the highest segment it absorbed, and compacted_segment_count marks it as consolidated rather than ordinary while saying how much of the spine it occupies. That count is transitive: a run can contain a consolidated row from an earlier round, and it is counted for the whole span it already stood for rather than as one row, so the number is always how many original segments the row speaks for.

flowchart LR
  subgraph before["before"]
    A["seg A, h=1<br/>spine S1"] --> B["seg B, h=2<br/>spine S2"]
    B --> C["seg C, h=3<br/>spine S3"]
  end
  subgraph after["after folding A and B"]
    X["consolidated, h=2<br/>spine S2, count 2"] --> C2["seg C, h=3<br/>spine S3"]
  end
  before --> after
  C2 --> R["S3 = H(3, S2, D_C)<br/>so an edited S2 fails here"]

The highest rather than the last in the run, because a run is grouped anchor to anchor in RANGE order while the spine runs in ARCHIVE order, and the two part company on a chain that has imported an archive or archived a later range before an earlier one.

A replay ADOPTS that head rather than deriving it, and nothing is lost by doing so. Deriving it would mean folding digests stored on the row itself, and no stored digest can be compared against the segment it describes once that segment is gone. What holds the consolidated row to account is the segment after it, whose own head was folded onto this one, and the lifecycle events above that.

For the same reason a consolidated row never appears in the verdict's covered map: its own range and anchors are values compaction restated, and editing them moves no digest, so the spine cannot speak for them. What accounts for those is the segment_compacted event, through SegmentIndex::compareIndexToChain().

What it is not

It is not a way to prove anything about purged rows. The rows themselves are destroyed, and nothing recovers what a hash was computed from. If archive files are exported to WORM storage, those files are what prove the rows, and the whole chain can be rebuilt from them without this. If they are not exported, the rows are gone and the only thing left to attest is the account of them: the ranges, the counts, the boundaries and the dates.

So what this adds is tamper-evidence for that account, at the cost of three columns and two hashes per archived segment. It is not a retention tier, not a proof of content, and not a defense against an operator who was dishonest when the account was written.

NDJSON envelope shape

Each archive file is a sequence of newline-delimited JSON envelopes. Three envelope types appear, in this order:

{"payload": <canonical>, "transient": <raw|null>, "type": "row"}
…
{"payload": <ack-canonical>, "type": "ack"}
…
{"payload": <archive-record-canonical>, "type": "archive_record"}
  • row envelopes carry one audit row each, ordered by ascending id. payload is the canonical the row's hash was computed over (version, channel, chain, severity, action, resource, context_permanent, context_transient_hash, created, secret_id, previous_hash): byte-identical to what the live verifier reconstructs from the DB row. version travels with the rest, so an archive read back years later says which digest to check it under rather than leaving the reader to assume. transient rides outside payload because the canonical / hash / HMAC layer was sealed at write time and cannot include bytes that the cron transient-purge stage may later drop. The verifier hash-binds the side-channel back to the canonical via the already-signed context_transient_hash: when transient is non-null, sha256(transient) == payload.context_transient_hash must hold, otherwise the file is tamper-evident at that row.

The transient field is optional. An envelope without it is read as "no forensic side channel available"; the verifier falls back to the segment-coverage rule on the live row.

There is no acknowledgment envelope. The file used to snapshot one line per acknowledgment covering the range, which was a copy of a derived index: an acknowledgment is an acknowledgment_recorded chain event, so it is already in the file wherever that event falls inside the archived range, and a restore rebuilds the index from those events rather than from a snapshot it has no way to vouch for.

  • archive_record envelope is the last line and carries the metadata the verifier needs to rebuild the audit_trail_segment row from the file alone after a total- DB-loss scenario: chain, range, anchors, row count, range timestamps, signing secret id, and the segment's identity HMAC.

The spine is deliberately NOT in the footer, and cannot be. A footer is part of the file whose digest the archive envelope signs, so every field in it has to be known before the file is closed, while the spine height is assigned in the phase after that, under the chain-write lock. It is also not the file's to carry: a file is one segment, and the spine is a statement about the whole chain of them. An imported archive therefore joins the spine at the end rather than reclaiming a position recorded in its bytes.

The whole file's SHA-256 (over data lines + footer) is signed into audit_trail_segment.archive_hmac at archive time, so while that row survives, any byte-level edit anywhere in the file is detectable independent of the per-row HMAC layer.

Import is the case where that row does not survive, and there the digest is recomputed from the file being imported, so it attests nothing on its own. What stands in for it is the file's own structure: the footer's two anchors are inside the identity HMAC, the rows each name the hash of the row before them, and import walks that run and refuses a file whose rows do not start at anchor_before, end at anchor_after, and number what the footer claims. Rows cannot be removed, added or edited without breaking a signature the importer cannot forge.

Forensic envelope

Every chained row carries a small envelope of forensic metadata that the framework attaches automatically, without a bridge or contributor needing to opt in. It lands in context_transient so the regular retention contract applies: a GDPR purge clears the actor / IP / referer along with the rest of the row's transient payload.

Key Source Purpose
uid currentUser->id() Actor identity at write time. 0 for anonymous.
request_uri $request->getUri() with credentials masked, empty string with no request Path the actor was on when the row landed. Empty on non-HTTP paths.
ip $request->getClientIp(), empty string with no request Client IP. Empty outside HTTP.
message_template The PSR-3 message the logger received, or record()'s own Lets the entry-detail page re-substitute placeholders at view time.

The stamp outranks the caller. Every key in the table is taken from what the framework observed, and it overwrites whatever the transient bucket already held under that name. A caller cannot decide what the row says about who acted, from where, or on what path, which matters because the row is then HMAC-signed and attested with those values: a context key named uid is an ordinary thing for a module to pass, and no compromise is needed to pass one.

Nothing is thrown away, with one exception below. A caller value that differs from the observed one is preserved under _audit_trail_caller_supplied (ForensicStamp::CALLER_KEY), so the entry shows both what was passed and what was recorded. Only a differing value counts: core's LoggerChannel::log() already writes the observed uid / ip / request_uri into the context before any logger sees the entry, so an ordinary PSR-3 row arrives carrying the same values the stamp is about to set and records nothing as displaced. ForensicStampTest pins each of these.

The exception is core's own copy of the request URI. The stamp masks credentials out of the URI it records (see below), and core copied the unmasked one into the context before the stamp ran, so on a redacted request the two differ by exactly the secret that was removed. Preserving that as a displaced value would put the secret back on the row under another key, so that one copy is dropped. A request_uri that is not core's copy of this request is somebody claiming a URI the request did not have, and that is still recorded.

The envelope is not pluggable: a bridge cannot register additional forensic fields. The fixed schema keeps the column shape predictable and the GDPR purge contract auditable: every transient column NULL-out wipes the same set of fields regardless of which bridge produced the row. Bridges that want additional attribution emit it through a contributor (subject to the same retention rules) or as plain context keys (which land in transient alongside the envelope).

URL-borne credentials

The URI is masked before it is stamped, because a one-time login link, a password-reset URL or a signed callback would otherwise land in the trail and stay there, readable by anyone with the entries permission rather than only by someone with database access.

Two fields carry a URL, and both are masked: the request_uri the stamp observes, and the referer core copies out of the request header, which is where a one-time login link shows up on the page a user lands on after clicking it. The referer stays caller context rather than becoming a stamp key; it is masked in place, so the envelope keeps its four observed fields. A referer parse_url() cannot parse is recorded as the placeholder alone: the header is whatever the client sent, and a string the module cannot examine is one it cannot call safe.

The rules are two lists in audit_trail.settings, shipped with defaults and meant to be edited: redact_query_keys names query parameters, matched case-insensitively and including nested ones such as ?filter[token]=, and redact_path_prefixes covers what a name cannot reach, because core puts the one-time login hash in the path. A pattern there keeps its own segments and masks whatever follows, so /user/reset/* records /user/reset/7/REDACTED while /user/reset/7, the "send me a link" form, is left alone; a star matches any one segment, which is what the account-cancellation shape needs. Patterns match whole segments, case-insensitively, and tolerate one leading segment, because getPathInfo() keeps a language prefix: inbound path processing rewrites the path Drupal routes on, not the request.

Both lists are on the settings form as well, one entry per line, and emptying both switches the whole thing off: no query bag walked, no referer parsed, and the URI passed through as getUri() built it.

Config rather than constants, and no hook over it. The lists are a site's to edit, including to shorten: code is shipped for OAuth callbacks and is a country code elsewhere, and a site that knows which it has should be able to say so. A module that serves its own credential-bearing URLs adds its pattern to that config in hook_install(), the way modules extend one another's configuration, so there is one place to look and it exports with the rest of the site. It appends to what ForensicStamp::getPathPrefixesToRedact() returns, because a site whose key is absent has the shipped patterns in force without having stored them, and appending to the raw key there would store a list holding nothing else.

Only the two fields the framework itself observes are masked. Payload a caller passes is its own: a message placeholder, a link, or a contributor's data is recorded as given, so a module that logs a credential of its own is the one that has to stop.

Masking happens at write time only, and it has to: context_transient_hash signs the transient bytes, so rewriting a stored value would make the row report itself as tampered. Values captured before the module started masking them age out through the transient purge, which NULLs the column, a state the verifier accepts.

Where nothing matches, which is nearly every write, the URI is recorded exactly as the request carried it: the guard is one hash lookup per query parameter and one route lookup, and the URI is not rebuilt at all, so an encoded ?destination= parameter stays encoded exactly as it arrived.

And it runs once per request rather than once per row. One request writes as many rows as it has bridges and entities, and the masked value is the same for all of them, so it is cached on the request and its route, the route included because it arrives late and a row written before routing must not answer for one written after. Measured, per additional row: 0.12 microseconds, against the 1.2 to 4.8 the getUri() behind the unmasked field cost on every row before this.

Snapshot-delta bucket format

Contributors that snapshot the state of an audited subject on each event (entity create/update/delete, webdav lock/copy/move, etc.) historically emitted dual snapshots under two top-level bucket keys:

{"before": { … }, "after": { … }}

The bucket is normalized on the write path into a compact self-describing shape that keeps every field value at most once and carries diff hints alongside the current state:

{
  "_v": 1,
  "state": {
    "extra": "x",
    "status": 1,
    "tags": ["a", "b"],
    "title": "New"
  },
  "delta": {
    "new": ["extra"],
    "original": {
      "title": "Old",
      "old_field": "old_value"
    }
  },
  "key_order": ["title", "status", "old_field", "tags", "extra"]
}
  • _v is the wire-format version. Readers consult it before trusting the rest of the shape so future format evolutions can ship without rewriting existing rows.
  • state is the current state of every field after the event. Updated fields keep their post-event value here; newly-added fields too; unchanged fields just sit alongside with no annotation.
  • delta.new is the list of field names that appeared for the first time on this event. Values for those fields live in state: the list is names-only.
  • delta.original is the sparse map of previous values for fields whose before-value differed from state. For updated fields, the current value lives in state; for removed fields, the key is absent from state entirely and the value lives only here.
  • key_order is the merged before/after iteration order : state keys plus any removed-field keys inserted at their natural before-position. Survives canonical JSON encoding because arrays preserve order; object keys would otherwise be alphabetized at storage time.

What counts as changed is one rule, in Snapshot\SnapshotDelta::hasValueChanged(), and a reader reconstructing a diff from a row needs it. Where both sides are strings they are compared byte for byte, so "01" and "1" are two values. Anything else is compared loosely, because Drupal field storage hands back one logical value with a different scalar type across a read/write round trip and reporting that would drown real edits in noise: a field holding 1 that comes back as "1" is not in delta.original, and a row that carries it reads as unchanged because nothing was edited. Arrays are walked by key set, since canonical JSON sorts keys at every depth and the chain therefore records no order inside a value; position still counts in a list, where the position is the key.

The same rule decides whether a row is written at all when skip_no_op_updates is on for a bundle, so the check that skips a row and the diff that renders one cannot disagree about what the same value is.

Write-path overview:

flowchart LR
  CB["Contributor returns<br/>{before, after, ...}"]
  EB["AuditTrail::record()<br/>snapshot fold"]
  V{has _v?}
  SC["SnapshotDelta::<br/>computeStateWithDelta()"]
  JSON["canonical JSON<br/>{_v, state, delta, key_order, ...}<br/>: context_*"]
  CB --> EB
  EB --> V
  V -->|yes,<br/>opt-out| JSON
  V -->|no| SC
  SC --> JSON

Field classification at render time, given the shape above:

Field is Classification
In delta.new AND state added
In delta.original AND state changed (before in delta.original, after in state)
In delta.original only removed (value in delta.original)
In state only unchanged

Per-action semantics:

  • Create: emitted as state only (no delta). The row's action=create column implies "everything is new".
  • Delete: emitted as state only, carrying the pre-delete snapshot. The row's action=delete column implies "everything was removed".
  • Update: emitted as state + delta blocks for whatever changed.

Equality

Field comparison uses loose !=. Drupal field storage routinely surfaces purely cosmetic scalar-type drift ('10000.00' vs 10000, '1' vs 1) across read/write round-trips that strict comparison would flag as a real change. PHP 8+ tightened loose equality enough that 0 == '' returns FALSE, so the historical type-juggling gotchas don't apply.

Contributors that need domain-specific equality semantics (fuzzy floats, whitespace-normalized strings, EXIF-stripped image bytes) pre-compute the snapshot-delta shape themselves and emit the resulting _v / state / delta / key_order keys directly. The framework's auto-fold leaves any bucket already carrying a _v key untouched.

Contributor contract

The simplest path: emit before and/or after as top-level peers in a bucket. The framework's SnapshotDelta::computeStateWithDelta() folds them into the canonical shape at write time, in AuditTrail::record(), which is also the only path contributors run on. Existing contributors that already use the dual-snapshot pattern get the new wire format with zero plugin code change.

Contributors that need stricter equality opt out by emitting the canonical shape directly (with _v at the top); the auto-fold leaves their bucket alone.

The HMAC secret

Secrets are first-class config entities: each audit_trail_secret entity carries an integer secret_id matching the per-row secret_id, computed from the bytes the Key holds so it can never point at other bytes, a key_id referencing a drupal/key Key entity holding the actual bytes, and a lifecycle status (pending / active / retired).

Why Key-backed: the module never stores the secret bytes in its own storage. The Key module's provider plugins decide where bytes live: config storage, environment variable, file outside the webroot, AWS Secrets Manager, HashiCorp Vault, HSM-backed providers, etc., and the module dispatches through that abstraction at write and verify time.

Changing the signing secret

The operator provisions a Key; the module points a secret at it and puts that secret in the signing role. It never handles the bytes. Either SecretRepositoryInterface::setSigningSecret() (reached from the admin form) or drush audit_trail:set-signing-secret --key=<key_id>, which also creates the audit_trail_secret record when there is not one, so the whole operation is available from a shell. The repository:

  1. Saves the named secret as active.
  2. Iterates other active entities and retires them.

The order is deliberate: a crash between the two saves leaves two active entities (both with valid Key bytes, so fresh writes still succeed) rather than zero, which would halt every chained write. getSigningSecretId() picks the active secret activated last when more than one exists, so writes land on the new secret even in that transient window.

Because writes carry on, the operation looks finished when it is not, so the status report raises an error while more than one secret is active. Operators converge by re-running the command or by retiring the leftover directly.

Rotation does not rewrite existing rows: every row carries the secret_id it was signed under, and the verifier dispatches per-row. A chain can span any number of rotated secrets without re-signing.

Drop-in via the logger service tag

AuditTrailLogger implements LoggerInterface and is registered with the logger service tag. Drupal's logger.factory collects every tagged logger and dispatches each log call to all of them: exactly the mechanism dblog and syslog use. No bespoke API, no replacement of core services.

Consumers do not need to know audit_trail exists. A call like:

\Drupal::logger('finance')->notice('Acte signed', [
  'audit_trail' => [
    'chain' => TRUE,
    'action' => 'state_change',
    'resource' => 'node/' . $nid,
  ],
]);

works whether audit_trail is installed or not. With the module installed and the entry's context flag set, the call additionally lands in the chain. Without the module, only dblog / syslog record it.

The canonical write path for modules that audit their own business events is the orchestrator service AuditTrailInterface::record(). It pre-buckets context across the two retention tiers via the ContextContributor plugin pipeline, then writes the row through the chain writer and makes one ordinary log call for the other sinks, marked as already chained so this module's own logger does not record it a second time. Buggy contributor plugins cannot cascade into the caller: every applies() / contribute() call is wrapped in try/catch and exceptions are surfaced to the audit_trail logger channel without aborting the event.

Hot-path resolution cache

Every PSR-3 log call that reaches Drupal's logger pipeline hits AuditTrailLogger::log(). The vast majority of those calls are NOT meant to chain (they have no chain key in context). To keep that fast path cheap, the logger lazy-builds a per-request chain registry on the first log() call and reuses it for every subsequent call.

The logger itself is built whenever logger.factory is, since the factory builds every logger-tagged service, and that is nearly every request Drupal handles in full. So it takes the chain writer, the forensic stamp and the chain filters as closures, built on the first entry a chain claims: the writer alone brings the secret repository, the lock and the chain repository with it. A request that logs nothing onto a chain builds none of them.

The registry holds:

  • entities: every active audit_trail_chain entity, indexed by id.
  • channel_claim: PSR-3 channel to chain id (first-match wins, deterministic via ksort).
  • auto_channels: subset of channel_claim restricted to channels that resolve to a mode: auto chain. Drives the implicit-write short-circuit. The default chain contributes its OWN claimed channels (and its id, via channel-as-id); it does NOT auto-claim every other channel via fallback, otherwise a default: mode: auto configuration would chain every log call on the site, including unrelated dblog noise like PHP deprecations.

The implicit-write short-circuit:

log(level, msg, context):
  statement = the chain named in context.audit_trail
  if statement === FALSE: drop
  if no statement:
    if not in auto_channels[$channel]: drop
  if context._audit_trail_already_chained: drop
  chain = registry.resolve(channel, id named in context)
  if no chain: drop, and report it if the caller named one
  if not opted in and chain is not mode: auto: drop
  if one of the chain's filters votes no: drop
  write

Both maps are built from the chain configuration objects, not from the entities built on top of them: status, channels, mode and the id are all on the stored object, and constructing an entity to read them costs more than reading them. The entities are loaded separately and only when a chain actually has to be returned, which is to say only for a log call that survived the short-circuit. On a site with zero mode: auto chains, no chain entity is ever constructed to answer a log call.

Within a request the first call builds the maps and every call after it is two array reads. Neither the maps nor the entities are cached beyond the request: a stale channel map would mean entries silently not chained, and that is the failure this module exists to prevent, so it is not a risk worth taking to avoid a read the configuration system already serves from its own cache.

Lifecycle: the registry follows Drupal's static config-entity cache. Saving a chain entity invalidates via the AuditTrailRegistryHooks hook so the next log() call in the same request picks up the change.