Architecture¶
The chain invariant¶
For each chain id, the audit_trail table holds rows in id
ascending order. Every row carries three signature columns:
previous_hash: the SHA-256 hash of the previous row in the same chain (empty string at genesis).hash: SHA-256 of this row's canonicalized payload (includingprevious_hash), publicly verifiable without any secret.hmac: HMAC-SHA-256(hash, secret), wheresecretis the operator's signing key resolved per row via thesecret_idcolumn.
Two invariants link rows together:
row[n].previous_hash = row[n-1].hash (linking constraint)
row[n].hash = SHA-256(CanonicalJson::encode(payload columns))
row[n].hmac = HMAC-SHA-256(hash, secret(row[n].secret_id))
with row[0].previous_hash = '' (genesis row of each chain).
Chain\CanonicalJson::encode() recursively ksort(SORT_STRING)s
the payload and encodes it via
json_encode(JSON_UNESCAPED_SLASHES | JSON_UNESCAPED_UNICODE |
JSON_THROW_ON_ERROR). The exact byte output is pinned by
tests/src/Unit/CanonicalJsonTest.php so any change to the
encoding fails CI before it can ship: otherwise every existing
row's stored hash would silently become unverifiable.
The columns the canonical payload is built from:
[
'channel' => '…', // PSR-3 channel
'chain' => '…', // chain id
'severity' => …, // RFC 5424 numeric severity
'action' => '…', // event verb (create / update / …)
'resource' => '…', // subject identifier
'context_permanent' => '…', // JSON-encoded permanent bucket
'context_transient_hash' => '…', // SHA-256 of the transient bucket
'created' => '…', // microsecond Unix timestamp
'secret_id' => …, // signing key id
'previous_hash' => '…', // chain link
]
That column set is defined once, in
Drupal\audit_trail\Chain\ChainPayload. The verifier at every
depth, the archiver writing NDJSON and the entry detail page all
read it from there, so none of them can drift from the others and
report an intact row as tampered.
Two retention tiers split the row's context across separate columns:
context_permanentis signed raw in the canonical payload. Kept forever; never purged by retention.context_transientis signed viacontext_transient_hashonly: the column itself is NULLed at retention by the transient-purge stage, and a coveringaudit_trail_segmentrow attests the transition (transient_purged_at != 0). The chain still verifies because only the hash is in canonical.
This shape lets GDPR right-to-erasure operate on the transient
column without breaking chain verification, while still
distinguishing legitimate purge from an attacker NULLing a
column to hide content: a bare NULL on a row whose write-time
context_transient_hash was non-empty AND that doesn't fall
within any covering segment with transient_purged_at != 0 or
archived_at != 0 fails the verifier. (Archive ops legitimize
the NULL because the archive NDJSON envelope captured the
row's transient state at archive time: see
NDJSON envelope shape below, so the
row's pre-purge bytes live in the file even after the live
column is gone.)
Why publicly-verifiable hash plus operator HMAC¶
Two layers:
- The
hashcolumn is recomputable by anyone with the row data: no secret required. A third-party auditor can walk the chain, recompute everyhash, and confirm the linking constraint. This is the public-verifiability anchor. - The
hmaccolumn proves the row was signed by an operator who held the matching secret bytes. Without the secret the hmac cannot be forged; with it, an attacker can sign rows but still cannot rewrite the existing chain (their forgery would need to also rewrite every downstreamprevious_hashand recompute everyhashandhmac, and the operator could detect that via a TSA-anchored snapshot).
Verification runs at one of three depths, given to
verifyChain() as a VerificationDepth: Strict requires every
signature to authenticate and treats an unresolvable secret as a
finding, Tolerant checks what the resolvable secrets allow and
names the rest, and None reads no signature at all. Whatever a
walk accepts without checking is listed on the verdict, and
authenticated_ok is NULL whenever anything was skipped rather
than checked: it reports what the walk did, so a Tolerant walk
that met no missing key says TRUE.
Multi-tamper walk¶
If the verifier sees a row that doesn't validate, it records
the broken range and keeps walking rather than stopping at
the first recovered range. The verdict's broken_ranges list
contains every contiguous failed range in chain order. The
message advertises additional ranges so operators don't ack
the first range and assume the chain is otherwise clean.
Why per-chain (not global)¶
A single global chain has appeal: it gives a total ordering over every event in the system. In practice the cost is too high:
- Every INSERT must read the global head, serialize behind it, and write. Under load (a busy WebDAV save plus a node update hitting at the same time) the lock acquisitions stack up.
- A single break: a corrupted row, an accidental schema migration that touched the table: invalidates verification for every row downstream, including unrelated subsystems.
- Cross-channel cause-and-effect is rarely the forensic question (you usually want "what happened to this acte", not "what happened to the system at 14:23").
The module's default routing maps each PSR-3 channel to a
chain of the same id. Operators that want to funnel related
channels into a single chain register an audit_trail_chain
config entity with the target chain id and a channels[] list
naming the source channels. The choice is per-channel on the
input side and per-chain everywhere else (storage,
verification, retention) so chain composition is a deployment
decision, not a code one.
Per-chain write serialization¶
INSERTs into the same chain need to be serialized: otherwise
two concurrent writers might both read the same head hash and
emit two rows with the same previous_hash, breaking the
chain.
The module acquires a Drupal lock named
audit_trail.write:<chain_id> for the duration of SELECT
head, compute hash + hmac, INSERT. Locks are per-chain, so
unrelated chains do not contend.
Schema-level defense-in-depth: a UNIQUE index on (chain,
previous_hash) makes chain forks impossible at the database
level. Even if the application-level lock fails to serialize,
the second row's INSERT fails with a uniqueness violation
rather than silently extending the chain in two places.
If the application lock cannot be acquired within 5 seconds,
the log entry is dropped from the chain rather than blocking
the request indefinitely: the underlying dblog or syslog
modules still record the entry on their side. Every drop
bumps audit_trail.dropped_under_contention in State, and a
non-zero counter surfaces on /admin/reports/status as a
warning so the gap is operator-visible.
The lock cannot reach inside somebody else's transaction¶
The serialization above holds only while the writer can read a head that reflects committed reality. Inside a caller's open transaction it cannot: the lock is released before the row is visible to anyone else, and the head read is bound to the caller's snapshot, so two callers whose transactions overlap read the same head however the lock is scheduled. The unique key does its job, the chain never forks, and the second entry is dropped instead.
An entry written inside an open transaction is therefore not chained there. It is buffered and chained from a post-transaction callback, once the caller's outermost transaction has resolved. The callback fires once, at that outermost resolution and never on a savepoint release, and it reports success for a committed transaction and for one a DDL statement voided, each of which did make the caller's work durable. The whole batch takes one lock hold and one head read, each subsequent entry linking to the hash the previous one just produced, so buffering makes a multi-entry transaction cheaper here rather than dearer.
Where the entry waits is
in_transaction_write_mode:
the audit_trail_outbox table by default, which is atomic with
the change it describes and survives a crash before the flush.
This applies only to writes that ask the writer for the lock.
The module's own lifecycle writes pass acquire_lock: FALSE
because they hold the chain lock themselves, across their whole
transaction, so they are serialized correctly already, and they
read the returned row id to bind segment bookkeeping to it.
Neither would survive being buffered, and the acquire_lock
test is what keeps them out of it.
Storage layout¶
A single Drupal-managed table audit_trail holds all chains.
Each row carries:
| Column | Notes |
|---|---|
id |
Serial primary key, monotonically increasing |
created |
Microsecond Unix timestamp (string, 16 chars) |
channel |
Originating PSR-3 channel. ASCII only, the requirement core makes of dblog.type with the same column type; on MySQL the column enforces it |
chain |
Chain group id (defaults to channel) |
severity |
RFC 5424 numeric (0=emergency .. 7=debug) |
action |
Event verb (create / update / delete / …); used by the entries listing |
resource |
Subject identifier (e.g. entity:node/42, webdav:files/contract.docx) |
context_permanent |
JSON-encoded permanent-bucket payload; signed raw in hash |
context_transient |
JSON-encoded transient-bucket payload; NULLed at retention by the transient-purge cron run (the cleared range is attested by a transient-purge segment). When the operator opts out of transient-purge, the raw bytes are preserved into the archive NDJSON envelope at archive time and live there until file-purge; the column itself is gone after live-purge regardless. Signed via context_transient_hash (only the hash is in the canonical, never the raw bytes) |
context_transient_hash |
SHA-256 of context_transient at write time; empty string when the bucket had no content |
secret_id |
Integer id of the audit_trail_secret entity that produced hmac |
previous_hash |
hash of the previous row in the same chain (empty at genesis) |
hash |
SHA-256 of the canonical payload: publicly verifiable |
hmac |
HMAC-SHA-256(hash, secret(secret_id)): operator-verifiable |
A compound (chain, id) index covers the hot read path
(fetching a chain's head for the next INSERT, walking a chain
for verification). A UNIQUE index on (chain, previous_hash)
enforces the no-fork invariant at the schema level. Secondary
indexes on channel, created, action, resource,
severity and correlation_id back the admin entries page and
scripted queries alike.
Which of the page's twelve filters an index can actually serve is worth spelling out, because they are not all matched the same way:
- Chain, action and correlation ID are equality matches, and each has an index that answers them. So is the created range behind the From and To fields.
- Severity is
severity <= non an indexed column, so whether the planner takes the index comes down to how much of the table the bound admits. A permissive bound reads as a scan. - Channel and resource are matched with a leading wildcard
(
LIKE '%needle%'), which no B-tree can serve. Their indexes are there for callers that match a prefix instead, such asresource LIKE 'entity:node/%'from a script. - User, message, client IP and request URI read JSON out of
context_transient, which no index covers at all.
Chain, action, correlation ID and a date range are therefore the ones to reach for first on a large table. Combine one of them with whichever of the rest you actually need, and the scan is bounded by the indexed one.
audit_trail_segment.lifecycle_hmac default value¶
The lifecycle_hmac column defaults to '' (empty string)
at the schema level. This is defensive: if a code path ever
inserts an archive bookkeeping row without computing the HMAC
the column ends up non-NULL but empty, and the verifier
treats empty lifecycle_hmac as an invalid signature (the
HMAC of any byte string under any secret is 64 hex
characters, never the empty string).
In practice every code path that creates an archive record also computes the HMAC inline, so the default is unreachable. It stays in the schema as belt-and-suspenders insurance.
Write-path index trade-off¶
audit_trail carries seven secondary indexes plus the (chain,
previous_hash) UNIQUE key. On high-write workloads (e.g.
1 000 rows/second sustained into a single chain) the bottleneck
is index maintenance: each INSERT updates eight B-trees. The
compound (chain, id) is one of the seven and the one every
write certainly touches, so leaving it out of the count
understated the cost this section exists to quantify.
For 1.0 the index set is what every operator wants by default
(filter by channel, severity, action, created, chain). Sites
that profile a hot write path and confirm specific filters
are never used can DROP those indexes via a custom
hook_update_N; the chain semantics don't depend on the
secondary indexes, only on (chain, previous_hash) (UNIQUE)
and (chain, id) (compound). A future settings flag for the
secondary index set is on the roadmap if real operators ask
for it.
id column width¶
The id column is declared serial at big size, which on
MySQL / MariaDB resolves to BIGINT UNSIGNED: maximum
18,446,744,073,709,551,615 rows. At a million rows a day
that is about 50 billion years, so no deployment reaches it.
The width is deliberate rather than defensive. One sequence is shared by every chain, and a value is never reused: auto-archive and live-purge move old rows out of the live table and reclaim its space, but not its numbers. So the ceiling counts every row the site has ever written, not the rows it currently holds. A 32-bit column would have run out after 4,294,967,295 of them, which is about twelve years at a million rows a day, and widening it at that point means rebuilding a table with billions of rows in it.
Every column that stores an audit_trail.id value is as
wide as the column it references: audit_trail_checkpoint's
last_id, both range bounds on audit_trail_segment and
audit_trail_acknowledgment, and each of the four
*_event_id attestation references on a segment. A narrower
one would fail before the id itself did.
It buys nothing cryptographic: id is not among the columns
ChainPayload signs, nor among the values a segment's
archive_hmac covers, and an archive records the ids it
holds as values. Nothing that has to keep verifying depends
on how wide the column is.
What the width costs¶
Index space, and almost nothing else. InnoDB appends the primary key to every secondary index entry, so the four extra bytes are paid once in the row and again in each of the eight index structures. Measured on MariaDB over 200,000 rows carrying this module's own average payload (533 bytes permanent, 298 transient):
| 32-bit | 64-bit | delta | |
|---|---|---|---|
| data (primary key leaf) | 260.77 MB | 260.83 MB | +0.0% |
| all indexes | 60.33 MB | 70.44 MB | +16.8% |
| whole table | 321.10 MB | 331.27 MB | +3.2% |
Four bytes on a row already carrying about 1.3 KB disappears
into existing page slack, which is why the data side does not
move. A secondary index entry is small enough that the same
four bytes is a third of it: on the severity index an entry
goes from 13.2 to 18.4 bytes, and that index alone grows 40%.
Two of the eight are load-bearing, and the write-path index
trade-off above says which: (chain, previous_hash) and
(chain, id). Keeping only those, the same measurement reads
+10.7% on indexes and +1.1% on the table, because 69% of
the increment sits on the six indexes a site is free to drop
when it does not filter on them. correlation_id is the
first to consider: it grew 2.00 MB, more than the fork guard
did on a structure four times its size, and the schema
describes it as observability metadata rather than part of
the integrity contract.
No timing cost was measurable in either direction. An index-only range scan over the same data measured 9.2 ms against 9.1 ms per query, and bulk insert throughput was indistinguishable across three alternating rounds.
Retention is the lever, not the schema¶
Index size is linear in the number of live rows, and the
number of live rows is what archive_after and
live_purge_after decide. So the way to keep the database at
a reasonable size is the one this module already ships:
archive, then purge the live rows an archive holds. Over a
live table holding ninety days (live_purge_after: P3M):
| events/day | live rows | all indexes | required only |
|---|---|---|---|
| 1 000 | 90 000 | 27 MB to 32 MB | 13 MB to 15 MB |
| 100 000 | 9 000 000 | 2.7 GB to 3.1 GB | 1.3 GB to 1.4 GB |
| 1 000 000 | 90 000 000 | 26.5 GB to 31.0 GB | 12.8 GB to 14.2 GB |
A site that finds its index tight shortens the window:
moving live_purge_after from three months to about eleven
weeks gives back the whole 16.8%, costs no schema change and
no downtime, and loses nothing, because live-purge only
removes rows an archive already holds. The stages and their
thresholds are described under Retention
lifecycle below.
Which is also why the width is worth paying for. Purging bounds the space; it cannot bound the numbers, because a purged id is never reissued. Retention answers how large the database gets, and the column width answers how long the chain can go on being written to. They are different questions, and only one of them has a configuration setting.
A second table, audit_trail_checkpoint, records verification
checkpoints: (chain, last_id, last_hash, created, hmac,
secret_id) rows minted each time the verifier walks a chain
cleanly to its head. The next incremental walk starts after
last_id and uses last_hash as its expected_previous,
bounding verification cost to events accumulated since the
last checkpoint. The checkpoint's own HMAC is recomputed
before trust; a forged checkpoint falls back to a full walk
from genesis. The checkpoints are an optimization layer for the rows that
survive: lose them and a full walk from genesis still proves
the integrity of every row still present. What it cannot
prove without them is that no rows were removed from the END,
because what remains is a consistent shorter chain and the
checkpoint is the only record of how far the chain reached. See verification for the
operational pattern.
A third table, audit_trail_segment, records the lifecycle
state of every contiguous chain id range that has been
processed in some way (archived, transient-purged, live-
purged, file-purged). One row per range, created lazily on
the first lifecycle op that touches the range and surviving
the full lifecycle (including file-purge) so the verifier can
bridge across long-purged ranges via the row's
anchor_before / anchor_after columns. Each lifecycle
transition records its timestamp + the audit_trail.id of
the chain event that attested it (segment_archived,
segment_transient_purged, segment_live_purged,
segment_file_purged); the chain event carries the segment id
in resource='segment:<N>' and the verifier cross-checks
both directions of the mutual reference. Three independent
HMACs cover different concerns on the row:
hmac (identity, sealed at creation), archive_hmac
(archive content, sealed at archive op), lifecycle_hmac
(re-signed on every state mutation under the currently-active
secret).
A fourth table, audit_trail_acknowledgment, indexes the
operator attestations that a specific row range is known to be
unverifiable (e.g. signed by a now-deleted secret, or restored
from a backup with no chain link). The verifier silently skips
acknowledged ranges and reports them in the verdict.
The record is the chain, not the table. Recording, moving and
deleting an acknowledgment each append an acknowledgment_*
chain event carrying the range, the reason, both anchors and the
creation stamp, signed and linked like any other chain row.
Replaying those events in id order rebuilds the table exactly,
which is what drush audit_trail:reindex-acknowledgments does
and what the status report checks on every build.
So the index rows carry no signature of their own. Signing the
copy while never reading the original was backwards: the copy
was what a single DELETE could revoke without trace, and the
tamper-evident original went unread. What is checked now is
whether the index still agrees with the chain, which is a
question a DELETE, a hand-edited range and a restore from a
partial backup all answer loudly.
A live-purge does not evict the index entries covering the range it takes. It used to, because the archive carried a snapshot of each one; with no snapshot, evicting would destroy the only live record of why a range reads broken. What is left is inert while the rows are gone, because its anchors point at rows that are not there, and it covers them again if they are restored.
An acknowledgment eventually outlives its own recording event,
which retention archives and purges like any other row. The
entry then cannot be rebuilt from the chain, and is not treated
as one nobody recorded: its id falls inside a range a signed
segment_live_purged event says was taken, and the archive of
that range still carries the event for anyone verifying against
the file.
The table holds no actor. The acting operator is in the
acknowledgment_* event's transient bucket. A uid on the
table could not be emptied at all: it would be copied into the
archive NDJSON, which no purge stage reaches.
The transient bucket is the one tier the retention lifecycle can
empty, which is what makes erasing an actor possible at all. It
empties on age, over a range of rows, attested by a segment: a
retention control rather than a per-subject erasure on request,
and one that does nothing until transient_purge_after is set,
which it is not by default. The acknowledgments listing links
each id to the entries listing filtered on its resource, which
is where that history is read.
Retention lifecycle¶
Five cron stages carry rows through retention, gated by five
thresholds clocked from the row's created timestamp, with a
sixth stage, coverage, ahead of them all to put rows into
segments in the first place:
transient-purge to archive to live-purge to file-purge to compaction
↑ ↑ ↑ ↑ ↑
transient_purge archive_after live_purge file_purge compact
_after (opt-in) (required) _after _after _after
(opt-in)
Drupal builds a hook class in order to call it, so a collaborator taken as an ordinary constructor argument is built before the flag that governs it can be read, and the switch stops switching anything off. Every cron hook here therefore takes what it needs to do the work as a closure, and reads its flag first:
- auto-archive takes the archiver and the chain registry, behind
cron_archive.enabled, which ships disabled; - auto-verify takes the verifier, behind
auto_verify_enabled, which ships enabled; - the TSA bridge takes the timestamper, the chain repository and
the chain registry, behind
audit_trail_tsa.settings'cron_enabled, which ships disabled.
Two of those three ship off, so a tick that builds them is what a site got by installing the module and leaving it alone: 7.2 ms and 3.8 ms of audit graph for the first two, measured on Drupal 11.3, on every tick of a site that wanted neither.
The settings form enforces transient_purge_after < archive_after
< live_purge_after < file_purge_after < compact_after at submit
time.
Transient-purge runs before archive on the same tick because
it must: once a row is archived, the NDJSON file has frozen
whatever transient state the live row was in. A late transient-purge
(after archive) would NULL only the live column while the file
still holds the bytes: defeating the operator's
data-minimization intent.
Each stage writes a chained event row attesting its transition
(segment_transient_purged, segment_archived,
segment_live_purged, segment_file_purged). The
audit_trail_segment bookkeeping row records the matching
*_at timestamp + the chained event's audit_trail.id in the
*_event_id column. Both writes happen inside one DB
transaction so a process crash between them can't persist a
"ghost segment" (state flag set, event id zero).
Two thresholds are optional, and both ship empty.
transient_purge_after disables its stage when empty; archives
written for that chain carry the raw context_transient bytes
into the NDJSON envelope as a sibling of the canonical payload
(see below). Operators who DO configure transient-purge can also
disable it per chain via the zero-duration sentinel PT0S on
the per-chain form. compact_after disables compaction the same
way, which costs nothing a verifier can observe: what it folds
is the bookkeeping tail behind file-purge.
Segments as the lifecycle unit¶
Each stage operates on segments: contiguous id ranges
recorded in audit_trail_segment, not on individual rows.
A segment is just a bookkeeping row pinning [from_id,
to_id] plus the lifecycle stamps (archived_at,
transient_purged_at, live_purged_at, file_purged_at).
Rows themselves stay in audit_trail; the segment table
records which ranges have been processed in which way.
A segment is bare when all four stamps are zero (it
exists but no lifecycle op has touched the range yet),
transient-purged when transient_purged_at != 0,
archived when archived_at != 0, and so on. Stamps
accrue in the order
transient_purged_at to archived_at to live_purged_at to file_purged_at;
each lifecycle stage advances segments that match its
prerequisite stamp pattern.
Chain writes are tied to segment existence as well as to
lifecycle transitions. Creating a segment writes a
segment_created event carrying the segment's whole identity
envelope, and the audit_trail.id that event lands on becomes
the segment's id: the record and the index entry are one
number. Each later stage (transient-purge, archive, …) emits
its own segment_<verb> event and stamps the matching
*_event_id column, and carries the identity envelope too,
because segment events age out under retention like any other
row and whichever survives longest is the one a rebuild reads.
audit_trail_segment is therefore an index over what the chain
attests, the same relationship audit_trail_acknowledgment
has. Two consequences follow.
A segment row that goes missing is detectable: its creation
event names a row that is not there. Two further signals cover
segments minted before creation was recorded, and damage the
chain cannot speak about at all: neighboring segments meet
exactly, the earlier one's anchor_after being the later
one's anchor_before, so SegmentIndex::findAnchorBreaks()
reports a segment removed from between two that remain; and
coverage mints per run of chain-linked rows, so a hole left by
a missing segment is never sealed over by a segment minted
straight across it.
And a segment row cannot be removed to suit another operation.
Compaction is the one path that removes rows, and it attests
what it absorbed in compacted_from; the verifier accepts a
reference to an absorbed segment on that evidence. There is no
"delete this segment" operation: the retention policy advances
segments, and an operator changes the policy rather than a
row.
The cron pipeline shares one mental model:
- Cover orphan rows (rows past the earliest-needed
retention cutoff not yet inside any segment). The coverage
stage groups them into closed granularity buckets and mints
one bare segment per uncovered gap via
ChainArchiver::ensureSegmentCoverage(). Coverage is the SOLE minter of bare segments in the cron pipeline. - Walk existing segments in the prerequisite lifecycle
state past each downstream stage's cutoff, and advance them
by calling the matching
ChainArchivermethod. Each downstream stage (transient-purge, archive, live-purge, file-purge) only walks; none mints.
Operator-driven entry points (drush audit_trail:archive,
the segments admin Save action, the tugboat seeder) also call
ensureSegmentCoverage() directly when they need bares
outside the cron cadence.
Restore is a one-way ratchet, not a lifecycle escape¶
SegmentRestorer::restoreSegment() brings an archived segment's rows
back into audit_trail for incident response or operator
inspection. On a successful restore, the segment row's
live_purged_at is cleared back to 0 and
live_purged_event_id is re-pointed at the freshly-minted
segment_restored chain event. lifecycle_hmac is re-signed
under the restored event's signing secret. Three things follow:
- Retention re-applies. The cron live-purge stage selects
segments via
live_purged_at = 0 AND to_created < cutoff; restored segments match again, so the next eligible tick re-DELETEs the rows and stamps a freshsegment_live_purgedevent. Restored rows can't dodge thelive_purge_afterceiling. - Operators inspecting restored rows must pause cron (or
temporarily lift
live_purge_after). The unpaused-cron behavior is to re-purge within one tick of the threshold window's cron cadence. - Lifecycle stamps don't progress strictly monotonically
anymore. The original
transient_purged_at to archived_at to live_purged_at to file_purged_atordering invariant pins the first time each transition fires; restore + re-purge can produce a chain with multiplesegment_live_purgedevents for the same segment, in id-ascending order, interleaved with asegment_restoredevent. The segment row carries the LATEST live-purge transition; the chain history carries the full sequence.
The verifier's segment-event cross check is aware of this
shape: live_purged_event_id on the segment row is allowed
to point at any later same-chain, same-segment event whose
action is segment_live_purged or segment_restored. See
docs/verification.md for the lookup rule and what tampers
it still catches. file_purged_at is left untouched on
restore: if the file was already file-purged before
restore (only possible via the override_path import path),
the cron file-purge stage's own file_purged_at = 0 filter
prevents a second file-purge from firing.
Restore ordering: event first, rows second, segment row last¶
restoreSegment() is a three-step orchestration behind three
pre-flight refusals; each step has a specific concurrency
contract.
- Pre-flight: three refusals, all of them before anything
is written. The segment already in the finalized-restored
state (
live_purged_at = 0ANDlive_purged_event_id != 0), which is a double-call on a segment that restored cleanly. The segment with asegment_restoredevent already in flight, which is a previous attempt that did not finalize; asked here as well as under the lock in Step 1, so that a half-finished restore is diagnosed by the event it left rather than by the rows that event put back. And the range whose own chain still has rows in it, which is every archived segment until live-purge takes it.
They exist because of what came after them. The PRIMARY KEY
rejects the replay either way, but only in Step 2, and Step 1
has written the segment_restored event by then:
SegmentIndex::replaySegmentEvents() reads that event as the
live-purge pointer move it would have been, so the chain
attests a restore that never happened and
compareIndexToChain() reports the segment altered from then
on. A refusal that arrives after the attestation is not a
refusal.
The range is counted on the segment's own chain, because the range is a span and not the set of ids the replay will use, and the two coincide on that chain alone: every archived id in the range was a row of it. Counting the range across all chains would block a restore that is free to proceed, since interleaved chains leave their own rows at ids inside it, untouched by its purge. What that leaves open is another chain's row on one of the archived ids, which needs the id set rather than the range: #3620412.
The importFromFile-then-restore path is preserved: a
freshly-imported segment has no live rows of its own chain and
live_purged_event_id = 0.
- Step 1: emit the
segment_restoredchain event under a brief chain-write lock. The lock guards the chain anchor + per-chain id sequence for the single INSERT, then releases. Step 0 (the in-flight-event check) runs inside the lock so two concurrent restoreSegment() calls on the same segment can't both pass the check and both emit; the second caller throws on the in-flight detection (the crashed-mid-restore signature that the pre-flight check cannot detect, because the segment row's lifecycle pointer hasn't moved yet). - Step 2: replay archived rows + acks inside a DB transaction with no chain-write lock held. Logger writes on the same chain proceed unblocked; their AUTO_INCREMENT ids never collide with the explicit archived ids (which are always below the live head). A mid-loop throw rolls the transaction back so partial INSERTs don't persist.
- Step 3: reacquire the chain-write lock briefly, reload
the segment row, sign
lifecycle_hmacagainst the FRESH lifecycle fields, then UPDATE the four restore-mutated columns (live_purged_at = 0,live_purged_event_id= restored event id,lifecycle_secret_id,lifecycle_hmac). The lock serializes Step 3 againstpurgeSegmentArchiveFile(), the only sibling lifecycle op whose preconditions are satisfied during the Step 1 -> Step 3 window: a concurrent file-purge would otherwise leavelifecycle_hmacsigned against stalefile_purged_at = 0values.
The chain attestation is therefore committed BEFORE rows land in the live table. This is the property that lets Step 2 run without holding the chain-write lock: every intermediate state (Step 1 done / Step 2 partially done / Step 2 fully done / Step 3 done) verifies cleanly against the existing verifier rules, because:
- The verifier does not require "rows present in
[from_id, to_id]whilesegment.live_purged_at != 0" to be absent. It walks visible rows and validates their hash + previous_hash linkage. segment_restoredis a non-cross-checkable lifecycle action; the verifier doesn't strict-equal it against any segment-row column.- The existing live-purge supersession rule already accepts
segment.live_purged_event_idpointing at a latersegment_restoredevent (post-Step-3 state). Pre-Step-3 (segment row still references the old purge event) is the baseline strict-match case.
If Step 1 succeeds but Step 2 or Step 3 fails, the segment is
in a partially-applied state: the chain records the restore
intent, but the rows either rolled back (Step 2 failure) or
landed without the segment-row finalization (Step 3 failure).
A retry of restoreSegment() is refused by Step 0: the operator
must verify chain integrity, audit_trail row presence in the
archived window, and the segment row state, then finalize
directly (re-run the missing step manually under SQL). The
restoreSegment() API does not offer an automated resume path; the
partial-state error message names the emitted event id and
the [from_id, to_id] window so the operator knows exactly
what to inspect.
Coverage stage: minting bares around granularity buckets¶
segment_granularity (hour / day / week / month)
controls the bare segment width minted by the coverage stage.
hour is intended for staging / testing workflows where
waiting a full day for the first segment to materialize is
impractical; production installs typically run with day or
coarser. The coverage stage groups orphan rows into closed
buckets: buckets whose end-of-window has passed the
coverage cutoff, and mints one bare segment per uncovered
gap inside each closed bucket via
ensureSegmentCoverage(). Downstream stages (transient-
purge, archive, live-purge, file-purge) operate on segments
coverage produced; none of them mint.
cutoff (earliest needed)
↓
time ──── bucket A ──── │ ── bucket B (still open) ── …
┌──────────────┐ │ ┌────────────────────┐
rows │ R1 R2 R3 R4 │ │ │ R5 R6 R7 │
└──────────────┘ │ └────────────────────┘
↓ closed │ ↓ open
mint one bare │ skip: a future row may
segment over │ still land in bucket B
[R1.id, R4.id] │ before its window closes
↓
segment A: bare (no lifecycle stamps yet)
↓
later cron tick past transient_purge_after:
runTransientPurge() advances segment A,
stamps transient_purged_at and emits chain event.
The coverage cutoff slides:
- With
transient_purge_afterset, coverage usestransient_purge_after_us: rows enter segments just before transient-purge needs them. - Without
transient_purge_after, coverage falls back toarchive_after_us: rows enter segments just before archive needs them.
The "closed bucket" requirement matters because a still-open
bucket may receive new rows in the next cron tick, which
would either land in a row id below the bare's from_id
(impossible: ids are monotonic per chain) or require
amending an immutable segment (forbidden by the chain
attestation). Closed-bucket minting guarantees the bucket
boundaries match what archive will later turn into a single
NDJSON file.
The segment spine¶
The rows of a chain link to each other: every row's hash
covers the previous row's, so altering one breaks the next.
Segments did not link to each other. Each one was attested
only by its own segment_created and segment_archived
events, and retention purges those like any other row. From
that point the segment row rested on its three HMACs, which
is the operator's word and unreadable to anybody without the
key.
The spine closes that, by doing to segments what the chain already does to rows. It has a name of its own because "chain" is already the most load-bearing word here, and a sentence mentioning both needs to say which one it means.
The construction¶
Each archived segment produces a DIGEST over its own fixed facts, and the digests link:
D_k = H(chain, from_id, to_id, anchors, file_hash, row_count)
S_k = H(k, S_k-1, D_k)
S_k is stored on the segment row as spine, with k as
spine_height, and written into the segment_archived
event's permanent bucket. A consolidated row has no archive
event of its own, so the head it adopts is written into its
segment_compacted event instead, and a replay reads both.
The digest covers immutable facts only. A segment's lifecycle
stamps are re-signed on every transition by design, and a
transient-purge months later must not move a value already
recorded elsewhere, so the purge stamps, the event ids and
verification_failed_at_purge stay outside it.
There is no separate digest over the rows, because
anchor_after already is one: a row's hash covers its
previous_hash, so the hash of the row at to_id
transitively covers every row back to anchor_before.
flowchart LR
subgraph segments["audit_trail_segment"]
S1["segment A, h=1"] --> S2["segment B, h=2"]
S2 --> S3["segment C, h=3"]
end
subgraph chain["audit_trail"]
E1["segment_archived, spine = S1"]
E2["segment_archived, spine = S2"]
E3["segment_archived, spine = S3"]
end
S1 -. "must agree" .-> E1
S2 -. "must agree" .-> E2
S3 -. "must agree" .-> E3
What it buys¶
A new archive re-attests the old ones. Segment 5's facts are folded into the head on segment 40, whose archive event is recent and still within retention. Evidence for an old segment is carried forward by later activity instead of dying with that segment's own events.
One value covers everything below it. Checking forty segments used to need forty surviving events; a head at height 40 commits to all forty.
A local edit becomes a global rewrite. Altering segment 5 changes every head above it, so every later segment row and every later archive event has to change too, and those are chain rows whose HMACs an attacker without the key cannot produce.
That last one is the whole of it, and its limit is the same as everywhere else in this module: someone holding the signing secret can perform the global rewrite. What they cannot do is a small one.
What compaction leaves¶
Compaction writes one consolidated row over a run of
file-purged segments and deletes every segment in the run,
which takes with it the fields their digests were derived
from. The consolidated row records the height and head of the
highest segment it absorbed, and compacted_segment_count
marks it as consolidated rather than ordinary while saying how
much of the spine it occupies. That count is transitive: a run
can contain a consolidated row from an earlier round, and it is
counted for the whole span it already stood for rather than as
one row, so the number is always how many original segments the
row speaks for.
flowchart LR
subgraph before["before"]
A["seg A, h=1<br/>spine S1"] --> B["seg B, h=2<br/>spine S2"]
B --> C["seg C, h=3<br/>spine S3"]
end
subgraph after["after folding A and B"]
X["consolidated, h=2<br/>spine S2, count 2"] --> C2["seg C, h=3<br/>spine S3"]
end
before --> after
C2 --> R["S3 = H(3, S2, D_C)<br/>so an edited S2 fails here"]
The highest rather than the last in the run, because a run is grouped anchor to anchor in RANGE order while the spine runs in ARCHIVE order, and the two part company on a chain that has imported an archive or archived a later range before an earlier one.
A replay ADOPTS that head rather than deriving it, and nothing is lost by doing so. Deriving it would mean folding digests stored on the row itself, and no stored digest can be compared against the segment it describes once that segment is gone. What holds the consolidated row to account is the segment after it, whose own head was folded onto this one, and the lifecycle events above that.
For the same reason a consolidated row never appears in the
verdict's covered map: its own range and anchors are values
compaction restated, and editing them moves no digest, so the
spine cannot speak for them. What accounts for those is the
segment_compacted event, through
SegmentIndex::compareIndexToChain().
What it is not¶
It is not a way to prove anything about purged rows. The rows themselves are destroyed, and nothing recovers what a hash was computed from. If archive files are exported to WORM storage, those files are what prove the rows, and the whole chain can be rebuilt from them without this. If they are not exported, the rows are gone and the only thing left to attest is the account of them: the ranges, the counts, the boundaries and the dates.
So what this adds is tamper-evidence for that account, at the cost of three columns and two hashes per archived segment. It is not a retention tier, not a proof of content, and not a defense against an operator who was dishonest when the account was written.
NDJSON envelope shape¶
Each archive file is a sequence of newline-delimited JSON envelopes. Three envelope types appear, in this order:
{"payload": <canonical>, "transient": <raw|null>, "type": "row"}
…
{"payload": <ack-canonical>, "type": "ack"}
…
{"payload": <archive-record-canonical>, "type": "archive_record"}
rowenvelopes carry one audit row each, ordered by ascendingid.payloadis the canonical the row'shashwas computed over (version, channel, chain, severity, action, resource, context_permanent, context_transient_hash, created, secret_id, previous_hash): byte-identical to what the live verifier reconstructs from the DB row.versiontravels with the rest, so an archive read back years later says which digest to check it under rather than leaving the reader to assume.transientrides outsidepayloadbecause the canonical / hash / HMAC layer was sealed at write time and cannot include bytes that the cron transient-purge stage may later drop. The verifier hash-binds the side-channel back to the canonical via the already-signedcontext_transient_hash: whentransientis non-null,sha256(transient) == payload.context_transient_hashmust hold, otherwise the file is tamper-evident at that row.
The transient field is optional. An envelope without it is
read as "no forensic side channel available"; the verifier
falls back to the segment-coverage rule on the live row.
There is no acknowledgment envelope. The file used to snapshot
one line per acknowledgment covering the range, which was a copy
of a derived index: an acknowledgment is an
acknowledgment_recorded chain event, so it is already in the
file wherever that event falls inside the archived range, and a
restore rebuilds the index from those events rather than from a
snapshot it has no way to vouch for.
archive_recordenvelope is the last line and carries the metadata the verifier needs to rebuild theaudit_trail_segmentrow from the file alone after a total- DB-loss scenario: chain, range, anchors, row count, range timestamps, signing secret id, and the segment's identity HMAC.
The spine is deliberately NOT in the footer, and cannot be. A footer is part of the file whose digest the archive envelope signs, so every field in it has to be known before the file is closed, while the spine height is assigned in the phase after that, under the chain-write lock. It is also not the file's to carry: a file is one segment, and the spine is a statement about the whole chain of them. An imported archive therefore joins the spine at the end rather than reclaiming a position recorded in its bytes.
The whole file's SHA-256 (over data lines + footer) is signed
into audit_trail_segment.archive_hmac at archive time, so
while that row survives, any byte-level edit anywhere in the
file is detectable independent of the per-row HMAC layer.
Import is the case where that row does not survive, and there
the digest is recomputed from the file being imported, so it
attests nothing on its own. What stands in for it is the file's
own structure: the footer's two anchors are inside the identity
HMAC, the rows each name the hash of the row before them, and
import walks that run and refuses a file whose rows do not start
at anchor_before, end at anchor_after, and number what the
footer claims. Rows cannot be removed, added or edited without
breaking a signature the importer cannot forge.
Forensic envelope¶
Every chained row carries a small envelope of forensic
metadata that the framework attaches automatically, without
a bridge or contributor needing to opt in. It lands in
context_transient so the regular retention contract
applies: a GDPR purge clears the actor / IP / referer along
with the rest of the row's transient payload.
| Key | Source | Purpose |
|---|---|---|
uid |
currentUser->id() |
Actor identity at write time. 0 for anonymous. |
request_uri |
$request->getUri() with credentials masked, empty string with no request |
Path the actor was on when the row landed. Empty on non-HTTP paths. |
ip |
$request->getClientIp(), empty string with no request |
Client IP. Empty outside HTTP. |
message_template |
The PSR-3 message the logger received, or record()'s own |
Lets the entry-detail page re-substitute placeholders at view time. |
The stamp outranks the caller. Every key in the table is
taken from what the framework observed, and it overwrites
whatever the transient bucket already held under that name. A
caller cannot decide what the row says about who acted, from
where, or on what path, which matters because the row is then
HMAC-signed and attested with those values: a context key named
uid is an ordinary thing for a module to pass, and no
compromise is needed to pass one.
Nothing is thrown away, with one exception below. A caller value
that differs from the observed one is preserved under
_audit_trail_caller_supplied
(ForensicStamp::CALLER_KEY), so the entry
shows both what was passed and what was recorded. Only a
differing value counts: core's LoggerChannel::log() already
writes the observed uid / ip / request_uri into the context
before any logger sees the entry, so an ordinary PSR-3 row
arrives carrying the same values the stamp is about to set and
records nothing as displaced. ForensicStampTest pins each of
these.
The exception is core's own copy of the request URI. The
stamp masks credentials out of the URI it records (see below),
and core copied the unmasked one into the context before the
stamp ran, so on a redacted request the two differ by exactly the
secret that was removed. Preserving that as a displaced value
would put the secret back on the row under another key, so that
one copy is dropped. A request_uri that is not core's copy of
this request is somebody claiming a URI the request did not have,
and that is still recorded.
The envelope is not pluggable: a bridge cannot register additional forensic fields. The fixed schema keeps the column shape predictable and the GDPR purge contract auditable: every transient column NULL-out wipes the same set of fields regardless of which bridge produced the row. Bridges that want additional attribution emit it through a contributor (subject to the same retention rules) or as plain context keys (which land in transient alongside the envelope).
URL-borne credentials¶
The URI is masked before it is stamped, because a one-time login link, a password-reset URL or a signed callback would otherwise land in the trail and stay there, readable by anyone with the entries permission rather than only by someone with database access.
Two fields carry a URL, and both are masked: the request_uri
the stamp observes, and the referer core copies out of the
request header, which is where a one-time login link shows up on
the page a user lands on after clicking it. The referer stays
caller context rather than becoming a stamp key; it is masked in
place, so the envelope keeps its four observed fields. A referer
parse_url() cannot parse is recorded as the placeholder alone:
the header is whatever the client sent, and a string the module
cannot examine is one it cannot call safe.
The rules are two lists in audit_trail.settings, shipped with
defaults and meant to be edited: redact_query_keys names query
parameters, matched case-insensitively and including nested ones
such as ?filter[token]=, and redact_path_prefixes covers what
a name cannot reach, because core puts the one-time login hash in
the path. A pattern there keeps its own segments and masks
whatever follows, so /user/reset/* records
/user/reset/7/REDACTED while /user/reset/7, the "send me a
link" form, is left alone; a star matches any one segment, which
is what the account-cancellation shape needs. Patterns match
whole segments, case-insensitively, and tolerate one leading
segment, because getPathInfo() keeps a language prefix: inbound
path processing rewrites the path Drupal routes on, not the
request.
Both lists are on the settings form as well, one entry per line,
and emptying both switches the whole thing off: no query bag
walked, no referer parsed, and the URI passed through as
getUri() built it.
Config rather than constants, and no hook over it. The lists are
a site's to edit, including to shorten: code is shipped for
OAuth callbacks and is a country code elsewhere, and a site that
knows which it has should be able to say so. A module that serves
its own credential-bearing URLs adds its pattern to that config
in hook_install(), the way modules extend one another's
configuration, so there is one place to look and it exports with
the rest of the site. It appends to what
ForensicStamp::getPathPrefixesToRedact() returns, because a site
whose key is absent has the shipped patterns in force without
having stored them, and appending to the raw key there would
store a list holding nothing else.
Only the two fields the framework itself observes are masked.
Payload a caller passes is its own: a message placeholder, a
link, or a contributor's data is recorded as given, so a module
that logs a credential of its own is the one that has to stop.
Masking happens at write time only, and it has to:
context_transient_hash signs the transient bytes, so rewriting
a stored value would make the row report itself as tampered.
Values captured before the module started masking them age out
through the transient purge, which NULLs the column, a state the
verifier accepts.
Where nothing matches, which is nearly every write, the URI is
recorded exactly as the request carried it: the guard is one hash
lookup per query parameter and one route lookup, and the URI is
not rebuilt at all, so an encoded ?destination= parameter stays
encoded exactly as it arrived.
And it runs once per request rather than once per row. One
request writes as many rows as it has bridges and entities, and
the masked value is the same for all of them, so it is cached
on the request and its route, the route included because it
arrives late and a row written before routing must not answer for
one written after. Measured, per additional row: 0.12
microseconds, against the 1.2 to 4.8 the getUri() behind the
unmasked field cost on every row before this.
Snapshot-delta bucket format¶
Contributors that snapshot the state of an audited subject on each event (entity create/update/delete, webdav lock/copy/move, etc.) historically emitted dual snapshots under two top-level bucket keys:
{"before": { … }, "after": { … }}
The bucket is normalized on the write path into a compact self-describing shape that keeps every field value at most once and carries diff hints alongside the current state:
{
"_v": 1,
"state": {
"extra": "x",
"status": 1,
"tags": ["a", "b"],
"title": "New"
},
"delta": {
"new": ["extra"],
"original": {
"title": "Old",
"old_field": "old_value"
}
},
"key_order": ["title", "status", "old_field", "tags", "extra"]
}
_vis the wire-format version. Readers consult it before trusting the rest of the shape so future format evolutions can ship without rewriting existing rows.stateis the current state of every field after the event. Updated fields keep their post-event value here; newly-added fields too; unchanged fields just sit alongside with no annotation.delta.newis the list of field names that appeared for the first time on this event. Values for those fields live instate: the list is names-only.delta.originalis the sparse map of previous values for fields whose before-value differed fromstate. For updated fields, the current value lives instate; for removed fields, the key is absent fromstateentirely and the value lives only here.key_orderis the merged before/after iteration order :statekeys plus any removed-field keys inserted at their natural before-position. Survives canonical JSON encoding because arrays preserve order; object keys would otherwise be alphabetized at storage time.
What counts as changed is one rule, in
Snapshot\SnapshotDelta::hasValueChanged(), and a reader
reconstructing a diff from a row needs it. Where both sides are
strings they are compared byte for byte, so "01" and "1" are
two values. Anything else is compared loosely, because Drupal
field storage hands back one logical value with a different
scalar type across a read/write round trip and reporting that
would drown real edits in noise: a field holding 1 that comes
back as "1" is not in delta.original, and a row that carries
it reads as unchanged because nothing was edited. Arrays are
walked by key set, since canonical JSON sorts keys at every
depth and the chain therefore records no order inside a value;
position still counts in a list, where the position is the key.
The same rule decides whether a row is written at all when
skip_no_op_updates is on for a bundle, so the check that
skips a row and the diff that renders one cannot disagree
about what the same value is.
Write-path overview:
flowchart LR
CB["Contributor returns<br/>{before, after, ...}"]
EB["AuditTrail::record()<br/>snapshot fold"]
V{has _v?}
SC["SnapshotDelta::<br/>computeStateWithDelta()"]
JSON["canonical JSON<br/>{_v, state, delta, key_order, ...}<br/>: context_*"]
CB --> EB
EB --> V
V -->|yes,<br/>opt-out| JSON
V -->|no| SC
SC --> JSON
Field classification at render time, given the shape above:
| Field is | Classification |
|---|---|
In delta.new AND state |
added |
In delta.original AND state |
changed (before in delta.original, after in state) |
In delta.original only |
removed (value in delta.original) |
In state only |
unchanged |
Per-action semantics:
- Create: emitted as
stateonly (nodelta). The row'saction=createcolumn implies "everything is new". - Delete: emitted as
stateonly, carrying the pre-delete snapshot. The row'saction=deletecolumn implies "everything was removed". - Update: emitted as
state+deltablocks for whatever changed.
Equality¶
Field comparison uses loose !=. Drupal field storage
routinely surfaces purely cosmetic scalar-type drift
('10000.00' vs 10000, '1' vs 1) across read/write
round-trips that strict comparison would flag as a real
change. PHP 8+ tightened loose equality enough that
0 == '' returns FALSE, so the historical type-juggling
gotchas don't apply.
Contributors that need domain-specific equality semantics
(fuzzy floats, whitespace-normalized strings, EXIF-stripped
image bytes) pre-compute the snapshot-delta shape themselves
and emit the resulting _v / state / delta /
key_order keys directly. The framework's auto-fold leaves
any bucket already carrying a _v key untouched.
Contributor contract¶
The simplest path: emit before and/or after as top-level
peers in a bucket. The framework's
SnapshotDelta::computeStateWithDelta() folds them into the
canonical shape at write time, in AuditTrail::record(), which is
also the only path contributors run on. Existing contributors that
already use the dual-snapshot pattern get the new wire format with
zero plugin code change.
Contributors that need stricter equality opt out by emitting
the canonical shape directly (with _v at the top); the
auto-fold leaves their bucket alone.
The HMAC secret¶
Secrets are first-class config entities: each
audit_trail_secret entity carries an integer secret_id
matching the per-row secret_id, computed from the bytes the
Key holds so it can never point at other bytes, a key_id
referencing a drupal/key Key entity holding the actual
bytes, and a lifecycle status (pending / active /
retired).
Why Key-backed: the module never stores the secret bytes in its own storage. The Key module's provider plugins decide where bytes live: config storage, environment variable, file outside the webroot, AWS Secrets Manager, HashiCorp Vault, HSM-backed providers, etc., and the module dispatches through that abstraction at write and verify time.
Changing the signing secret¶
The operator provisions a Key; the module points a secret at it
and puts that secret in the signing role. It never handles the
bytes. Either SecretRepositoryInterface::setSigningSecret()
(reached from the admin form) or
drush audit_trail:set-signing-secret --key=<key_id>, which also
creates the audit_trail_secret record when there is not one, so
the whole operation is available from a shell. The repository:
- Saves the named secret as active.
- Iterates other active entities and retires them.
The order is deliberate: a crash between the two saves leaves
two active entities (both with valid Key bytes, so fresh
writes still succeed) rather than zero, which would halt every
chained write. getSigningSecretId() picks the active secret
activated last when more than one exists, so writes land on the
new secret even in that transient window.
Because writes carry on, the operation looks finished when it is not, so the status report raises an error while more than one secret is active. Operators converge by re-running the command or by retiring the leftover directly.
Rotation does not rewrite existing rows: every row carries the
secret_id it was signed under, and the verifier dispatches
per-row. A chain can span any number of rotated secrets
without re-signing.
Drop-in via the logger service tag¶
AuditTrailLogger implements LoggerInterface and is
registered with the logger service tag. Drupal's
logger.factory collects every tagged logger and dispatches
each log call to all of them: exactly the mechanism dblog
and syslog use. No bespoke API, no replacement of core
services.
Consumers do not need to know audit_trail exists. A call
like:
\Drupal::logger('finance')->notice('Acte signed', [
'audit_trail' => [
'chain' => TRUE,
'action' => 'state_change',
'resource' => 'node/' . $nid,
],
]);
works whether audit_trail is installed or not. With the
module installed and the entry's context flag set, the call
additionally lands in the chain. Without the module, only
dblog / syslog record it.
The canonical write path for modules that audit their own
business events is the orchestrator service
AuditTrailInterface::record(). It pre-buckets context across
the two retention tiers via the ContextContributor plugin
pipeline, then writes the row through the chain writer and
makes one ordinary log call for the other sinks, marked as
already chained so this module's own logger does not record it
a second time. Buggy
contributor plugins cannot cascade into the caller: every
applies() / contribute() call is wrapped in try/catch and
exceptions are surfaced to the audit_trail logger channel
without aborting the event.
Hot-path resolution cache¶
Every PSR-3 log call that reaches Drupal's logger pipeline
hits AuditTrailLogger::log(). The vast majority of those
calls are NOT meant to chain (they have no chain key in
context). To keep that fast path cheap, the logger lazy-builds
a per-request chain registry on the first log() call and
reuses it for every subsequent call.
The logger itself is built whenever logger.factory is, since
the factory builds every logger-tagged service, and that is
nearly every request Drupal handles in full. So it takes the
chain writer, the forensic stamp and the chain filters as
closures, built on the first entry a chain claims: the writer
alone brings the secret repository, the lock and the chain
repository with it. A request that logs nothing onto a chain
builds none of them.
The registry holds:
entities: every activeaudit_trail_chainentity, indexed by id.channel_claim: PSR-3 channel to chain id (first-match wins, deterministic via ksort).auto_channels: subset ofchannel_claimrestricted to channels that resolve to amode: autochain. Drives the implicit-write short-circuit. Thedefaultchain contributes its OWN claimed channels (and its id, via channel-as-id); it does NOT auto-claim every other channel via fallback, otherwise adefault: mode: autoconfiguration would chain every log call on the site, including unrelated dblog noise like PHP deprecations.
The implicit-write short-circuit:
log(level, msg, context):
statement = the chain named in context.audit_trail
if statement === FALSE: drop
if no statement:
if not in auto_channels[$channel]: drop
if context._audit_trail_already_chained: drop
chain = registry.resolve(channel, id named in context)
if no chain: drop, and report it if the caller named one
if not opted in and chain is not mode: auto: drop
if one of the chain's filters votes no: drop
write
Both maps are built from the chain configuration objects,
not from the entities built on top of them: status,
channels, mode and the id are all on the stored object, and
constructing an entity to read them costs more than reading
them. The entities are loaded separately and only when a chain
actually has to be returned, which is to say only for a log
call that survived the short-circuit. On a site with zero
mode: auto chains, no chain entity is ever constructed to
answer a log call.
Within a request the first call builds the maps and every call after it is two array reads. Neither the maps nor the entities are cached beyond the request: a stale channel map would mean entries silently not chained, and that is the failure this module exists to prevent, so it is not a risk worth taking to avoid a read the configuration system already serves from its own cache.
Lifecycle: the registry follows Drupal's static config-entity
cache. Saving a chain entity invalidates via the
AuditTrailRegistryHooks hook so the next log() call in
the same request picks up the change.