Configuration¶
Three things to configure:
- Chain entities:
audit_trail_chainconfig entities that declare which PSR-3 channels chain, what retention windows apply, and which context contributors run. - Secret entities:
audit_trail_secretconfig entities that referencedrupal/keyKey entities holding the actual HMAC bytes. - Settings:
audit_trail.settingsglobal tunables for the auto-archive lifecycle and auto-verify cron.
Chain entities¶
Each chain is a audit_trail_chain config entity. CRUD through
/admin/config/system/audit-trail/chains. The default chain
ships in config/install/ so a fresh module enables with a
working chain in place. While it exists, default also acts as
the catch-all for AuditTrailInterface::record(): an
record() on a channel no chain claims by id or channels[]
lands in default rather than falling through to plain dblog,
which is why entries can appear in default for channels you
never configured. But default is an ordinary, removable chain:
if it is disabled or deleted, record() on an unclaimed channel
resolves no chain and flows out to dblog / syslog instead. A
channel claimed by a chain that exists and is inactive does
not reach the catch-all either: that chain holds its channels,
so the entry is refused rather than filed under default. See
consumers.md.
Per-chain settings:
mode, which entries chain¶
mode governs the generic logger path only: what
happens to a plain \Drupal::logger() call on a channel this
chain claims. It does not gate AuditTrailInterface::record():
once a chain resolves for the channel, record() writes the row
regardless of that chain's mode (see
consumers.md).
Two values, both respecting an explicit
'audit_trail' => FALSE in the PSR-3 context as a universal
opt-out:
flag(default): only logger calls that pass'audit_trail' => TRUEin their context land in the chain. Everything else stays a normal log entry, reachingdblog/syslogwithout chaining. Use this when you want explicit opt-in per call. (This opt-out applies to the logger path;record()calls still chain.)auto: every entry on a channel claimed by this chain chains automatically. The chain claims a channel either by matching its id (e.g. thewebdavchain claims thewebdavPSR-3 channel) or via itschannels[]list. Amode: autochain on thedefaultid does NOT wildcard every other channel: it claims only the channels it explicitly names, so unrelated dblog noise (PHP deprecations, core notices) stays out.
channels[]: additional channels routed into this chain¶
By default each PSR-3 channel chains to its own
same-id chain. channels[] lets you funnel several channels
into one chain, for installations that want a single
auditable sequence across related subsystems.
# config/install/audit_trail.chain.notarial.yml
status: true
id: notarial
label: 'Notarial workflow'
mode: auto
channels:
- webdav
- finance
- workflow
With the above, a save by webdav and a node update by
finance land in the same notarial chain in id (≈
arrival) order, and verifying notarial walks the whole
sequence.
Removing a channel from channels[] does not retroactively
move or rewrite rows that already chained under it. Those rows
keep their chain column pointing at this entity, stay
verifiable, and remain visible on the entries list. Only
future entries on the removed channel stop landing here:
they flow through to whatever chain claims that channel next:
the default catch-all when nothing else does, and plain
dblog when no default chain is active either. To wipe
historical entries on a channel, use Clear chain data on the
chains list, which keeps the chain and its overrides, or delete
the chain from the same list, which removes its configuration as
well. Both wipe the rows, segments, acknowledgments, checkpoints,
any staged entries the trail has given up on, and on-disk archive
files tagged with that chain id.
Either can refuse, and it is worth knowing why rather than retrying blindly. An entry written inside a caller's open transaction waits in the staging table until that transaction resolves, and the request that made it then records it from its own memory. Deleting such an entry out from under that request would not stop it being recorded: it would land on the chain that was just emptied, as a perfectly valid first entry, with nothing to distinguish it from a new one. So a clear declines while any entry is still waiting to be recorded, which is the status report's phrase for the same entries, and names how many. The remedy is to try again in a moment: they are recorded within seconds of the request that made them, or by the next cron run. Entries the trail has given up on do not hold a clear up, since nothing is coming to record those and nothing else removes them.
chain_only: keep the chain's events off the other loggers¶
A chain records through audit_trail.logger like any other,
and the same event reaches dblog and syslog as a normal log
entry. Check Log to the audit chain only on the chain form,
or set chain_only: true in the chain's configuration, and
the other loggers are left out.
# config/install/audit_trail.chain.notarial.yml
status: true
id: notarial
label: 'Notarial workflow'
chain_only: true
Worth setting when the entries carry content the chain is meant to keep tamper-evident rather than readable: the dblog tail is visible to anyone with access site reports, and its rows age out on dblog's own row cap instead of on the chain's retention windows.
What it reaches is the module's own write API. An event written
through AuditTrail::record() is dispatched by this module, so
it can hold the other loggers back. An event logged straight
through a PSR-3 channel cannot be held back at all:
\Drupal::logger($channel) hands the record to every
registered logger in one pass, dblog included, and this
module's logger is one of those registered loggers rather than
the one deciding the order. By the time the setting could be
read, dblog already has the row. So a chain that must stay off
the other loggers has to be written to through record(), and
a chain fed by plain PSR-3 calls keeps a dblog row per entry
whatever this setting says.
A caller can ask for the same thing one event at a time,
without the chain being configured for it, by passing
chain_only: TRUE to AuditTrail::record(). An operator's
filter does not override that; the caller's
intent wins. If no chain is accepting writes for that channel,
such an event is recorded nowhere at all, since the other loggers
are what it rules out, and a warning on the audit_trail channel
names the channel it came from.
Filtering on a chain-only chain covers the other loggers whatever the filter's own Keep the event out of dblog and syslog too setting says, on the path where the setting applies: there is no other logger left for it to reach.
Per-chain retention overrides¶
Each chain inherits the global cron_archive retention
windows from audit_trail.settings and can override any of
them:
The five thresholds have to increase strictly, in the order the
stages run: transient_purge_after < archive_after <
live_purge_after < file_purge_after < compact_after. Both forms
refuse anything else, and the cron worker re-checks the rule for
the configuration imports and drush config:set calls that never
saw a form. What it does about a breach depends on which one. The
first three order stages that destroy something, so a chain
breaking one of them is skipped whole and the reason is logged. A
compact_after inside the file-purge window refuses the
compaction stage alone: it destroys nothing file-purge has not
already destroyed, so the rest of that chain, the erasure stages
included, carries on. Listing them in that order makes the rule
visible:
transient_purge_after: P7D # signed-purge the transient bucket first, before an archive freezes it
archive_after: P30D # archive after 30 days
live_purge_after: P75Y # keep archived rows queryable for 75 years (notarial)
file_purge_after: P100Y # effectively never: has to exceed live_purge_after
compact_after: P101Y # optional: fold contiguous file-purged segments into one bridging row
segment_granularity: month # hour | day | week | month: one WORM archive per chain per bucket
transient_purge_after and compact_after both ship empty, so
each stage runs only for the chains that set one (or for every
chain, once the global form does). shortened_original_length is
the seventh value a chain can restate; it is a write-time cap
rather than a lifecycle window, and
its own section
says what it does. The chains list reports all seven under
Overrides, so a chain deviating from site policy in any of
them is visible without opening its form.
Why ISO 8601 rather than seconds: durations on the scale of
years are mis-typed easily as raw seconds. P75Y is
unambiguous, future-readable, and matches what regulatory
documents use.
contributors[]: context contributor pipeline¶
Each chain runs its contributor plugins in ascending weight
order at write time. Each contributor inspects the subject
plus the caller context and returns a two-bucket payload
(permanent / transient) that the orchestrator merges per
tier. Higher-weight contributors can overwrite lower-weight
ones key-by-key.
Plugins discoverable under
<module>/src/Plugin/ContextContributor/ and tagged with the
#[ContextContributor] attribute. Enable them via the chain
edit form; per-instance settings render in-form for plugins
that declare them.
Buggy contributor plugins cannot cascade into the entity-save
hook that triggered the audit event. Every applies() /
contribute() call is wrapped in try/catch; throws surface
to the audit_trail logger channel with audit_trail: FALSE
and the row still lands minus the offending contribution.
Bundled: paragraph ancestry¶
audit_trail_entity_paragraphs ships the only bundled
contributor with a setting. It walks a paragraph's parents to
the host entity and stamps the path, so an audited change to a
nested paragraph records what it was nested in.
- Maximum ancestry depth to walk (
depth_cap, default 32, range 1 to 256): how many parent hops the walk follows before it stops. It doubles as the cycle guard, since a corrupt parent pointer would otherwise loop until this cap. The ceiling exists because the point of the setting is to bound the work one audit event can cause. A value below 1 is read as 1.
filters[]: per-chain filter pipeline¶
Filter plugins run before contributors and decide whether an event is even allowed to reach the chain. Useful when a channel is almost what you want to chain but a few noisy verbs / actions don't carry audit signal.
Bundled: request_method¶
Skips events based on the current HTTP request method. Configurable per-instance:
- Mode (
allow/disallow): allow: only events whose request method is one of the checked verbs pass through; everything else is skipped. Useful when the legitimate-traffic shape is small and known (e.g. onaudit_trail_file, a GET is the one method that serves a private file; a HEAD answers with the headers and no body, and everything else reaching that hook is protocol housekeeping).disallow: events whose method is in the list are skipped; everything else passes through. Useful when the noise pattern is small and known (e.g. quiet PROPFIND / LOCK / UNLOCK on a chain that also audits regular file activity).- Request methods: checkbox list covering standard HTTP
- WebDAV verbs (GET, HEAD, POST, PUT, PATCH, DELETE,
OPTIONS, PROPFIND, PROPPATCH, MKCOL, COPY, MOVE, LOCK,
UNLOCK). Custom verbs already in saved config (e.g.
REPORTfor CalDAV) appear as already-checked options so they can be unchecked through the UI. - Channel allow-list: scope the filter to specific channels (one per line). Empty applies to every channel the chain handles. Lets one chain hold filters whose policies differ per channel.
- Action allow-list: scope the filter to specific
actions (one per line, e.g.
file_downloaded). Empty applies to every action. - Keep the event out of dblog and syslog too: when checked,
an event this filter drops is kept out of the other loggers as
well (no chain row AND no dblog row). Default (unchecked) lets
it still reach dblog, as on any chain that is not chain-only,
so operators can see what got filtered. How far that
reaches depends on the writing API:
AuditTrail::record()dispatches to the other loggers itself and can hold them back, while an event logged straight through a PSR-3 channel has reached dblog before this module's logger is called at all, so for those the setting only skips the chain row.
Empty methods list is a no-op in both modes: events pass
through. CLI / cron requests always pass through, regardless
of mode (there's no HTTP method to match against).
Adding your own¶
Custom filters: ship a class extending AuditTrailFilterBase
under src/Plugin/AuditTrailFilter/ annotated with
#[AuditTrailFilter]. Drupal's plugin discovery finds them;
they show up in the chain edit form under the filter list.
See the bundled RequestMethodFilter for the reference
shape (shouldRecord() returns FALSE to skip;
buildConfigurationForm() provides the per-instance UI).
When a plugin's module is uninstalled¶
A chain names its contributors and filters by plugin id, and records which module provides each one. Uninstalling that module therefore has a visible effect, and which effect depends on whose chain it is.
- A bridge's own chain goes away with the bridge. The three
bridge submodules each ship a chain, and that chain belongs
to them, so uninstalling
audit_trail_entity,audit_trail_fileoraudit_trail_user_authremoves the matching chain and its settings. Installing the module again brings the shipped chain back in its default state; any edits made to it in between are gone. Export the chain first if you want to keep them. - Any other chain keeps everything except that plugin. A chain you built yourself that merely uses a contributor or a filter from the departing module loses that one component. Its channels, its retention settings and its other plugins are untouched, and rows already in the chain are not affected at all.
Neither case leaves a gap in the hash chain: nothing that was already recorded is altered, and the chain continues to verify. What changes is what gets recorded from then on: a chain that lost a contributor stamps less context than it used to, and a chain that lost a filter starts recording the events that filter was there to decide on.
When a plugin stops resolving without an uninstall¶
Uninstalling is the tidy case, because the chain is edited as the module leaves. A plugin can also stop existing while its module stays installed, which is what a module update does when it drops a plugin class. Nothing is edited then: the chain still names the id, and both write paths skip a component whose id no longer resolves rather than failing the write.
/admin/reports/status carries an entry for that, one per
affected chain, naming the ids it could not load, for
contributors and filters alike. Both are worth acting on, for
different reasons: a missing contributor costs every row a
payload it was configured to carry, and a missing filter changes
which events are recorded at all. Edit the chain to remove or
reconfigure the component the entry names.
Secret entities¶
Each audit_trail_secret config entity carries:
secret_id: integer matching the per-rowsecret_idcolumn, computed from the bytes the Key holds: the first four bytes ofHMAC-SHA-256(secret, "audit_trail secret id"), read as an unsigned 32-bit integer and shown as eight hexadecimal digits. An id is what a signed record names to say which bytes signed it, and an id computed from the bytes can never point at other ones. The same bytes give the same id under any Key, so a secret recreated from its Key after a configuration loss gets its id back, and the rows and archive files naming it verify again. A second secret over the same bytes is refused, since it would carry the same id.key_id: id of adrupal/keyKey entity holding the bytes. Editable while the secret ispending, which is the status that says it has never signed; the secret then takes the id of the new bytes. Past that, a different Key is accepted only when it gives the same id, which means it holds the same bytes. That is what makes a provider migration (configuration to file, file to KMS) safe, and what lets a secret whosekey_ida partial restore left empty be given its Key back. Every record signed under the id would otherwise verify as tampered. To sign with different bytes, create a new secret and rotate to it.secret_status:pending/active/retired.created/activated/retiredtimestamps.
CRUD through
/admin/config/system/audit-trail/secrets. The module never
stores secret bytes itself; the drupal/key Key module's
providers decide where bytes live.
Provisioning a fresh secret¶
- Provision a Key entity holding 32 bytes of CSPRNG output via your chosen Key provider (config provider for testing, file / env / cloud-managed for production).
- Create an
audit_trail_secretentity referencing that Key. - Choose "Use for signing" on the secret list page (or
SecretRepositoryInterface::setSigningSecret($id)from PHP).
The active secret is what getSigningSecretId() returns to
the logger for fresh writes. The verifier dispatches per-row
to the row's stored secret_id, so a single chain can span
any number of rotated secrets.
Changing the signing secret¶
From a shell, once the Key exists, it is one command:
drush audit_trail:set-signing-secret --key=<key_id>
--key is required. It creates the audit_trail_secret record
for that Key if there is not one already, makes that secret the
signing secret, and retires the outgoing one. A Key that already
backs a secret is reused when that secret is pending, and
refused when it is active (nothing to change) or retired
(material taken out of service is never put back).
Through the UI it is two steps: provision the Key, then add a secret referencing it from the secrets list and choose Use for signing.
Either way the repository saves the new secret active first,
then retires the one it replaces. That order is deliberate: a
crash between the two saves leaves two active entities rather
than zero, and zero would halt every chained write.
getSigningSecretId() returns the active secret activated last
when more than one exists, so fresh writes land on the new
secret even in that window.
Writes surviving is not the same as the operation having finished, though, so the status report raises an error while more than one secret is active: the change looks complete, the list shows two active rows, and a retirement is still owed. Converge by re-running the command for the Key backing the signing secret, or by retiring the leftover directly.
Retiring without replacement¶
The "Retire" operation on the secrets list flips an entity to retired status without promoting a new active. This is the emergency-isolation primitive: chained writes immediately start failing with the no-active-secret diagnostic, which is the loud-fail desired state in a key-compromise incident.
Key provider trade-offs¶
| Provider | Material location | Operator effort | Survives DB-read compromise? |
|---|---|---|---|
config |
Drupal DB (key_config_override) |
Zero (the Key form generates and stores) | No |
| File | Outside the webroot, owned by a separate UID | Provision per rotation | Yes if FS permissions are correct |
| Environment variable | Process memory only | Set in deploy / systemd / k8s | Yes |
| AWS Secrets Manager / GCP Secret Manager / Azure Key Vault | Managed cloud service | Provision in the cloud console | Yes |
| HashiCorp Vault | Vault | Provision via Vault CLI | Yes |
| HSM-backed providers | HSM | HSM ceremony | Yes |
audit_trail.settings: global tunables¶
cron_archive: automatic retention lifecycle¶
When enabled, a cron-driven worker progresses chain rows through their lifecycle automatically. The lifecycle is deterministic: every row passes through the same stages as it ages past the configured thresholds.
cron_archive:
enabled: false # off by default: opt in explicitly
cron_interval_seconds: 3600 # minimum gap between cron runs per chain
transient_purge_after: P3D # signed-purge the transient column at this age (empty = disabled); must come before archive_after
archive_after: P7D # rows older than this get archived
live_purge_after: P3M # archived rows older than this lose their live copy
file_purge_after: P2Y # archived files older than this get unlinked
compact_after: null # off by default: compact contiguous file-purged segments at this age; must come after file_purge_after
segment_granularity: week # hour | day | week | month: one archive per chain per bucket; `hour` exists for staging/test workflows
The duration thresholds all clock from each row's own
created timestamp. The settings form refuses any
configuration that breaks the archive_after <
live_purge_after < file_purge_after invariant.
transient_purge_after is optional: when empty (per-chain or
globally) the transient-purge cron run simply doesn't run
for that chain. compact_after is optional the same way and
ships empty, which disables the compaction stage: it folds
contiguous runs of file-purged segments into one bridging row,
so the bookkeeping tail stops growing on an install that has
been purging files for years. Set it later than
file_purge_after, since a segment has to reach that stage
before there is anything to fold.
in_transaction_write_mode: where an entry waits during a save¶
An audit entry is often written in the middle of a larger
change: a hook firing while an entity is being saved, a
consumer writing inside its own transaction. Such an entry
cannot be chained where it happens. The head a writer reads for
previous_hash comes from the caller's transaction snapshot,
so two overlapping callers read the same head however the
per-chain lock is scheduled, and the unique key on
(chain, previous_hash) then refuses the second row. The
chain stays intact; the entry is lost.
So the entry is buffered and chained once the caller's outermost transaction resolves. This setting says where it waits in the meantime.
in_transaction_write_mode: outbox # outbox | memory | inline
keep_rolled_back_writes: false # memory only
| Value | Where the entry waits | What it costs |
|---|---|---|
outbox (default) |
The audit_trail_outbox table, written inside the caller's transaction |
One extra INSERT and one DELETE per entry. Atomic with the change it describes, and durable, so a crash between the commit and the flush leaves it for cron |
memory |
The request that wrote it | Nothing. A fatal before the flush loses the entries, and nothing was written inside the caller's transaction that could have aborted it |
inline |
Nowhere: chained immediately | Entries lost under concurrency, per the defect above. Kept for a consumer that needs a failed chain write to abort its own transaction |
Two things are unaffected by all three values. An entry written outside any transaction is chained immediately, as it always was. And the module's own lifecycle writes hold the per-chain lock across their own transaction, so they are already serialized and are never buffered.
Buffering makes a batch cheaper, not dearer: the per-chain lock and the head read are paid once for a whole transaction rather than once per entry, which is most of what a chain write spends its time on.
keep_rolled_back_writes only applies to memory, because it
is the only value where the choice exists: outbox entries
roll back with the transaction that staged them, and inline
entries were never buffered. Off by default. When on, entries
buffered by a transaction that rolled back are chained anyway,
each marked _audit_trail_rolled_back in its permanent bucket
so the chain records the attempt rather than asserting a
transition that never happened.
Two entries appear on /admin/reports/status when they are
non-zero: the number of entries still waiting in the outbox
(normally zero, since an entry waits microseconds; a climbing
number means cron is not running or the flush is failing), and,
on inline only, the number of entries recorded during a save,
which is what makes the dropped-entry counter structural rather
than a capacity problem.
Rolling a savepoint back¶
A caller that opens a savepoint, writes audit entries inside it, rolls that savepoint back and then commits its outer transaction is the one case the transparent path cannot get right on its own. The database layer reports the outcome of the outermost transaction only, and once a savepoint is gone a released one and a rolled-back one look identical, so the entries written inside it would be chained at the outer commit and describe work that was undone.
Only a caller that does this has anything to do, and the API is two calls:
$marker = $this->auditTrail->markPendingWrites();
$savepoint = $this->database->startTransaction();
try {
// ... work that writes audit entries ...
}
catch (\Throwable $e) {
$savepoint->rollBack();
$this->auditTrail->discardPendingWritesSince($marker);
}
shortened_original_length: what a cut value leaves behind¶
An entry holds 64 characters of the channel it came from, 64 of
the action and 255 of the thing acted on. A longer value is cut
to fit rather than refused, because refusing it would take down
the operation the entry is recording. All three columns are
covered by the row's hash, so the value that gets signed is the
cut one: the row verifies perfectly while naming something the
caller never sent, and for resource, the column that says what
the event was about, nothing else remembers what it named. A
WebDAV path over 255 characters is cut, signed, and the record
now names a different file.
So the writer keeps what it had to cut, in the row's permanent bucket, under a key of its own:
{"_audit_trail_shortened": {"resource": "webdav:/the/full/original/path/..."}}
Keyed by column, and written only for a column that overflowed:
a row where everything fit carries no such key, so a site where
nothing is ever too long pays nothing at all. It is signed with
the rest of the permanent bucket, never purged, and goes into the
archive, which is exactly why it is capped:
shortened_original_length is how many characters of the
original are kept, and without it this would move the unbounded
string one column over, out of a varchar and into a text.
The shipped default is 1024, deep enough for a long file path.
0 keeps nothing, which is what the module did before the
setting existed.
A column is only recorded when the cap exceeds that column's own
width, since a prefix of what the row already carries says
nothing new. Both forms refuse anything from 1 to 255 on that
basis: a form is filled in long before anyone knows which column
a future entry will overflow, and below 255 nothing at all is
recovered for resource, which is the column the setting is
mostly for.
A single chain can state its own figure, on its edit form or as
shortened_original_length on the chain entity. Empty there
inherits the site default; 0 switches it off for that chain
alone.
redact_query_keys / redact_path_prefixes¶
Every entry records the URL the request came from, and a
one-time login link, a password-reset URL or a signed callback
carries a working credential in it. These two lists say which
query parameters and which URL shapes are masked before an entry
is written: the first names parameters, matched case-insensitively
and including nested ones such as filter[token], the second
names path patterns, for the credentials Drupal puts in the path
itself. Both are on the settings form under Credentials in
recorded URLs, one entry per line, and both are a site's to
edit.
Emptying both switches masking off entirely. A key left out altogether, which is what an install predating these settings has, falls back to the shipped list instead, and the form offers that list rather than an empty field.
security.md
has the rest: what ships in each list, why removing a name is a
deliberate decision, and what a module serving its own
credential-bearing URLs adds in hook_install().
auto_verify_enabled / auto_verify_max_age_hours¶
Cron-driven incremental verification. When enabled, each
cron run walks every chain via verifyChainIncremental()
and writes the results to State. The runtime status report
(/admin/reports/status) surfaces stale runs (cron hasn't
verified in N hours) and broken chains as
RequirementSeverity::Warning / RequirementSeverity::Error.
auto_verify_max_age_hours is how old the last stored result may
get before that page turns yellow. Staleness is read from the
timestamp of the last cron run alone and never from its verdict,
so it answers "has anything verified lately?" rather than "did it
verify clean?". A broken chain is reported as an error, and that
error is returned before the staleness check is reached, so the
yellow warning is what a site with clean chains and a stuck cron
sees. Default 24; raise it only if cron genuinely runs less
often than that, because a threshold longer than the gap it is
meant to notice reports nothing.
auto_verify_full_walk_every_days¶
How often a cron tick does a full walk instead of an incremental one. Every other tick resumes from the chain's last signed checkpoint; once per this interval the next tick re-walks every chain from genesis, re-deriving every hash and HMAC rather than trusting the checkpoint it would otherwise start from.
That is what catches a tampered row behind the checkpoint, and
a checkpoint whose signature is valid but whose chain has been
rewritten under it. Default 7. 0 disables the periodic full
walk entirely, which leaves nothing re-reading the rows an
incremental walk has already passed: set it only where a full
walk is driven some other way, such as
drush audit_trail:verify --full on a schedule of its own.
A break the full walk finds stays reported. The incremental walks that follow do not read the rows below the checkpoint, so each of them reports its own result and adds that the last full walk found those rows broken. That part goes when a later full walk finds the chain clean, typically the first one after the range has been acknowledged, or when an incremental walk that started below the broken rows does, as it does from genesis on a chain with no checkpoint it can resume from.
A full walk run by hand counts the same as cron's:
drush audit_trail:verify --full and a full verify from the
chains list record a break they find, clear one they no longer
find, and log the change. So an operator who acknowledges a range
can verify the chain straight away rather than wait for cron's
next full walk, and a site whose full walks all run from drush,
with this setting at 0, gets its breaks recorded and cleared by
them. Only a walk at the default --depth=strict counts: the other
depths read fewer signatures.
The cost is one full walk of every chain per interval, which is the cost the checkpoints exist to avoid the rest of the time. On a multi-million-row chain, lengthen the interval rather than switching it off.
checkpoint_min_rows / checkpoint_retain¶
The two tunables on checkpoint minting, described in verification.md § "Performance: incremental verification + checkpoints".
checkpoint_min_rows is the throttle: a clean incremental walk
mints a new checkpoint only once at least this many rows have
accumulated since the last one. Without it, every cron tick that
found a single new row would mint one, and the checkpoint table
would grow nearly as long as the chain it indexes. Default
100. Lower it on a chain that writes rarely and whose walks
should still resume close to the head; raise it on a busy one.
checkpoint_retain is the cap: each successful mint prunes
older checkpoints past this count. Only the most recent one is
ever read at verify time, so the rest are history an operator
can inspect. Default 50.
Neither accepts 0. The settings form refuses it on both fields
and the schema constrains both to a minimum of 1, because a
zero throttle mints on every tick and a zero retention makes the
prune that follows every mint ask the database for a negative
offset.
Lifecycle stages¶
| Stage transition | Trigger | What happens |
|---|---|---|
| Live to Archived | row.created + archive_after < now | Cron picks closed calendar buckets and calls ChainArchiver::archiveChainRange(), landing an NDJSON under <archive_directory>/<chain>/<year>/<YYYY-MM-DD>--<id>.ndjson. The audit_trail_segment bookkeeping row is HMAC-signed. |
| Archived to Live-purged | row.created + live_purge_after < now | Cron calls ChainArchiver::purgeSegmentLiveRows(). Live rows are deleted from audit_trail; the NDJSON archive remains. audit_trail_segment.live_purged_at gets stamped. |
| Live-purged to File-purged | row.created + file_purge_after < now | Cron calls ChainArchiver::purgeSegmentArchiveFile(). The NDJSON file is unlinked (after a SHA-256 sanity check that refuses to delete tampered files). The bookkeeping row stays with its anchor_before / anchor_after hashes so the verifier keeps bridging across the now-empty range. audit_trail_segment.file_purged_at gets stamped. |
| Transient column purge | row.created + transient_purge_after < now | Cron NULLs context_transient on eligible rows in place, creates a new audit_trail_segment row attesting the transition (transient_purged_at != 0), and emits a segment_transient_purged chain event. The chain still verifies because only context_transient_hash is signed; the new segment legitimizes the NULL via the segment-coverage rule in verifyTransientColumn. |
| File-purged to Compacted | row.created + compact_after < now, and off unless set | Cron calls ChainArchiver::compactFilePurgedSegments(). A contiguous run of file-purged segments is replaced by one row spanning it, taking anchor_before from the first and anchor_after from the last, so the verifier bridges the same id range with fewer rows. Contiguity is anchor equality (the earlier row's anchor_after is the later's anchor_before), not adjacent ids. A segment_compacted chain event lists every absorbed id in compacted_from. |
Each lifecycle transition writes a chained log row (action
= segment_archived / segment_live_purged /
segment_file_purged / segment_transient_purged /
segment_restored) back into the chain it affects, with
resource = 'segment:<id>' (segment_created carries
segment:<from>-<to>, its id being its own row's). The matching
segment row records
the event's audit_trail.id in its *_event_id column,
signed into lifecycle_hmac. The verifier walks the chain
and cross-checks both directions of the mutual reference, so
tampering with either the chain event or the segment row's
state flags surfaces as a verifier error.
Stage ordering: transient-purge runs before archive¶
On any cron tick that runs both stages, transient-purge always executes first. This is load-bearing: once a row is archived, its NDJSON file captures whatever bytes the transient column held at that moment, signed into the file's SHA-256. A later transient-purge can NULL the live row's column but cannot retroactively scrub the bytes that already sealed into the NDJSON.
The settings-form invariant (transient_purge_after <
archive_after) exists for this reason: if archive could win
the race, raw transient bytes would leak into long-retention
WORM archives. Operators should keep that ordering in mind
when picking thresholds: if transient purge needs to run at
least N days before archive, configure thresholds so the gap
exists.
Misconfigured chains skip and log¶
When audit_trail_chain config carries a malformed ISO 8601
duration (typo, hand-edited YAML, partial config import) or
a segment_granularity outside the accepted value set,
RetentionThresholds::resolve() returns NULL and cron skips that chain
for that tick: no archive, no purge, no transient purge.
The settings forms validate every duration at save time and
refuse invalid values, so the form path is safe. That now
holds with the lifecycle switched off as well: each field is
bound to its config key, so saving the form validates the
result against the schema, which holds these keys to a
duration whether the stage runs or not. A recipe is refused
the same way. A hand-edited YAML file and drush config:set
still bypass both. Whenever the skip fires, the cron worker
writes a WARNING-level entry to the
audit_trail logger channel naming the chain and the
admin-edit URL. CI / monitoring that alerts on
audit_trail-channel WARNINGs catches it; otherwise watch
the channel during cron windows.
Segments a stage keeps refusing¶
A stage that cannot process one segment logs the reason and moves to the next. That is deliberate: with the per-chain resume points, one segment nobody can archive no longer holds up retention for everything behind it. What it costs is that the skip leaves nothing an operator reads. The refusal is one line in a log of thousands, the settings page still shows the schedule that was configured, and the segment stays where it is.
The status report counts them, under Audit trail retention backlog. It re-runs no operation and diagnoses no cause: a segment a stage keeps refusing and one a stage has not reached are the same row from the outside, and telling them apart would mean performing the operation.
What makes the count meaningful is the deadline it measures against, and where that comes from. The schedule is the one the cron worker itself acts on, so a chain it refuses is never measured against a schedule nothing is following, and a stage it runs is never treated as switched off. On top of the threshold goes one granularity period, because the closed-bucket rule waits for a row's bucket to finish after the cutoff passes, which is the worst case the effective delays preview above shows. A segment still unprocessed at twice that has missed it with as much time again to spare, on any schedule and at any granularity. No fixed span is assumed: a chain that archives after two years is judged on two years, not on a number the module picked.
Only so many segments are examined per chain per stage, because each costs an acknowledgment lookup and the answer an operator acts on is the same at two hundred as at two thousand. When the count stops there it reads as "200 or more", so a truncated figure never looks like a finished one, and the list of chains and stages is bounded the same way the verdict messages are.
The entry is an error when what is stuck is a stage that destroys something, meaning transient-purge or file-purge, and a warning otherwise. A stalled archive costs table size and every row is still readable; a stalled transient-purge means context an operator asked to be dropped is still on disk, which is a promise the module made and did not keep.
Three things it does not count. An inactive chain is not visited at all: the worker asks which chains are active and pausing the lifecycle is what making one inactive is for, so reporting its segments would turn a deliberate choice into an error a fortnight later and keep it there. A stage switched off has no deadline to miss either, so a chain that never asked for transient-purge is not reported for failing to perform it. And an acknowledged range is stuck but not unexplained: some stalls cannot be repaired, and an acknowledgment is already this module's word for that, recorded as a chain event so the operator's reason is as tamper-evident as what it explains. That test runs through the acknowledgment repository, which re-checks each one's anchors against the chain, so a row inserted straight into the index cannot silence the report.
Where the cause is an index that no longer matches the chain,
drush audit_trail:reindex-segments --dry-run reports what
differs and rebuilding from the chain re-seals the signature.
Where it is a missing archive file, drush
audit_trail:rewrite-archive --id=N rebuilds it against the
digest recorded when the segment was archived.
A segment that cannot be repaired is acknowledged, never
deleted. Its lifecycle stamps are the evidence that a range was
removed lawfully, and a verification walk crosses that range on
the segment's signed anchor_before and anchor_after.
Deleting one leaves a hole nothing accounts for, the chain
reads as broken from there on, and no restore brings the bridge
back.
Chains cron will not run at all¶
The extreme of the same silence, and its own entry: Audit trail
retention configuration. RetentionThresholds::resolve() refuses a chain
whose durations do not parse, or whose granularity is not one of
the four, or whose thresholds are out of order in a way that would
have a stage destroy something before the stage it depends on has
run, and skips every stage on it for that tick. Nothing is archived, nothing is
purged and nothing is erased there for as long as the value stays
wrong.
The backlog count above cannot be the signal for this, and not because it was overlooked: with no threshold there is no deadline to measure a segment against, so a refused chain drops out of that count and reads exactly like a healthy one. It is reported separately for the same reason the outbox splits its two entries, too: the repair is one configuration value rather than a segment at a time, and once it is fixed the stages work through whatever accumulated on their own.
Both retention forms validate durations and refuse to save an
unreadable one, so a value that reaches this state was written by
a configuration import or drush config:set, which is the same
pair of paths named under Misconfigured chains skip and log
above. That log line is still written every tick; this entry is
what an operator sees without going looking for it.
Chains cron will not compact¶
The fourth ordering rule, and the one that does not take the chain
out of retention: Audit trail compaction threshold. A
compact_after that is not later than file_purge_after refuses
the compaction stage alone, every tick, and leaves the other five
running. Nothing is kept longer than configured and nothing is
erased later, so this is a warning rather than an error; what is
lost is the folding that keeps the segment index small, and
folding a run the moment its files go widens the range any later
acknowledgment has to speak for.
It arrives the way the entry above does, from a configuration
import or drush config:set, and the same repair fixes it: both
retention forms refuse the pairing, so saving the chain through
its own edit form is enough.
Bucket granularity¶
hour / day / week (ISO) / month selects how cron
slices time into archive buckets. A bucket is only archived
once it has fully closed (the calendar window has ended)
AND every row in it has aged past archive_after. Both
conditions must hold.
- hour: one archive per chain per UTC hour. Intended
for staging / test workflows where waiting a full day
for the first archive to materialize is impractical.
Production installs typically run with
dayor coarser. - day: best for high-volume chains; one archive per chain per UTC day.
- week: sensible default for moderate traffic; one archive per chain per ISO week (Monday 00:00 UTC to next Monday 00:00 UTC).
- month: best for very-low-volume chains where annual archives would be too coarse. Aligns to UTC calendar months.
Buckets align to UTC, not site timezone¶
All bucket boundaries are computed in UTC, regardless of the
Drupal site's system.date:timezone setting. A site in
Asia/Tokyo (UTC+9) running day granularity produces
archives that close at 09:00 local time, not local midnight.
This is intentional: UTC is the only reproducible boundary
across servers in different zones and a deployment that moves
hosts. Operators planning daily WORM exports should align
their off-host transfer cadence to UTC midnight, not local.
Granularity must not exceed the smallest applicable threshold¶
The closed-bucket rule has a non-obvious consequence: when
segment_granularity is larger than archive_after (or
transient_purge_after), the granularity dominates and the
threshold is effectively ignored.
Concrete example: archive_after = P3D and
segment_granularity = month on a row written May 1.
- The row is past
archive_afterby May 4. - The May bucket end is June 1 00:00 UTC.
- The bucket isn't closed until June 1, and won't be archive-eligible until June 4 (June 1 + 3-day cutoff).
So the row sits in the live table until June 4: ~34 days
after writing, not the 3 days the threshold suggested. Same
effect on transient_purge_after: the transient column lives
in the DB up to one granularity period longer than the
threshold value declares.
This matters for GDPR commitments. If transient_purge_after
exists because raw bytes can't legally linger beyond N days,
pairing it with a segment_granularity larger than N silently
breaks the commitment. The chains forms warn on this case at
save time; production deployments should pick granularity ≤
the smallest threshold that has a compliance constraint.
Switching granularity mid-chain¶
Changing the value on a chain that already has archives:
existing segments keep their original bucket size and stay
exactly as written. Newly-eligible rows after the change
bucket at the new granularity. No retroactive re-bucketing
happens (segments are immutable post-write). Operators
switching from week to day on a long-running chain will
see weekly archives in their old WORM exports and daily
archives in new ones, both verify, the bucket distinction
is purely operational.
Manual operation¶
Auto-archive can be triggered out of band:
drush audit_trail:auto-archive --uri=https://example.com
drush audit_trail:auto-archive --chain=default --uri=https://example.com
Runs the lifecycle immediately (coverage, then transient-purge,
archive, live-purge, file-purge, and compaction where the chain
configures one), bypassing the per-chain throttle that
hook_cron honors and without stamping
last_run_at (so the next scheduled tick fires on its own
cadence). Honors every other config knob and per-chain
retention override. Useful for smoke-testing after editing
thresholds or for moving a specific chain through the
lifecycle on demand. The manual
drush audit_trail:archive, audit_trail:purge, and
audit_trail:archive-restore commands remain available for
range-specific operations the cron flow doesn't cover.
A full walk of every chain (cold verification from genesis on
each chain) can be triggered on demand via
/admin/reports/audit-trail/verify-all-full. On a
multi-million-row chain this is minutes-to-hours of CPU: the
trigger is gated behind the separate run audit trail
verification permission so read-only viewers can't accidentally
launch it.