Skip to content

Roadmap

Tracked here so the design discussions that fed into this module don't get lost between 1.0 and 2.0.

In the 1.0 line

  • Chained PSR-3 logger plugged into Drupal's standard logger pipeline via the logger service tag.
  • Per-chain audit_trail_chain config entities with mode: flag (default) / mode: auto selection, channel claims via channels[] list, per-chain retention overrides, and per-chain contributor pipeline configuration.
  • Two-tier retention model (context_permanent / context_transient) with hash-only signing of the transient column for GDPR-compatible purge. Segment-row attestation legitimizes the NULL'd column at verify time so attacker-NULL is distinguishable from operator purge.
  • Publicly-verifiable SHA-256 hash chain plus operator HMAC-SHA-256 layer per row.
  • Per-row contract version: every row records a version pinning the column set, the canonical form and the digest that produced its hash and keyed the hmac over it, inside the signed payload, and every reader takes it from the row. A chain can carry rows written under more than one primitive, and nothing is rewritten when one has to be replaced, which is what multi-decade retention needs.
  • Segment spine: a hash chain over archived segments, mirroring the one over rows, so a segment stays attested after its own lifecycle events have aged out and altering an old one means rewriting every head above it.
  • Versioned archive contract: version on the segment row and in the file's own footer, covered by the signature at both ends. Version 1 is NDJSON with a signed footer digested with SHA-256. A reader refuses a version it does not know, which is what lets a compressed, sharded or differently digested archive be added later without making every archive already written unreadable.
  • Multi-tamper detection in a single verifier walk: every contiguous broken range surfaces in verdict.broken_ranges without operators needing to ack-and-re-verify iteratively.
  • Schema-level chain-fork prevention via UNIQUE(chain, previous_hash).
  • 64-bit row ids: audit_trail.id and every column that stores one are declared at big size, so the sequence cannot run out on a site that writes for decades.
  • Multi-secret rotation via audit_trail_secret config entities backed by drupal/key. Activate-new-first / retire-old-second ordering so a crash mid-rotation leaves two actives (benign) instead of zero (halts writes). Per-row secret_id lets a chain span any number of rotated secrets.
  • Integrated Key-backed secret storage: the module consults drupal/key providers at write / verify time. Provider choice (config / file / env / cloud-managed / HSM-backed) is the operator's, not the module's.
  • Signed verification checkpoints in audit_trail_checkpoint for incremental walks.
  • Cron says when a chain stops verifying: critical on the audit_trail channel naming the chain and its first broken range, and notice when it verifies again. Once per change of answer, so a break that stands for a week is one line rather than one per tick, and whatever the site ships logs with carries it.
  • Auto-archive lifecycle: cron-driven staged transitions (coverage, transient-purge, archive, live-purge, file-purge and compaction) with per-chain overrides and WORM-archive bridging in the verifier.
  • Operator acknowledgments in audit_trail_acknowledgment for known-unverifiable row ranges (e.g. secrets lost in an incident).
  • Cron auto-verify with checkpoint minting and a runtime-requirements surface on /admin/reports/status.
  • Silent-row-drop counter on chain-write lock contention: surfaces as a status-report warning so sustained contention can't be exploited to hide rows.
  • Permission split: view audit trail reports (read-only) vs run audit trail verification (gates the CPU-expensive verify routes) vs administer audit trail (secrets / chains / acks).
  • Admin pages: entries listing with side-by-side diff, chain collection, secret list, archive list with per-archive SHA-256 + lifecycle HMAC verdicts, ack list, settings form.
  • Bundled submodules: audit_trail_entity (generic entity events bridge), audit_trail_entity_paragraphs (paragraph ancestry contributor), audit_trail_file (file lifecycle events), audit_trail_user_auth (authentication events), audit_trail_tsa (RFC-3161 TSA timestamping). WebDAV bridging is a separate top-level contrib module (audit_trail_webdav) released independently.
  • Drush commands: audit_trail:verify, audit_trail:acknowledge-reset, audit_trail:reindex-acknowledgments, audit_trail:reindex-segments, audit_trail:auto-archive, audit_trail:archive, audit_trail:archive-import, audit_trail:archive-restore, audit_trail:archive-verify, audit_trail:rewrite-archive, audit_trail:purge, audit_trail:compact, audit_trail:set-signing-secret, audit_trail:retire-secret, and from the TSA submodule audit_trail:timestamp and audit_trail:verify-timestamp.
  • Published docs site: mkdocs build deployed from the default branch to project.pages.drupalcode.org/audit_trail, so operators get a navigable site instead of raw markdown.
  • Comprehensive test coverage: over a thousand tests across unit / kernel / functional, counted on the metrics page, covering happy-path writes, every tampering shape (row edit / delete / insert / hmac forge / chain-link break), multi-tamper detection, secret-rotation atomicity, contributor isolation, byte-stability of CanonicalJson::encode(), TSA auth-options threading, lock-contention drop counter, chain-archive lifecycle, ack workflow.

Near-term (1.x point releases)

  • audit_trail:purge --resume: pick up after a partial purge run that crashed mid-way, instead of forcing operators to manually inspect the audit_trail_segment table to figure out the right next id range.
  • JSON-extract functional indexes: for sites that filter the entries list by uid or ip on a multi-million-row chain, add generated columns + indexes in a settings-toggle-driven update hook. Default off (most installs don't filter at that scale). Those two are equality matches, which is what a B-tree on a generated column answers. The message and request-URI filters are substring matches with a leading wildcard, so they need a different tool entirely, or a narrower match to be worth indexing.
  • TSA end-to-end test fixture: containerized TSA in audit_trail_tsa/tests/fixtures/ so the cron-driven TSA anchor flow gets covered without mocking the openssl pipeline.
  • Multi-chain archive interleaving test, coverage gap: no kernel test currently exercises archive ops on two chains running concurrently. Pin the chain-write lock's per-chain isolation under interleaved access.

Medium-term (1.x / 2.0 candidates)

  • drupal/audit_log StorageBackendInterface integration so users of that module can opt into HMAC chaining without leaving their workflow. Bridges the two modules; collaboration rather than competition.
  • A second digest algorithm: the version column and the dispatch around it already ship, so what is left is the successor itself. It arrives as version 2, a line in ChainPayload::VERSION_ALGORITHMS for rows and in SegmentVersion's for segments, stamped by the writer from then on. A chain carries records under both and nothing already signed is rewritten. Held until there is a reason to pick one: the value of the column was never the second algorithm, it was being able to add one without invalidating what is already signed.
  • Multi-instance distributed locking: Drupal's default database lock backend works fine for single-node deployments. Sites scaling horizontally with no central database lock need a Redis-backed lock. Document the configuration, ship a services.yml example.
  • Plugin-instance cache in AuditTrail::record(): the contributor-pipeline orchestrator currently re-instantiates every enabled #[ContextContributor] plugin on each record() call. Cache instances per chain so a high-volume caller doesn't pay the createInstance() cost per row. Drop in when telemetry shows the per-event microsecond budget matters.
  • Field-level redaction (GDPR-grade data minimization): wholesale archive + purge handles row-level retention, but doesn't help when ops needs to drop a SINGLE field on a SINGLE row (e.g. erase a PII column under a right-to- erasure request) while keeping the rest of the row verifiable. The plan: a new audit_trail_redaction table recording (chain, row_id, field, original_value_hash, redacted_at, uid, secret_id, hmac). The redacted field is NULLed in the row's context and the verifier rebuilds the canonical payload by reading the redaction record's original_value_hash in place of the missing field, so the row's stored hash still derives. Operator-driven via drush / UI; redactions are themselves chain-anchored evidence. Cleaner long-term GDPR posture than archive + purge for surgical erasure requests.
  • Storage backend abstraction: pull the DB write behind a StorageBackendInterface. Plug-ins to consider:
  • Append-only file backend (logrotate-friendly, harder to tamper at OS level than a DB table).
  • KMS-sealed backend that signs each row with an HSM-held key (escrowed material).
  • WORM cloud bucket backend for direct write-once persistence without local archives.
  • Conditional "forget" stage for old archive bookkeeping rows (see brainstorming below).

Brainstorming (not committed to a timeline)

Conditional "forget" stage for old archive bookkeeping

The shipped lifecycle ends at compaction (compact_after), which folds a contiguous run of file-purged segments into a single bridging row. That already removes most of the bookkeeping tail: what it deliberately keeps is one row spanning the run, so the verifier still has anchors to bridge the empty range with.

What remains open is going further and deleting the bridging row itself. That is technically safe in one configuration: when the live segment is fully detached from genesis (every row before the live segment has been live-purged), only the archive immediately preceding the live segment is consulted by the verifier. Older archive bookkeeping rows are ornamental: their anchor_after hashes don't line up with any live row's previous_hash.

What you give up by forgetting: public-hash linkage back to genesis. With only the bridge archive kept, you can HMAC-verify "the live chain plus its immediate-pre-live anchor are operator-authentic", but you can't independently reconstruct the SHA-256 hash chain from genesis anymore. For operators who don't care about multi-decade chain walkability, it's a clean trade.

Why it isn't on the roadmap yet: storage savings are modest (a 10-year monthly-archived chain = ~120 bookkeeping rows, total ~10 KB); the bookkeeping is forensic evidence of how the chain got from genesis to here; the use case is rare. If a concrete operator need surfaces, the rule above is a clean starting point.

Hash-salt secondary HMAC (defense in depth)

Today each row's hmac column signs (row.hash, secret). A secondary HMAC over (row.hash, salt) where salt is a per-secret random byte string would let an operator detect secret-extraction attacks: if an attacker steals the secret bytes and starts forging rows, their forgeries wouldn't carry the secondary HMAC (the salt isn't in the key material; it lives in a separate Key entity). Detection-not-prevention; useful for incident-response posture. Not on the roadmap proper because the threat model it addresses (attacker holds the secret but not the salt) is narrow and operationally fragile (rotating the salt requires re-signing every row, same shape as a full key rotation). Captured here so the conversation doesn't get lost.

Submitting to drupal.org

A separate axis from version progression. The project is published and releasing; what is left is what a published project still has to settle:

  • Drupal.org Security Advisory opt-in: coordinate with the drupal.org security team for SA-CONTRIB coverage. The module's tamper-evidence promise makes covered-status particularly valuable. Requires maintainer approval workflow, and coverage applies to stable releases only: an alpha, a beta and a release candidate each say on their own page that they are not covered.
  • Update hooks from the first 1.0.0 beta: from there on a schema change carries stored data forward, reversible where it can be and documented as one-way where it cannot. The alpha line ships none: reinstalling is how you move between alphas.
  • README badges / project metadata for the drupal.org project page: pipeline status, latest tagged release, PHPUnit coverage, Drupal core compatibility matrix.

Out-of-scope (other tools do it better)

  • Real-time stream into a SIEM: for sites already running Splunk / ELK / Sentry / Datadog Logs, the existing core/syslog + a syslog to SIEM pipe handles this. The chain is a complement (integrity proof on a small subset of events), not a replacement.
  • Entity revision tracking: that's audit_log's expertise. The two modules can coexist.
  • Application-layer access control: chain integrity doesn't say anything about whether the actor was authorized to perform the recorded action. That stays the responsibility of the consumer module.
  • Asymmetric / third-party-verifiable signatures: the module's HMAC layer is symmetric (operator proves authenticity to themselves). For third-party-verifiable provenance, layer TSA timestamps (already shipped) and WORM archival on top. A future module could add an asymmetric signature column for full PKI-grade provenance, but that's a different problem domain.