Architecture¶
Request flow¶
POST /ai-answers/question requires the use ai answers permission and a
valid X-CSRF-Token request header, both enforced at the routing layer
(ai_answers.routing.yml), not in the controller. AiAnswersController::question()
decodes the JSON body, then content-negotiates on the Accept header: a
request for text/event-stream streams Server-Sent Events via
streamAnswer(); anything else returns a single JSON Answer via
jsonAnswer(). Both paths call the same AnswerService::answer(), differing
only in whether streaming callbacks are passed. Before the SSE stream starts,
the controller explicitly releases the session write lock
($session->save()) so it doesn't block other concurrent requests in the
same session.
Exceptions from answer() map to HTTP status in the controller:
DomainException → 409, InvalidArgumentException → 400,
OutOfBoundsException → 404, anything else → 500 (logged, generic message
returned).
Answer pipeline¶
AnswerService::answer():
- Validates the question (rejects empty/whitespace-only input).
- Loads the conversation record from
keyvalue.expirable, collectionai_answers.conversations. Follow-up turns are pinned to the record's agent. - Gates:
AgentSettingsRegistry::getSettings()must show the agent enabled for answers;RagSettingsResolver::resolve()must return a non-null result.RagSettingsResolveritself never throws. It returns?array,NULLwhen the agent has noai_search:rag_searchtool configured or no forced index.AnswerService::answer()is what turns aNULLresolve into a thrownDomainException. prepareConversation()authorizes the caller against the conversation's owner uid. Anonymous-owned conversations are bearer-by-id (anyone holding the conversation id can resume them); authenticated-owned conversations only resume for the same uid. Unknown or mismatched-owner conversations throwOutOfBoundsException.generate()runs the agent throughAiAgentEntityWrapper(plugin.manager.ai_agents):setChatInput()with the threaded history, an explicitsetRunnerId()(a UUID generated byAnswerServiceitself, since the wrapper's own tagging has no other injection point),determineSolvability(), thensolve().- On a brand-new conversation (no prior turns), the service forces
retrieval directly via
forceRagSearch()rather than leaving the choice to the agent. On a follow-up turn, retrieval is entirely the agent's own tool-call decision: if it doesn't call the tool, the previous turn's held sources are reused. - The stream is consumed inside the same
try/finallyblock that ownsAgentRunContext;finallyclears the context. This matters because a tool-calling agent's second round only executes lazily, while the caller iterates the stream. Clearing the context any earlier silently drops the sources block before round 2 runs. - After
generate()returns, the no-answer message is substituted for$textwhenever nothing has actually reached the client yet: either the non-streaming path found no usable sources, or generation produced no text at all on either path.$onTokenis only ever invoked with the same non-empty pieces$textaccumulates, so$text === ''means notokenframe was sent either, making the substitution safe even when streaming. A streamed answer that already reached the client with some text is left alone, even with zero usable sources, since it may already be visible; the citation contract handles that case instead.
Subscriber bridge¶
AgentRunContext is a single-slot service (not keyed by runner id) storing
the current run's agent id, base system prompt, an optional sources block,
and a tool-results callback.
AnswerSystemPromptSubscriberlistens toBuildSystemPromptEvent(ai_agents.pre_system_prompt, defined in theai_agentsmodule) at priority 10. It guards on the event's agent id matching the run context's agent id: if they don't match, it returns without touching the prompt, protecting against event bleed across unrelated concurrent runs. When they match, it appends the citation contract, guidance, and sources block to the agent's own system prompt; it never replaces it. Priority 10 is chosen becauseai_context's equivalent subscriber runs at priority 0 and must layer on top of this module's additions, not be overwritten by them.AnswerToolResultSubscriberlistens toAgentToolFinishedExecutionEvent(ai_agents.tool_finished_executed, also defined inai_agents) at the default priority. Like the subscriber above, it first guards on the event's agent id matching the run context's agent id, returning early on a mismatch. It then filters for aStructuredExecutableFunctionCallInterfaceinstance that is eitherai_search:rag_search(itsgetStructuredOutput()['results']) or one of the agent's configured source tools (its whole structured output), and reports it to the run context along with the tool's plugin id.- The same subscriber then replaces the tool message the model reads next
(
setOutput()) with the sources that call added, numbered with their final citation numbers.ai_agentsbuilds a round's system prompt before running that round's tools, so the sources block in the system prompt is always one call behind a tool the agent calls itself; without this, the model cites numbers that point at the wrong references. A source tool's own summary message is kept above the numbered sources.
Source processing¶
sourcesFromToolResults() applies the score gate, then maybeRerank()
reorders using the site-default rerank provider if one is configured, a
stopgap pending ai_reranker (see the
rerank known gap).
applyEntityCap() caps by distinct entity, skipping rather than breaking so
extra chunks of already-included entities aren't lost. renderReferences()
dedupes per entity, does a translation-aware load with an access re-check,
and renders each in the agent's configured view mode. It prefers the
visitor's current content language when the entity has that translation,
and otherwise keeps the language the source was retrieved in.
A configured source tool's output goes through resolveListedSources()
instead. sourcesFromListedOutput() accepts rows shaped like rag_search's
(entity_type, entity_id, langcode, content) or like Tool API entity
output wrapped by tool_ai_connector (outputs.results[], with the entity
under _metadata.type/_metadata.id). Without a content value, each
source's chunk is one line: the entity's label in the current language, with
the row's other field values beside it, so two entities sharing a label
(one program at two degree levels) stay distinguishable by number. A listing
is complete by design, so it skips the score threshold and maybeRerank();
only the tool's own max_results cap applies. On the SSE path the
references event fires before generation completes, so it can only list
every retrieved source at that point.
A tool-calling turn can call rag_search more than once if the model judges
one round's results insufficient. resolveRetrievedSources() takes an
$existingSources parameter for exactly this: each round's own candidates are
merged onto whatever the turn already gathered, ahead of renderReferences()'s
own dedupe, so an entity already assigned a citation number keeps that same
position (and therefore that same number) no matter how many further rounds
run. Without this, a later round would replace $sourcesUsed outright and a
citation the model wrote against an earlier round's numbering would resolve
against a completely different round's sources by the time the answer
finishes.
Once generation finishes, finalizeCitations() drops any source that never
got an inline [n] marker in the text and renumbers the survivors in
first-citation order, returning an empty sources array if the text cites
nothing at all. On the JSON path this is the only Sources list the client
ever sees. On the SSE path the done event separately carries this
corrected text/references pair, and ai_answers.answer.js swaps to it
once done arrives. The earlier references event's list is provisional.
Persistence & feedback¶
Each turn stores log and trace ids captured once, via a single query for the
newest ai_agents_runner_<runnerId> ai_log entry. FeedbackLogger
authorizes against the same conversation record and owner rule as
AnswerService, requires the agent's feedback_enabled setting, annotates
the stored ai_log entry (deduping any prior feedback:* tag rather than
accumulating them, and writing an ai_answers_feedback extra-data payload),
and optionally scores a Langfuse trace when a trace_id was captured.
Front end¶
All three blocks render only configuration and drupalSettings into a cacheable
static shell. Answer content always arrives afterward via the API, never
baked into cached markup. The Answer block's DOM id is
Html::getUniqueId('ai-answers-answer-' . $instanceUuid), which deduplicates
a colliding id within the same request rather than emitting it verbatim. The
Question block doesn't recompute this formula at render time; its admin form
enumerates placed Answer blocks and stores the already-computed id string as
the target setting, which its JS resolves at runtime with
document.getElementById().
Wire protocol: { agent, question, conversation_id? } in; SSE events
references / token / done / error out. Drupal.aiAnswers.answer.ask()
is the sole function that actually issues the request. There is no parallel
implementation of the fetch itself. Chip click and same-page form submit
(ai_answers.question.js's submit()) go through dispatchAsk() first,
which calls ask() directly when the target Answer block is on the same
page, or falls back to an ai-answers:ask CustomEvent that the Answer
block's own listener turns back into an ask() call. The cross-page
fragment handoff and the same-block follow-up form call ask() directly,
bypassing dispatchAsk() entirely, since both already run on the Answer
block's own page.
CSS :not([hidden]) scoping is required only on elements with a
non-none display override that would otherwise fight the [hidden]
attribute: currently the feedback and follow-up containers; the references
list has no such override and needs no guard.
token/done frames carry raw Markdown; ai_answers.markdown.js converts it
to sanitized HTML before it ever reaches innerHTML. Two independent layers
enforce that boundary, since the text is model output and therefore
untrusted: markdown-it parses
with html: false, so any raw HTML the model writes is escaped to inert text
rather than parsed as markup; the resulting HTML is then run through
DOMPurify.sanitize() with an explicit
tag allowlist (h1–h6, p, br, strong, em, code, pre, ul, ol,
li, a, table/thead/tbody/tr/th/td), an attribute allowlist
(href, data-align), and a URI allowlist restricted to http(s):,
mailto:, and relative/fragment URLs. Both libraries are vendored, pinned,
minified builds under js/vendor/ (see
Vendored JS libraries
for versions and how to update them), not pulled from a CDN or installed via
the Libraries API.
Drupal.aiAnswers.markdownToBlocks() is the other half of this file: it
groups markdown-it's own block token stream into top-level blocks (heading,
paragraph, list, table, fenced code, …) by tracking nesting depth, and
renders and sanitizes each block independently. renderStreamingBlocks() in
ai_answers.answer.js diffs against this block list on every token frame
and only replaces the block that actually changed (normally just the growing
tail block), so a settled block keeps its DOM element identity — a text
selection survives, and a CSS entrance animation runs once instead of
retriggering on every token. Drupal.aiAnswers.markdownToHtml() (used once,
on the done frame) is just markdownToBlocks().join('').
List items are always rendered tight (no wrapping <p>, regardless of blank
lines in the model's raw Markdown between bullets): markdownToBlocks()
forces hidden = true on any paragraph token nested inside a list before
rendering. Without this, CommonMark's "loose list" rule — which markdown-it
follows correctly — would wrap each item's text in <p>, and list items
would inherit the browser's default paragraph margin once per item, since
this module has no p { margin: 0 } reset of its own.
A references section (local or a paired Sources block) is shared across every
turn in a conversation, not rebuilt per question: ask() assigns each call
its own turnId (a client-side counter, independent of the backend's own
turn/conversationId, which only arrive on the done frame — too late for
the references frame's own render) and every citation anchor, reference
<li> id, and pending-placeholder element carries that turn as a
data-ref-turn attribute. renderReferences() and renderReferencesPending()
only ever add or remove elements tagged with the current turn, so an earlier
turn's already-rendered references, and a still-open earlier turn's citation
links, survive later turns starting to stream, erroring, or re-rendering
their own references twice (once on the references frame, again on done).
The one exception is a genuinely new (non-follow-up) question, where every
section is fully cleared: there is no earlier turn's state left to preserve.
Each reference <li> shows only the <ol>'s own running count, not a
per-turn [n] marker: an index that restarts every turn would collide with
the shared list's numbering. data-ref-index/data-ref-turn carry the
per-turn identity for anchoring.
A reference that a follow-up turn re-retrieves is not given a second card:
renderReferences() identifies each reference by referenceKey() (its
entity_type and entity_id, falling back to url, then label, and never
matched when none is present) and tracks the first turn/index that got a card
for a key in state.seenReferenceKeys, a Map that persists for the whole
conversation and is reset only when a new (non-follow-up) question clears
everything. A duplicate still gets its own citation index and its own
clickable [n] in the answer text, but that citation's anchor points at the
earlier occurrence's existing <li> instead of a new one being built. The
entity id comes first because neither a URL nor a label is unique per
entity: two nodes can share a path alias or a title.
Every link in a reference opens in a new tab: the label link, and any link inside the rendered entity (for example a card whose whole surface is one link), so following a source never navigates away from the answer.
The Sources block follows the same targeting pattern as the Question block.
Its admin form stores the target Answer block's DOM id as the target
setting, rendered into the data-ai-answers-sources-target attribute on
its own section (see SourcesBlock::build()). There's no shared registry
linking the two. ai_answers.answer.js's referenceSections() finds every
matching Sources-block section by DOM query
([data-ai-answers-sources-target="<answer-root-id>"]) at render time,
alongside the Answer block's own inline references list if it has one, and
fills all of them with the same reference data. A Question/Answer pair can
have zero, one, or several standalone Sources blocks placed anywhere on the
page.
Both admin forms build their "Target Answer block" options from
AnswerBlockTargetTrait::getAnswerBlockOptions(), which lists every placed
Answer block regardless of how it was placed: a classic block config
entity (Block Layout), or a Drupal Canvas component instance in a
canvas_page's content, a content_template, or a page_variant. Each
Canvas entity type is checked for with hasDefinition() and skipped
independently, so the trait works unmodified whether Canvas is absent,
partially used, or fully adopted. Canvas does not support Drupal 10, so
there it is always absent. A component tree item's inputs arrive
as a plain array on config entities but as a JSON-encoded string on the
canvas_page content entity's components field; the trait normalizes
both shapes before reading instance_uuid. Every option, classic or
Canvas, is keyed by the same ai-answers-answer-<instance_uuid> DOM id, so
downstream targeting code never needs to know which kind of placement it
resolved.