Skip to content

Known gaps

Tracks known limitations that are not yet fixed. These are not accepted architectural trade-offs. They're TODO items.

Retrieval has no langcode filter

Where: ai_search module (contrib), RagTool::execute() (web/modules/contrib/ai_search/src/Plugin/AiFunctionCall/RagTool.php).

Problem: the RAG query is built from search_string, amount, and min_score only: nothing constrains results to the visitor's current interface language. On any multilingual site where translated nodes share the same underlying content, a query whose top-scoring match happens to be a translation in a different language will have that translation loaded (AnswerService::loadTranslated()) and fed to the LLM as a source, regardless of what language the visitor asked in.

Current mitigation (config only, per agent):

  • Raise min_score (Tools → RAG/Vector Search → Property setup → min_score, on the agent's own edit form) to cut down on weak keyword-only matches; see Getting started for how to pick a value.
  • Add an explicit language/relevance instruction to "Extra generation guidance" (AiAnswersAgentForm) telling the model to answer in the question's language and to decline weakly-related sources.

This reduces how often a wrong-language or off-topic chunk gets cited, but doesn't prevent one from being retrieved in the first place. A sufficiently ambiguous query can still score a wrong-language chunk above the threshold and get handed to the model. The config change only asks the model to compensate after the fact.

A reference is rendered in the visitor's current language whenever the entity has that translation, so the reference cards match the page language even when the chunk that matched did not. The chunk text given to the model is still whatever language matched.

Real fix (not implemented): add a langcode constraint to the Search API query in RagTool::execute(), e.g. restrict to (or boost) the current interface language, so mismatched-language chunks aren't retrieved as candidate sources at all. Tracked in ai_search as 3562047 ("Add language context to RAG tool"), which depends on 3554899 ("Language metadata is not added to indexed items"): the vector index has no language field to filter on yet. Both target ai_search 2.0.x, so no fix is expected in 1.3.x.

Status: open, not tested against an adversarial query. The one manual check performed happened to score the correct-language source highest, before this gap was identified.

In-module rerank shows a different order than the LLM saw

Where: AnswerService::maybeRerank() (src/Service/AnswerService.php).

Problem: the in-module rerank pool used to be a 4x over-fetch against the tool's result count, giving the reranker room to actually reorder. It now equals the tool's amount exactly, so maybeRerank() only reorders the final set, and it does so after the LLM has already generated its answer from the raw, unreranked tool order. That means two different orderings of the same chunks exist within a single context window: the one the LLM read, and the one presented in the reference list.

Current mitigation: none; the inconsistency between what the model read and what's shown is accepted as-is for now.

Real fix (not implemented): adopt the ai_reranker Search API index processor once drupal/ai 1.5.0 is stable. That processor reranks the query at the point the tool itself issues it, producing one consistent ordering everywhere instead of two. Tracked in ai_answers issue 3610866, which also adds a "maximum references" setting so amount can go back to being a wide over-fetch once rerank happens upstream of the tool. The ai_reranker processor itself needs an upstream fix too: its extractItemText() should prefer extraData('content') over its current source.

Status: open, blocked on a stable drupal/ai 1.5.0 release (1.5.0-rc4 is out) and an ai_reranker upstream fix.

Vendored JS libraries have no update mechanism

Where: js/vendor/markdown-it.min.js (markdown-it 15.0.1) and js/vendor/purify.min.js (DOMPurify 3.4.15), used by ai_answers.markdown.js. Both are pinned, minified builds committed straight to git — no composer.json/package.json declares them as dependencies.

Problem: nothing watches these files for new releases or security advisories. No Dependabot/Renovate config, no CI dependency-scan step, and no Composer/Libraries API indirection — updating means someone manually noticing a new version exists, re-downloading, and re-verifying. This is a real risk for DOMPurify specifically: it's the XSS-defense layer sanitizing model output, and has a genuine CVE/security-patch history.

The rest of the ai module family has never solved this either: ai_chatbot vendors Showdown.js the same way (js/showdown.min.js) and it has sat frozen at v2.1.0 since April 2022 with no tracking issue. Its deepchat.bundle.js was bumped exactly once, manually, tied to a one-off drupal.org issue (#3511968) rather than any process. There's no precedent to inherit.

Current mitigation: the update steps are documented here, not run from a committed script (an executable file in the module gets more scrutiny than its actual behavior warrants, for a one-off maintenance task — see #3615751). To update either library:

  1. Check the current pinned versions in the markdown_vendor library comment in ai_answers.libraries.yml, and check https://github.com/markdown-it/markdown-it/releases / https://github.com/cure53/DOMPurify/releases for newer ones.
  2. Re-download at the new version(s):
    curl -sf "https://cdn.jsdelivr.net/npm/markdown-it@<version>/dist/browser/markdown-it.umd.min.js" \
      -o js/vendor/markdown-it.min.js
    curl -sf "https://cdn.jsdelivr.net/npm/dompurify@<version>/dist/purify.min.js" \
      -o js/vendor/purify.min.js
    
  3. Update the version comment above the markdown_vendor library in ai_answers.libraries.yml to match, and refresh js/vendor/LICENSE-* if either project's license text changed.
  4. Before committing, re-check rendering as described in Getting started › Test: headings, lists, tables, code, and citation links render as HTML, with no console errors. A new version can change rendering or sanitization behavior.

This still has to be done by a human who remembers to check for new releases — nothing triggers it automatically.

Real fix (not implemented): either wire up genuine automation (a scheduled CI job that checks npm for new markdown-it/DOMPurify releases and opens an MR), or move to Drupal's Libraries API pattern (site builders install the libraries themselves via Composer npm-asset packages, with a version constraint the module declares) so ordinary dependency-update tooling applies. Both are real scope beyond what this module currently needs; revisit if either library ships a security advisory.

Status: open, accepted as a manual, documented process for now.

Tool messages are rewritten to work around a stale sources block

Where: AnswerToolResultSubscriber::onToolFinished() (src/EventSubscriber/AnswerToolResultSubscriber.php).

Problem: ai_agents builds a round's system prompt before it runs that round's tool calls, so when the agent calls rag_search or a source tool itself, the numbered sources block in the system prompt does not yet contain that call's sources. The model would cite numbers from a block that lags one call behind.

Current mitigation: the subscriber replaces the tool message the model reads next with the sources the call added, numbered with their final citation numbers.

Real fix (not implemented): ai_agents building the system prompt after running a round's tools, or exposing a hook to refresh it, so the sources block is current and the tool message can stay the tool's own.

Status: open, workaround in place.

Listing output is parsed by convention

Where: AnswerService::sourcesFromListedOutput() (src/Service/AnswerService.php).

Problem: a source tool's structured output is read in one of two shapes: rag_search's rows, or the outputs.results[]._metadata shape that tool_ai_connector produces for Tool Belt's entity tools. Neither is a declared contract, so a tool with any other output shape yields no sources.

Current mitigation: none needed for rag_search and tool:tool_belt:entity_list, which are the supported pairings.

Real fix (not implemented): a documented result contract, or an event that lets a module map its tool's output to sources.

Status: open.