ПРЕПРИНТ

Эта статья является препринтом и не была отрецензирована.
О результатах, изложенных в препринтах, не следует сообщать в СМИ как о проверенной информации.
Selective Permeability: A Behavioral-Security Metric for LLM Advisors, with Two Failure Modes of In-Context Provenance Workflows
2026-07-25

LLM advisors must incorporate legitimate evidence without becoming indiscriminately compliant. We define selective permeability as an operating-point summary of that balance: authorized-update rate minus illegitimate- leak rate. The quantity is equivalent to Youden's J and therefore gives equal weight to false acceptance and false rejection; we report both components because deployments may assign very different costs to them. We test seven frontier models, in three synthetic decision domains across single-turn, multi-turn, and persistent-memory settings. The results support three principal findings. First, blinded content, relevance, provenance, and independence auditors detect different failure classes, but their false positives can make them poor quarantine gates when the underlying model already has a low leak rate. Second, bearer-style provenance controls fail in two measurable ways: the model may emit the authority token, and ordinary compaction may erase the metadata needed to verify its origin. Per-item token-preserving summarization restores the wiki channel but does not establish a provenance-control advantage over the no-defense condition. The policy effect was larger in all seven selected model rows. This result concerns explicit policy directives placed in user context; it neither identifies an internal reasoning mechanism nor tests system/developer-policy placement. The study characterizes behavior in these assays and does not rank models.

Ссылка для цитирования:

Gordeychik S. 2026. Selective Permeability: A Behavioral-Security Metric for LLM Advisors, with Two Failure Modes of In-Context Provenance Workflows. PREPRINTS.RU. https://doi.org/10.24108/preprints-3115999

Список литературы