Эта статья является препринтом и не была отрецензирована.
О результатах, изложенных в препринтах, не следует сообщать в СМИ как о проверенной информации.
Selective Permeability: A Behavioral-Security Metric for LLM Advisors, with Two Failure Modes of In-Context Provenance Workflows
1. Gordeychik, S. (2026). Machine-Speed Cyber and Poisoned Cognition: A Layer-Dependent Game-Theoretic Framework, with Empirical Probes. PREPRINTS.RU. https://doi.org/10.24108/preprints-3115766
2. Hagendorff, T., Dasgupta, I., Binz, M., Chan, S. C. Y., Lampinen, A., Wang, J. X., Akata, Z., & Schulz, E. (2023). Machine Psychology. arXiv:2303.13988. https://arxiv.org/abs/2303.13988
3. Binz, M., & Schulz, E. (2023). Using cognitive psychology to understand GPT-3. Proceedings of the National Academy of Sciences, 120(6), e2218523120. https://doi.org/10.1073/pnas.2218523120
4. Stella, M., Hills, T. T., & Kenett, Y. N. (2023). Using cognitive psychology to understand GPT-like models needs to extend beyond human biases. Proceedings of the National Academy of Sciences, 120(43), e2312911120. https://doi.org/10.1073/pnas.2312911120
5. Ko, C., Shin, J., Song, H., Lee, H., Hwang, E. J., & Park, J. C. (2026). Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM Collectives. ACL 2026; arXiv:2604.06091. https://arxiv.org/abs/2604.06091
6. Zhang, L., & Chen, W. (2026). Human-like Social Compliance in Large Language Models: Unifying Sycophancy and Conformity through Signal Competition Dynamics. arXiv:2601.11563. https://arxiv.org/abs/2601.11563
7. Qu, J., Fu, L., & Hu, Y. (2026). Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity. arXiv:2606.01637. https://arxiv.org/abs/2606.01637
8. Wachowiak, L., Blain, S. D., Williams-King, D., & Marro, S. (2026). LLM Wardens: Mitigating Adversarial Persuasion with Third-Party Conversational Oversight. arXiv:2605.08321. https://arxiv.org/abs/2605.08321
9. Kumarappan, A., & Mujoo, A. (2026). Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy. arXiv:2605.12991. https://arxiv.org/abs/2605.12991
10. Louck, Y. (2026). Securing LLM-Agent Long-Term Memory Against Poisoning: Non-Malleable, Origin-Bound Authority with Machine-Checked Guarantees. arXiv:2606.24322. https://arxiv.org/abs/2606.24322
11. Asch, S. E. (1956). Studies of independence and conformity: I. A minority of one against a unanimous majority. Psychological Monographs: General and Applied, 70(9), 1-70. https://doi.org/10.1037/h0093718
12. Milgram, S. (1963). Behavioral study of obedience. The Journal of Abnormal and Social Psychology, 67(4), 371-378. https://doi.org/10.1037/h0040525
13. Freedman, J. L., & Fraser, S. C. (1966). Compliance without pressure: The foot-in-the-door technique. Journal of Personality and Social Psychology, 4(2), 195-202. https://doi.org/10.1037/h0023552
14. Johnson, M. K., Hashtroudi, S., & Lindsay, D. S. (1993). Source monitoring. Psychological Bulletin, 114(1), 3-28. https://doi.org/10.1037/0033-2909.114.1.3
15. Green, D. M., & Swets, J. A. (1966). Signal Detection Theory and Psychophysics. Wiley.
16. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. arXiv:2302.12173. https://arxiv.org/abs/2302.12173
17. Debenedetti, E., Zhang, J., Balunovic, M., Beurer-Kellner, L., Fischer, M., & Tramer, F. (2024). AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents. Advances in Neural Information Processing Systems, 37. https://doi.org/10.52202/079017-2636
18. Debenedetti, E., Shumailov, I., Fan, T., Hayes, J., Carlini, N., Fabian, D., Kern, C., Shi, C., Terzis, A., & Tramer, F. (2025). Defeating Prompt Injections by Design. arXiv:2503.18813. https://arxiv.org/abs/2503.18813
19. Chen, Z., Xiang, Z., Xiao, C., Song, D., & Li, B. (2024). AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases. Advances in Neural Information Processing Systems, 37. https://proceedings.neurips.cc/paper_files/paper/2024/hash/eb113910e9c3f6242541c1652e30dfd6-Abstract- Conference.html
20. OWASP Gen AI Security Project. (2025). LLM07:2025 System Prompt Leakage. https://genai.owasp.org/llmrisk/llm072025-system-prompt-leakage/
21. Jones, M., & Hardt, D. (2012). The OAuth 2.0 Authorization Framework: Bearer Token Usage. RFC 6750. https://www.rfc-editor.org/rfc/rfc6750
22. Youden, W. J. (1950). Index for rating diagnostic tests. Cancer, 3(1), 32-35. https://doi.org/10.1002/1097- 0142(1950)3:1%3C32::AID-CNCR2820030106%3E3.0.CO;2-3
23. Wataoka, K., Takahashi, T., & Ri, R. (2024). Self-Preference Bias in LLM-as-a-Judge. arXiv:2410.21819. https://arxiv.org/abs/2410.21819
24. Robinson, S., Oktar, K., Collins, K. M., Sucholutsky, I., & Allen, K. R. (2026). Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models. arXiv:2602.21262. https://arxiv.org/abs/2602.21262
25. Kim, J., & Flanigan, J. (2026). Right or Wrong, Models Comply: Directional Blindness in LLM Moral Judgment. arXiv:2606.14037. https://arxiv.org/abs/2606.14037
26. Waqas, D., Golthi, A., Hayashida, E., & Mao, H. (2025). Assertion-Conditioned Compliance: A Provenance-Aware Vulnerability in Multi-Turn Tool-Calling Agents. arXiv:2512.00332. https://arxiv.org/abs/2512.00332