AAAI 2026 大模型安全相关论文整理

AAAI 2026 大模型安全相关论文整理

总目录 大模型安全研究论文整理 2026年版:https://blog.csdn.net/WhiffeYF/article/details/159047894

https://claude.ai/chat/916dfe36-9753-4199-baa2-44fc2f709fb6

统计:共收集 27 篇论文,来自 AAAI 2026(第40届,2026年1月,新加坡)

主要来源:AI Alignment 特别 Track (Vol 40 No.44)和主技术 Track(NLP / ML 等)

分类概览:

  • 越狱攻击方法(Jailbreak Attack):10 篇
  • 安全防御与对齐(Defense & Alignment):10 篇
  • 安全评估与基准(Benchmark & Evaluation):4 篇
  • 隐私与数据安全:2 篇
  • 智能体安全(Agent Safety):1 篇

1 越狱攻击方法(Jailbreak Attack)

id 论文名 Track 链接
1 MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs AI Alignment PDF
2 Differentiated Directional Intervention: A Framework for Evading LLM Safety Alignment AI Alignment PDF
3 HumorReject: Decoupling LLM Safety from Refusal Prefix via a Little Humor AI Alignment PDF
4 StyleBreak: Revealing Alignment Vulnerabilities in Large Audio-Language Models via Style-Aware Audio Jailbreak AI Alignment PDF
5 STACK: Adversarial Attacks on LLM Safeguard Pipelines AI Alignment PDF
6 Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment AI Alignment PDF
7 Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models AI Alignment PDF
8 Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models Main Track (NLP) PDF
9 Multi-Turn Jailbreaking Large Language Models via Attention Shifting Main Track (NLP) PDF
10 Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs (CognitiveAttack) Main Track (NLP) PDF

2 安全防御与对齐(Defense & Alignment)

id 论文名 Track 链接
1 AlignTree: Efficient Defense Against LLM Jailbreak Attacks AI Alignment PDF
2 EASE: Practical and Efficient Safety Alignment for Small Language Models AI Alignment PDF
3 Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training AI Alignment PDF
4 STAR-1: Safer Alignment of Reasoning LLMs with 1K Data AI Alignment PDF
5 CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing AI Alignment PDF
6 WALKSAFE: Risk-aware Graph Random Walk with Bi-GRPO for LLM Safety Main Track (NLP) PDF
7 DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt AI Alignment PDF
8 Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks AI Alignment PDF
9 MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror Crafting Main Track (NLP) dblp
10 AntiDote: Bi-level Adversarial Training for Tamper-Resistant LLMs Main Track (NLP) dblp

3 安全评估与基准(Benchmark & Evaluation)

id 论文名 Track 链接
1 Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models AI Alignment PDF
2 MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks AI Alignment PDF
3 MMJ-Bench: A Comprehensive Study on Jailbreak Attacks and Defenses for Vision Language Models Main Track PDF
4 Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding AI Alignment PDF

4 隐私与数据安全

id 论文名 Track 链接
1 CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense AI Alignment PDF
2 Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models AI Alignment PDF

5 智能体安全(Agent Safety)

id 论文名 Track 链接
1 Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems AI Alignment PDF

6 其他相关论文(对齐理论 / 推理安全 / 幻觉检测 / 可解释性等)

以下论文虽然不直接属于"攻击/防御",但与大模型安全密切相关:

id 论文名 方向 链接
1 DNR Bench: Benchmarking Over-Reasoning in Reasoning LLMs 推理冗余/Overthinking PDF
2 Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning 推理安全/CoT可靠性 PDF
3 Bolster Hallucination Detection via Prompt-Guided Data Augmentation 幻觉检测 PDF
4 Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Models 幻觉/可靠性 PDF
5 Silenced Biases: The Dark Side LLMs Learned to Refuse 对齐副作用/过度拒绝 PDF
6 Beyond I'm Sorry, I Can't: Dissecting Large-Language-Model Refusal 拒绝机制分析 PDF
7 Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation 微调导致的对齐失效 PDF
8 AdvBDGen: A Robust Framework for Generating Adaptive and Stealthy Backdoors in LLM Alignment 后门攻击 PDF
9 Editing as Unlearning: Are Knowledge Editing Methods Strong Baselines for Large Language Model Unlearning? 机器遗忘 PDF
10 Polarity-Aware Probing for Quantifying Latent Alignment in Language Models 对齐可解释性 PDF
11 FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight 推理缺陷检测 PDF
12 Backdoor Attacks on Open Vocabulary Object Detectors via Multi-Modal Prompt Tuning 多模态后门攻击 PDF
13 Security Attacks on LLM-based Code Completion Tools 代码工具安全 PDF
14 MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control 智能体安全评估 PDF
15 MAJIC: Markovian Adaptive Jailbreaking via Iterative Composition of Diverse Innovative Strategies 越狱攻击 dblp
16 From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification Framework 去毒化防御 dblp
17 An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains 供应链后门 AAAI 2026

备注

相关推荐
momodira81 天前
暑期安全短视频大赛线上投票制作教程
安全
humors2211 天前
新手笔记本/路由/手机安全问题小记
安全·华为·电脑·手机·路由器·笔记本·tplink
廋到被风吹走1 天前
【AI】从“卖能力“到“卖信任“,合规与安全成为新战场
人工智能·安全
FrameNotWork1 天前
HarmonyOS 6.0 文件加密与安全存储:从哈希到硬件级密钥管理全链路实战
安全·哈希算法·harmonyos
栩栩云生1 天前
命令行的门槛从"会写"变成了"会拦"
安全·ai编程·命令行
GitLqr1 天前
别再盲目复制了:彻底搞懂 CORS 的本质与那些“神坑”
安全·http·面试
xian_wwq1 天前
【案例分析】Hugging Face生产基础设施入侵攻击分析
网络·安全
2401_873479401 天前
如何识别C2通信中的恶意出站IP?IP离线库+威胁情报融合方案
网络·tcp/ip·安全·ip
OpenAnolis小助手1 天前
Anolis OS 23.5 发布:全新平台支持、DDE 桌面升级,安全与多架构能力再度跃升
安全·操作系统·龙蜥社区·anolis os·anolis os 23.5
北冥you鱼1 天前
OpenZeppelin Contracts 完全指南:从入门到精通,构建安全的智能合约
安全·区块链·智能合约