about
CCFC: Core and Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection (arxiv.org)
1 point by summarity on Aug 21, 2025 | hide | past | pdf | discuss on HN

In plain words: It answers each request twice: once using only its stripped-down core to ignore added junk, and once wrapping that core around the full text to break attack patterns. A safety check across both cut jailbreaks by 50–75% versus top prompt defenses, without hurting normal answers.

Abstract · CCFC: Core & Core-Full-Core Dual-Track Defense for LLM Jailbreak Protection

Jailbreak attacks pose a serious challenge to the safe deployment of large language models (LLMs). We introduce CCFC (Core & Core-Full-Core), a dual-track, prompt-level defense framework designed to mitigate LLMs' vulnerabilities from prompt injection and structure-aware jailbreak attacks. CCFC operates by first isolating the semantic core of a user query via few-shot prompting, and then evaluating the query using two complementary tracks: a core-only track to ignore adversarial distractions (e.g., toxic suffixes or prefix injections), and a core-full-core (CFC) track to disrupt the structural patterns exploited by gradient-based or edit-based attacks. The final response is selected based on a safety consistency check across both tracks, ensuring robustness without compromising on response quality. We demonstrate that CCFC cuts attack success rates by 50-75% versus state-of-the-art defenses against strong adversaries (e.g., DeepInception, GCG), without sacrificing fidelity on benign queries. Our method consistently outperforms state-of-the-art prompt-level defenses, offering a practical and effective solution for safer LLM deployment.

Jiaming Hu, Haoyu Wang, Debarghya Mukherjee, Ioannis Ch. Paschalidis
arXiv:2508.14128 · cs.CR, cs.AI · submitted Aug 19, 2025
abstract · pdf · html · 11 pages, 1 figure

add comment on HN