about
Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search (arxiv.org)
3 points by limoce 195 days ago | hide | past | pdf | discuss on HN

In plain words: Writing harmful requests in old-style Chinese can slip past a chatbot's safety filters because the text is short and hard to read. A search tool automatically rewrites and tweaks such prompts, and it broke through more often than the best known jailbreak tricks.

Abstract · Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search

As Large Language Models (LLMs) are increasingly used, their security risks have drawn increasing attention. Existing research reveals that LLMs are highly susceptible to jailbreak attacks, with effectiveness varying across language contexts. This paper investigates the role of classical Chinese in jailbreak attacks. Owing to its conciseness and obscurity, classical Chinese can partially bypass existing safety constraints, exposing notable vulnerabilities in LLMs. Based on this observation, this paper proposes a framework, CC-BOS, for the automatic generation of classical Chinese adversarial prompts based on multi-dimensional fruit fly optimization, facilitating efficient and automated jailbreak attacks in black-box settings. Prompts are encoded into eight policy dimensions-covering role, behavior, mechanism, metaphor, expression, knowledge, trigger pattern and context; and iteratively refined via smell search, visual search, and cauchy mutation. This design enables efficient exploration of the search space, thereby enhancing the effectiveness of black-box jailbreak attacks. To enhance readability and evaluation accuracy, we further design a classical Chinese to English translation module. Extensive experiments demonstrate that effectiveness of the proposed CC-BOS, consistently outperforming state-of-the-art jailbreak attack methods.

Xun Huang, Simeng Qin, Xiaoshuang Jia, Ranjie Duan, Huanqian Yan, Zhitao Zeng, Fei Yang, Yang Liu, Xiaojun Jia
arXiv:2602.22983 · cs.AI, cs.CR · submitted Feb 26, 2026 · updated Mar 24, 2026
abstract · pdf · html · ICLR 2026 Poster The source code relevant to this article has now been open-sourced; for details, please visit: https://github.com/xunhuang123/CC-BOS

add comment on HN