Don't Command, Cultivate: An Exploratory Study of System-2 Alignment
Yuhang Wang, Yuxiang Zhang, Yanxu Zhu, Xinyan Wen, Jitao Sang

TL;DR
This study explores how System-2 thinking patterns influence AI model safety, showing that deliberate reasoning can improve safety but still faces vulnerabilities, and proposes methods to enhance alignment.
Contribution
It investigates the impact of System-2 reasoning on model safety and introduces prompt engineering and supervision techniques to improve safety alignment.
Findings
o1 models show improved safety but remain vulnerable to jailbreak attacks
Prompt engineering helps models scrutinize user requests more carefully
Proposed process supervision could further enhance safety alignment
Abstract
The o1 system card identifies the o1 models as the most robust within OpenAI, with their defining characteristic being the progression from rapid, intuitive thinking to slower, more deliberate reasoning. This observation motivated us to investigate the influence of System-2 thinking patterns on model safety. In our preliminary research, we conducted safety evaluations of the o1 model, including complex jailbreak attack scenarios using adversarial natural language prompts and mathematical encoding prompts. Our findings indicate that the o1 model demonstrates relatively improved safety performance; however, it still exhibits vulnerabilities, particularly against jailbreak attacks employing mathematical encoding. Through detailed case analysis, we identified specific patterns in the o1 model's responses. We also explored the alignment of System-2 safety in open-source models using prompt…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsSystems Engineering Methodologies and Applications · Complex Systems and Decision Making
