Entity
Jailbreaking
Techniques used to override an AI system's ethical guidelines and content restrictions
术语属性
- category:AI Security Attack
- target:Safety Filters
- mitigation:Robust Alignment
Entity
Techniques used to override an AI system's ethical guidelines and content restrictions