Safety evaluations of large language models (LLMs) typically report binary outcomes such as attack success rate, refusal rate, or harmful/not-harmful response classification. While useful, these can ...
Abstract: The rapid evolution of Large Language Models (LLMs) has increased the need for better ways to support human AI interaction. Prompt engineering is the most common method for guiding model ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results