← Back to feed
5

Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity

Researchers have introduced a new benchmark designed to evaluate how language models handle ambiguous instructions, policy conflicts, and adversarial commands.

Impact
25/100
Current rank score
5.38
Source tier
Tier 1
Category
Research
Read the full story at arxiv.org

Firefly links to the original publisher. The summary above is AI-generated for orientation and may differ from the source. The “current rank score” decays over time so newer significant stories surface first.