AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SSP-Bench: A Hybrid Data Generation Framework for Safety, Security, and Privacy Evaluation

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Evaluation of large language models (LLMs) for safety, security, and privacy (SSP) relies heavily on static benchmarks, which suffer from score saturation, data contamination, and aggregation artifacts, and fail to capture sensitivity to linguistic variation. As a result, models that perform well on fixed test sets often fail under semantically equivalent rephrasings. We introduce SSP-Bench, a dynamic benchmarking framework that generates evaluation instances on demand while preserving domain consistency. The framework ensures label validity through externally grounded sources, enforces scope

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T06:11:12.848Z. This is not the publication date.