Google's Android Bench 2.0 replaces binary pass/fail grades with continuous scoring to test AI models on complex multi-day ...
AWS’ Deception Benchmark tests whether AI models can tell real security vulnerabilities from safe code that looks risky.