TypeSafe AI's judgment-only model Jev lost to Claude Haiku 4.5 on a single phishing-detection question in a third-party benchmark released two days after launch. Splitting the task into 5 questions ...