It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
… But Gleave also believes that the findings show that models can be systematically tested for safety. “There's an optimistic angle here,” he says. “Defense and safety really are possible.” Rohin Shah, the director of AGI safety and alignment at Google DeepMind, says the results of the report “should… …