I feel we should have benchmarks for each model to show if they are capable of being undetected as this is a real world use case (able to blend in genuinely with humans and even pretend to be them).