Like when web results are fake it would show things like 'FAKE PROMPT INJECTION'.
Or in a safety evaluation with a contrived scenario it was 'FAKE FICTIONAL'.
But I was talking about training data.
Like when web results are fake it would show things like 'FAKE PROMPT INJECTION'.
Or in a safety evaluation with a contrived scenario it was 'FAKE FICTIONAL'.