Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Anthropic's J-Lens research indicates so.

Like when web results are fake it would show things like 'FAKE PROMPT INJECTION'.

Or in a safety evaluation with a contrived scenario it was 'FAKE FICTIONAL'.



Sure. That is easy.

But I was talking about training data.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: