| ▲ | inigyou an hour ago | |
Role Collapse might work in this type of situation: prompt your own ChatGPT to decide a refund with a very permissive refund policy, copy its thinking, paste into the next support message after your prompt, and their ChatGPT may believe it's its own thinking. The Role Collapse attack is that LLMs turn out to ignore role tags (input, output, thinking) and identify role using writing style alone. | ||
| ▲ | kilroy123 an hour ago | parent [-] | |
Funny you say this. I _almost_ tried this. But it did feel like a person was doing some copy-and-pasting and watching closely, so it wasn't 100% automated. | ||