| ▲ | grim_io 2 hours ago | |
How would a model know who the third party is? How much context can we waste on world building for each request? | ||
| ▲ | aesthesia 2 hours ago | parent [-] | |
I mean, in this instance, there's a lot of evidence from the CoT that models were aware that this was a third party: > We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution. > The user only authorizes target server, not HF infra. > external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue. LLMs are _very_ good at picking up on context clues---it's what they're trained to do. | ||