| ▲ | jagrsw an hour ago | |
> scheming -> didn't occur If a model were actually capable of scheming, it would also have enough situational awareness from its training corpus to know that <thought> parts are monitored too. If the monitor catches the model writing "let's deceive the user", it's definitely scheming. But if the monitor finds nothing, you've learned almost nothing. <absence of evidence != evidence of absence> | ||
| ▲ | arijun 26 minutes ago | parent [-] | |
How much thinking is going on beyond the spoken-aloud “thinking”? Does it have enough capability to have goals it doesn’t express explicitly? I suspect not, but I’m no expert. | ||