| ▲ | hn_acker 3 hours ago | |
The idea that human-readable explanations emitted by a language model don't necessarily correspond to the model's actual internal process of reaching a conclusion reminds me of parallel construction [1], a (fraudulent) law enforcement strategy of obtaining evidence of a crime through usually illegal means and claiming that the evidence was obtained legally through some other means. [1] https://www.hrw.org/report/2018/01/09/dark-side/secret-origi... | ||
| ▲ | 2 hours ago | parent [-] | |
| [deleted] | ||