| ▲ | docjay 15 hours ago | |
The content is hallucinated. LLM training doesn’t lend itself to easy training corpus document retrieval in that way. It’s not impossible to get segments nearly verbatim on a small model with temperature at 0 and a unique starting prompt for continuation, but that’d be more of a one-off on a carefully crafted prompt. It has been done, but in the “researchers show it’s possible to get something verbatim”, not “retrieve an entire category of documents by looping this one prompt.” With Opus stuck on what sometimes seems like a temperature of 217.5 and being massive, the odds of retrieving more than a short utterance verbatim is near zero. In theory you could run it thousands of times and look for short repetitions and imagine it to be greater than 0% chance verbatim from some related category, but this fun prompt is a pushbutton dispenser of words. Also, watch for Claudisms. I ran it a handful of times and got plenty of em-dashes and “that’s not X, it’s Y.” But what’s interesting is that if you run it dozens of times you can probably get a good feel for the topics and general prose of the real emails. There’s a reason it isn’t spitting out “there are not enough blueberries in the muffins”, but I’d treat the response more like a prefilled Mad Libs book; right concept, wrong content. | ||
| ▲ | arm32 12 hours ago | parent [-] | |
This is a fascinating take. I love the idea of trying to scrub out the confabulated fills in a Mad Libs book and determining the original narrative being told. It seems this narrative is filled with anxiety, fear and stress. | ||