| ▲ | danielmarkbruce 4 days ago | ||||||||||||||||
Ouput: 128 logits. Input: maybe 10 samples of plaintext,ciphertext (using the same key), so maybe a 2560 length tensor. Loss function: binary cross entropy on the true key bits. Architecture: anyone's guess. If you were in a place to debate this, you would have known the above (or something similar) is what I was suggesting when i said train on plaintext, cipertext -> key, and you'd have some deep mathematical insight as to why no architecture known is likely to work. And you would also know I wouldn't be here talking to you about it if I really had a solid idea of an architecture that is likely to work. | |||||||||||||||||
| ▲ | insanitybit 3 days ago | parent [-] | ||||||||||||||||
I'm not debating you at all. I'm asking what the model looks like since you've stated (and I've agreed) that a language model wouldn't work. I think it would make sense to explain how a theoretical model could do better than SAT. Otherwise, is the idea here just "magic is possible"? | |||||||||||||||||
| |||||||||||||||||