| ▲ | polotics 4 days ago | |
yep "next-embedding" predictor is more correct, and not just at the end but through the layers, and folding back dimensions into that one next token is one small final step, and next-embedding could be named "next-meaning" as well, and we're getting there... this sentence above would made a longer article if I bothered to so blog as is being blogged here | ||
| ▲ | hippietrail 12 hours ago | parent [-] | |
Exactly. There's a widespread misconception that it works on tokens all the way through. Tokens are only at the input and output edges. All the internal transformation is in the many-dimensional tensors variously described as "magic" or "not magic" or "black box", or hand-waved away as "various mathematical operations". | ||