| ▲ | andy12_ 5 hours ago | |
You can achieve open-vocabulary classification by making the final weights in the softmax come from a category encoder instead of being fixed learned weights. So instead of softmax(encode(input)*learned_weights) You have softmax(encode(input)*encode(categories)) I'm not sure if Jev does it this way, but it's how you get open-vocabulary zero-shot image classification with models like CLIP [1]. | ||