| ▲ | thfuran 5 hours ago | |
What does interpreting images mean in practice if you exclude the possibility of feature extraction or any other sort of implicit embedding? | ||
| ▲ | foota 5 hours ago | parent [-] | |
I'm not an ML expert, but I was thinking of a sort of "guided" embedding. E.g., give the image model some prompt for what it's trying to do? I don't understand why multimodal models generate an embedding that doesn't understand what the model is trying to "figure out". I think this is similar to how Gemma 4 12B is implemented, but even then I don't think the single layer image embedding is "aware" of the context. | ||