ive seen that a lot in recent gpts and bonsai/qwen models when they invoke their vision system/modality , or when they ask their harness to do so for them.