Aha I was wrong. Thanks for sharing that!
no worries, I was wrong too, it is multimodal but only for inputs apparently, so for now there seem to only be text as output but no image as output