> Just write a prompt. Text-to-image with strong prompt following and a native understanding of composition.
How is this not clear?