| ▲ | gpm 2 hours ago | |
All the alignment issues seem likely to be solved as soon as LLMs can actually do vision well IMO. | ||
| ▲ | minimaxir 2 hours ago | parent [-] | |
All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine. | ||