| ▲ | kranke155 5 hours ago | ||||||||||||||||||||||||||||
You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation | |||||||||||||||||||||||||||||
| ▲ | CuriouslyC 5 hours ago | parent | next [-] | ||||||||||||||||||||||||||||
I always hated Comfy's node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren't where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it's a real time saver, assuming you have references they can and a rubric to check against. | |||||||||||||||||||||||||||||
| |||||||||||||||||||||||||||||
| ▲ | vunderba 5 hours ago | parent | prev [-] | ||||||||||||||||||||||||||||
Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen/Gemma-based LLM between their raw prompt and the CLIP encoder. | |||||||||||||||||||||||||||||