| ▲ | vunderba 6 hours ago |
| One of the things they seem to be emphasizing here is the UX around being able to place specific elements where you want them in an image. If the positions of the components in the overall composition are very important, this seems to make that a lot easier and kind of reminds me of InvokeAI. Ideogram V4, an open-weight model released back in June can also do this [1], but you have to use a relatively cumbersome JSON structure to describe all the different bounding boxes. So it’s definitely a bit of a hassle. I'll probably be waiting until it goes open-weight (hopefully soon) like they did with Flux.2 / Klein. [1] - https://docs.ideogram.ai/using-ideogram/getting-started/prom... |
|
| ▲ | kranke155 5 hours ago | parent | next [-] |
| You just get an LLM to do the bounding box stuff or use the ComfyUI node that provides a GUI for bounding box generation |
| |
| ▲ | CuriouslyC 5 hours ago | parent | next [-] | | I always hated Comfy's node based UI, but agents make it tolerable. Now I just have them set up a workflow and I go in and tweak it manually if the results aren't where I want them. I even have agents cherry doing multiple runs and cherry picking the best outputs, models have gotten good enough that it's a real time saver, assuming you have references they can and a rubric to check against. | | |
| ▲ | jarjoura an hour ago | parent | next [-] | | It's definitely not for me. From where I'm sitting, it's just turning python functions into boxes and instead of write the function yourself, you drag from the output of one box to the input of another. For 2 or 3 boxes, this is cool, but I opened up a professional workflow and was taken into a view with 100s of boxes and wires all over the place. Uhh, ok? For myself, I'd rather just create my own python environment, write some quick pytorch or mlx calls, wire up some cli to it and share that in a GitHub. | |
| ▲ | swiftcoder 3 hours ago | parent | prev | next [-] | | > I always hated Comfy's node based UI It's one of the most uniquely hostile user experiences I've ever had the (dis)pleasure of working with | |
| ▲ | hdjrudni 3 hours ago | parent | prev [-] | | How do you get agents to set up a workflow? You just get them to modify the JSON directly and then import it, or do you have a tighter integration (e.g. in the UI)? | | |
| ▲ | CuriouslyC 3 hours ago | parent [-] | | The agents can interact with Comfy via API pretty well, which afaik ends up being directly with JSON. |
|
| |
| ▲ | vunderba 5 hours ago | parent | prev [-] | | Yeah, given how much better Ideogram v4 outputs are when you use the proper structured JSON (background, elements, etc.) I think most users probably have stuck some kind of Qwen/Gemma-based LLM between their raw prompt and the CLIP encoder. |
|
|
| ▲ | NBJack 3 hours ago | parent | prev | next [-] |
| Adobe has had something like this for a while with Generative Fill. Drag a box, specify contents. |
|
| ▲ | popalchemist 5 hours ago | parent | prev [-] |
| Ideogram 4.5 released 2 days ago with more features along these lines https://www.youtube.com/watch?v=2mecWZgbaEg |