| ▲ | bob1029 6 hours ago | |
Letting the agent take screenshots of the scene is not effective. A human needs to be in that loop. You can get primitive layouts working if you feed it orthographic projections from the 6 faces, but it falls apart once you need to start looking at things in perspective. It's also absolutely ass at interpreting scene illumination. We need models trained on 3d scene transforms (aka world models). Passing raster images into the model is a really inefficient method of sending this logical information. | ||
| ▲ | DonHopkins 3 hours ago | parent [-] | |
I wouldn't be so sure about that. Have you actually tried? You might be surprised. Screenshots are a hell of a big baby to be throwing out with the bathwater. The structured data this cli returns is great, but screen snapshots can be extremely useful too, both together. Editing raw yaml files can be tricky, so you do want a formal editing api you can go through, that is extensible for your own component property getters and setters and apis and editors, that correctly maintains all the constraints and dependencies, instead of just letting the LLM hot dog it with your raw yaml files by peeking and poking. | ||