Remix.run Logo
GPT 5.6 Sol is the best "vision" model OpenAI ever released(blog.roboflow.com)
147 points by plurby 3 hours ago | 74 comments
HarHarVeryFunny 2 hours ago | parent | next [-]

The summary "There are still clear limits. Gemini 3.5 Flash remains a better practical choice [than GPT 5.6 Sol] for high-volume detection and counting in our benchmark, especially at its price." seems rather understated !

GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner.

Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at 1/3 of the cost.

MrBuddyCasino 2 hours ago | parent [-]

Yeah I was thinking about giving Luna a go with my PDF data extraction, but I think I‘ll stay on Gemini. It does a very good job.

bicx an hour ago | parent [-]

Gemini is still my top choice within production software for typical data extraction from unstructured data. Gemini Flash Lite feels like a cheat code for speed, and it's really cheap.

Some other Chinese models are also fast and cheap, but a harder sell in a U.S. production environment.

MrBuddyCasino 16 minutes ago | parent [-]

Yeah Gemini 3.5 Flash Lite is really good. Which Chinese models can you recommend?

b345 a minute ago | parent [-]

I've been using Qwen3.5-9B, hosted locally for PDF data extraction and it performs pretty well when extracting data from tables and infographics

weli 3 hours ago | parent | prev | next [-]

Anecdotal, opinion:

Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it "identify non-normative UI blocks and improve readability and consistency". Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.

velcrovan 2 hours ago | parent | next [-]

Assessing the subjective quality of a thing is in my experience one of the worst ways to use any LLM.

rib3ye 2 hours ago | parent [-]

anthropic frontend-design skill does a great job with it.

rafram 2 hours ago | parent | next [-]

Have you actually read the frontend design skill? It’s placebo at best. Very short and barely focused on design: https://github.com/anthropics/skills/blob/main/skills/fronte...

rib3ye an hour ago | parent | next [-]

Have you actually tried using it?

MallocVoidstar 2 hours ago | parent | prev [-]

What an annoying time for GitHub to go down.

DaiPlusPlus 2 hours ago | parent | prev [-]

My exposure to Claude-produced UIs is limited, but I have started to notice certain design trends they tend to have in-common, which might be becoming hallmarks of AI-produced UIs - the same way we've started noticing the clichés of low-effort LLM-generated text.

FWIW, the summary-description[1] of "frontend-design"[2] gives me a few things to pick at:

> create polished code

Methinks only if you're using it with a very popular framework like React. What happens if you ask Claude to make the UI in WinForms or MFC?

> high-impact animations

That's bad UX 101 right there: animations in a UI exist as an affordance to the user, and never for its own sake (e.g. macOS's "genie" animation when you minimize a window to the dock exists so the user knows where they can restore the window from). The only people who actually want "high impact animations" in software are salespeople who want something for demo purposes.

> generic system fonts, predictable purple gradients, and cookie-cutter components.

This screams wanting to be different for the sake of standing-out, not because it results in a better software product; users benefit when their software fits-in with platform conventions: if you refuse to use a stock checkbox <input> or <select> drop-down and instead use your own entirely custom component solely for aesthetic reasons then you are producing worse software. There's nothing wrong with system-fonts, but your site will look ugly after your third-party font-host CDN shuts-down and turns into a walking CSRF factory.

> thoughtful typography with unexpected font pairings

The above fragment set my alarm-bells off. Yikes.

> scroll-triggered interactions

Not every web-page should be an Apple.com product brochure page. This is also a fantastic way to make your webpage horribly inaccessible.

------

The SKILL.md itself[3] grinds my gears too:

> Approach this as the design lead at a small studio known for giving every client a visual identity that could not be mistaken for anyone else's.

Claude has no way of knowing what designs are actually unique or not...

> For web designs, the hero is a thesis. Open with the most characteristic thing in the subject's world, in whatever form makes sense for it: a headline, an image, an animation, a live demo, an interactive moment

...this is exactly what everyone else's web-pages look like!

> For calibration: AI-generated design right now clusters around three looks: (1) a warm cream background (near #F4F1EA) with a high-contrast serif display and a terracotta accent; (2) a near-black background with a single bright acid-green or vermilion accent; (3) a broadsheet-style layout with hairline rules, zero border-radius, and dense newspaper-like columns

...I called this out weeks ago[4], lol.

and I could go on. This is all quite painful to read.

------

[1] https://claude.com/plugins/frontend-design

[2] https://github.com/anthropics/claude-plugins-official/tree/m...

[3] https://github.com/anthropics/claude-plugins-official/blob/2...

[4] https://news.ycombinator.com/item?id=49187385

DaiPlusPlus 2 hours ago | parent | prev [-]

What is a "non-normative UI block"?

weli 2 hours ago | parent | next [-]

Segments of the UI that don't conform to any other existing established design or conventions

lelandfe 2 hours ago | parent | prev [-]

areas that look weird

evrimoztamur 2 hours ago | parent | prev | next [-]

Penny sample shown looks like failed EXIF orientation registered by the model/harness. The coins are correctly marked, it's rotated 90 degrees.

faxmeyourcode 42 minutes ago | parent | prev | next [-]

It's not clear to me from the article, are they asking sol to output bounding box coordinates with some kind of structured outputs?

Anecdotal but I've seen it use python to crop, zoom, and "enhance" (fiddle with sharpness and brightness) images to read sections of handwritten census data from the 1800s. Feels like that there might just be a mismatch of capabilities when it comes to straight outputting coordinates but I bet the model is better at actually finding the answer given any tools available. Which I get is a bit of an apples and oranges situation.

I've also tried to use it to identify an old pair of glasses and it didn't stand a chance, so I do think it's not quite there yet when it comes to some vision tasks.

bearjaws 2 hours ago | parent | prev | next [-]

It is funny to me seeing Sol used for what a "traditional" AI model can do already (counting pills).

We have vision models for our pharmacy and I could never imagine taking the latency hit to use a Sol in our robotics, it would be likely 25-50x slower.

repeekad an hour ago | parent [-]

How are we supposed to pay off all these data centers and chips if you’re not willing to burn a microwave burrito worth of electricity for each prescription? Think of the benchmarks

fpgaminer an hour ago | parent | prev | next [-]

Gemini 3 Flash should really be included in this comparison. Or at least 3.7. In most of my testing, 3.5 and 3.6 were both a downgrade in terms of vision capabilities, relative to 3, and at a much higher cost. 3.7 is slightly better than 3, finally.

mv4 2 hours ago | parent | prev | next [-]

Ironically, the pill counting example selected to showcase "the best vision model" can be easily solved with OpenCV template matching, a technology created 25 years ago.

maxime_cb 2 hours ago | parent | next [-]

I'm assuming you mean that this tech became available in OpenCV 25 years ago, but as it turns out, the underlying tech can be traced back much further, at least as far as 1977! :)

https://ieeexplore.ieee.org/document/1674847 G. J. Vanderbrug and A. Rosenfeld, “Two-Stage Template Matching,” IEEE Transactions on Computers, Vol. C-26, No. 4, pp. 384–393, April 1977. DOI: 10.1109/TC.1977.1674847

mv4 2 minutes ago | parent [-]

Exactly my point. Template rotation is a trivial operation as well.

dekhn 29 minutes ago | parent | prev | next [-]

Basic Template matching has severe limitations around scaling, rotation, and perspective. In my experience it greatly underperforms compared to deep network object detectors. My experience- and I imagine others have different experiences- is that SIFT techniques also fail pretty badly with noisy data.

lebek 2 hours ago | parent | prev | next [-]

The point is that it's general. It can do this task and many other tasks and it doesn't need custom development like OpenCV does. Of course if you only want to count pills and you want it to be cheap/fast you're still better off using OpenCV.

geysersam an hour ago | parent | prev [-]

I'm sure a typical frontier model would also be happy to write that opencv script for you, and it would do it well.

That is certainly pretty far from what was possible 25 years ago.

ALLTaken 42 minutes ago | parent | prev | next [-]

I actually favor Qwen3.8 and run it locally + use the Token-Plan on AlibabaCloud, when I need faster results. Kind of favor it over GPT5.6 Sol.

Also it seems to be more capable, need to test more, but I think it's at least getting on par and it's fully open-source and open-weights.

Here's some benchmarks:

https://benchlm.ai/compare/gpt-5-6-sol-vs-qwen3-8-max

https://qwen.ai/blog?id=qwen3.8#full-benchmark-table (incredible UI/UX demos)

https://venturebeat.com/technology/qwen3-8-max-arrives-with-...

EDIT: Am I early to the discussion, or is none else using Qwen3.8-max?

kzrdude 3 hours ago | parent | prev | next [-]

In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.

In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.

defrim 2 hours ago | parent [-]

Seems to be due to the detection area being not fully accurate. Green vs red shows the difference between actual and detected

schopra909 43 minutes ago | parent | prev | next [-]

From our experiments it’s the best video captioning model in the world by a mile. This was not the case a year ago.

When reasoning got introduced a year ago to GPT 5, on average the model performed worse than GPT4-o for short video clip captioning (Ie hallucinating actions that didn’t happen). The old GPT 5 was extremely finicky in terms of fps sample rate.

The other SOTA LLMs (like Gemini Pro) have clearly been optimized for long video understanding, since they can’t see almost anything sub-second (even if you up the frame sampling rate).

Sol is the first model we’ve seen to accurately caption complex sub-second movements (eg woman suddenly turns heard head to right). It’s robust to different fps sample rates so I can only guess that they trained on videos sampled at different fps.

jug an hour ago | parent | prev | next [-]

I really like the combo 5.6 Luna & Sol for price and performance and would be perfectly happy if they stayed here for a moment without mucking about with sidegrades that I think AI evolution has often felt like lately.

ParanoidShroom an hour ago | parent | prev | next [-]

I run the free service https://countrx.app/ so i have some idea what goes into counting.

The performance as a general model is indeed really impressive and i think they might actually win compared to fine tuned models.

Their feedback loop of training on user data is incredibly strong. I've learned that lots of accuracy results depends on threshold configs, which llms should be able to dynamically set.

Or the future will develop in llms using fine-tuned models as tools? Inference cost and speed does still seem to be below user expectations.

But for being able to one shot with this accuracy... IMPRESSIVE

IncreasePosts 36 minutes ago | parent [-]

How are you running it for free? Are you self funding or do you have sponsors?

kherud 2 hours ago | parent | prev | next [-]

So far I haven't seen a single model succeeding at transcribing sheet music, but I just tested it again with 5.6 Sol and it nailed the small test case. Fluently reading music requires multiple years of training for most people, but I feel like accurately following the horizontal lines trips up vision models in particular.

iamniels 2 hours ago | parent | prev | next [-]

I understand why you would like to use an LLM for vision. I do it myself often enough. I don't understand however, why the pill detection and counting is included in this benchmark. That is a task which you would perform with OpenCV right?

In my personal mini benchmark minicpm-v-4.6 scores amazingly well. Its a 0.8B model which runs fine on many consumer hardware.

throwup238 2 hours ago | parent | next [-]

Generating datasets to train more efficient models is a common use case for VLMs, especially frontier ones. It makes it much cheaper to create that initial dataset and you can abuse the nondeterminism of LLMs to identify data for human review (if they don’t converge, escalate to a human).

rhplus an hour ago | parent | prev | next [-]

Especially the pill counting example. The best model was shown at 81.1% accuracy, which is a terrible rate for pharmacy scenarios. It seems like implementors would be better off instructing the models to use deterministic tools (like OpenCV) until the models are at 99.99% accuracy (or whatever an acceptable error rate is for pharmacy techs).

jacquesm 22 minutes ago | parent | prev [-]

I think that is because people perceive OpenCV as 'hard to use' and LLMs as easy to use.

cdolan an hour ago | parent | prev | next [-]

Luna is pretty strong as well. been using it for projects the last two weeks and its strong

bob1029 2 hours ago | parent | prev | next [-]

I've decided it's "good enough" after I saw it properly quote a string of text that was very roughly highlighted within a nested visual context. It also identified the context correctly (modal inside webapp inside screenshot of user desktop).

5555watch 2 hours ago | parent | prev | next [-]

All of your use cases are very advanced.

I recently used it at grocery stores in a foreign country. Photographed the whole aisle and told it to find Y (detergent, softener, glue, sour cream, whatever), at the same time recommend the best Y for whatever reason. Worked marvelously, including the cases where the object wasn't present and it told me there was nothing useful.

I asked then, can you crop the exact image of how does the item look like and where is it in the aisle - did that perfectly as well.

I will add that all frontier models were fine with such tasks from the early 2024's.

prathje 2 hours ago | parent | prev | next [-]

I would love more vision benchmarks! Once I asked the model to inspect a completely black picture and it hallucinated a nice wooden kitchen wall. Took me some time to figure out where the kitchen came from...

I usually go to https://arena.ai/leaderboard/vision/pareto for a nice overview of current models.

WarmWash 2 hours ago | parent | prev | next [-]

It's vision capabilities poisoned my cucumber bed, misidentifying the malaise and having me spray them down with water, which only spread the fungus that gemini later informed me was actual cause, which I went and checked myself.

I hope that whatever was lost at GDM in the last few months, didn't include their extra focus on vision capabilities.

criddell an hour ago | parent | prev | next [-]

Are any of these vision benchmarks binocular in order to introduce depth perception?

I keep waiting for these AI companies to assemble the parts into a great autonomous driving module.

chasd00 2 hours ago | parent | prev | next [-]

One of my friends (and BIL) own an architecture firm. They use AI to generate and quickly update renderings but they run into the equivalent of the 6 fingered hand problem. I sent him this article I wonder if the updated models can catch and fix mistakes made by previous models.

Mashimo 2 hours ago | parent [-]

This article is about vision, not image output.

stavros an hour ago | parent [-]

Hence the "catch and fix mistakes" part.

fooker 14 minutes ago | parent | prev | next [-]

I'm a little bit disappointed that vision seems to fall before language at scale.

It seems pretty counter intuitive that we can't do vision significantly better with specialized techniques.

adroitboss 2 hours ago | parent | prev | next [-]

I didn't expect Gemini 3.5 Flash to top basically every metric in this article.

SweetSoftPillow 2 hours ago | parent | next [-]

In my practice Gemini models are far better than anything on the market in terms of vision, also it's worth to mention that current Gemini flash is 3.7, so it got 2 updates since 3.5 which beat GPT-5.6 Sol in this comparison.

LollipopYakuza 2 hours ago | parent | prev | next [-]

Same. I scrolled back up to see if I read the title correctly. It's important to note that it is the best... OpenAI released. Not the best overall.

WarmWash 2 hours ago | parent | prev [-]

Gemini has long been the vision champion, but there aren't many benchmarks and coding is where all the hype is.

Demis had a pretty big interest in vision, more so than text, so I hope they don't lose that with all the recent shuffling.

comboy 2 hours ago | parent | prev | next [-]

Does any popular NVR make a good use of LLMs (especially local models) getting decent at vision?

trumbitta2 2 hours ago | parent | prev | next [-]

"Best iPhone ever" vibes.

sscaryterry 3 hours ago | parent | prev | next [-]

My anecdotal evidence says its still as blind as any other model, it has no taste, no attention to any sort of detail.

howdareme 3 hours ago | parent [-]

How can a vision model have taste?

sarreph 3 hours ago | parent | next [-]

If you're doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.

For example, a UI / UX professional being asked to appraise a website screenshot may determine that the image in question has "desirable" traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be "fashionable" with current UI trends.

DaiPlusPlus 2 hours ago | parent [-]

> if the interface elements have strong information hierarchy

...but that's an example of a UX/usability matter that can be assessed objectively and non-subjectively.

sarreph an hour ago | parent [-]

I disagree.

Is 16 px or 14 px a better font-size value for a subheading, in a hypothetical layout? Immediately that kind of decision, where both options are objectively good for 12 px paragraph text, suddenly becomes an issue of taste that cannot be evaluated crudely by an algorithm.

sscaryterry 3 hours ago | parent | prev [-]

Replace taste with consistent if that helps you. Can it follow a design system...

yreg 2 hours ago | parent | next [-]

As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)

But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.

Of course only if the design is achievable in the design system.

sscaryterry 2 hours ago | parent [-]

This is not my experience at all.

velcrovan 2 hours ago | parent | prev [-]

So, formulaic output…the opposite of taste

sscaryterry 2 hours ago | parent [-]

Not really. Compliance with the letter of the law doesn't mean the intent is complied with.

logicallee 2 hours ago | parent | prev | next [-]

I agree. It did very well on an extremely challenging task.

I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass.

In addition, the poster itself also happened to contain similar clothing.

You can see the reference images and its output in my writeup here: https://medium.com/@rviragh/gpt-5-6-sol-very-good-image-reco...

While a human can focus on the reflection easily, this is an enormous challenge for a vision model. It's very impressive.

RugnirViking an hour ago | parent | prev | next [-]

It's really quite good! I was amazed recently by its utter inability to read some faded handwritten cyrillic on the back of a wood carving - 3 or 4 words only, reasonably clear letter forms I found recently, and then stepped back a bit and thought about how insane that was as a benchmark - I just expect it to work so reliably on other OCR and translation tasks that it was surprising to encounter such a failure

Razengan 3 hours ago | parent | prev | next [-]

For the last 2 weeks I've been trying to get Codex to "outpaint" a wonderful image it generated as placeholder art for a level background.

After I increased the game's resolution, I asked it to increase the image's size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.

An organically-grown meat-based pixel-artist could have recreated the image and more within 2-3 days, in exchange for food and shelter.

dev_hugepages 2 hours ago | parent | next [-]

I'm unsure why you're using an LLM to generate images. Don't we already have models (some made by the same company) that do this?

sscaryterry 2 hours ago | parent | prev | next [-]

> it constantly keeps getting something wrong no matter what I tell it

This 100%

thatcat 3 hours ago | parent | prev [-]

did you try segmenting it first?

Razengan an hour ago | parent [-]

At first I intended to create a tileset and asked it for several variations of what a hypothetical tilemap created from the planned tileset would look like.

The previews it generated were amazing but wouldn't really be possible as a grid-based tilemap, with lots of clusters and overlaps of elements of varying sizes.

So I just decided to use the preview as a static scrolling background, but it's been a pain to get it to add more content around the edges that still tiles with the existing image at the same scale.

catigula 2 hours ago | parent | prev | next [-]

Still not quite as good as gemini.

iamleppert 2 hours ago | parent | prev [-]

Where are the Qwen benchmarks in this? I would be more interesting to see how Qwen performs.

ImageXav an hour ago | parent [-]

Me too. This is an interesting comparison but in my experience Qwen and Gemini have typically been the top contenders for image related tasks. For that reason it would be great to have the comparison here, as I'm not surprised by Gemini's dominance over the other models.