| ▲ | DeepSeek-v4-flash-vision-exp(api-docs.deepseek.com) |
| 396 points by dares2573 9 hours ago | 134 comments |
| |
|
| ▲ | ciberado 8 hours ago | parent | next [-] |
| DS being unable to precisely view Playwright screenshots is the only thing I really miss from Sonnet. This is promising. > Images are converted into tokens based on their dimensions, and these tokens are billed together with your text tokens. > Before inference, every image is automatically resized: > - Images with a total pixel count below roughly 384×384 are scaled up while preserving their aspect ratio. > - Larger images are scaled down while preserving their aspect ratio so that the total pixel count after resizing is roughly that of an 800×800 image. > As a result, there is an upper bound of 384 tokens per image: for example, a 2000×2000 image and a 5000×5000 image consume the same number of tokens after resizing. When a request contains multiple images, each image is counted independently under the same rule—there is no separate calculation for multi-image requests. 400 tokens per image results in 2,500 images per dollar, if I’m not mistaken. edit: format. |
| |
| ▲ | knollimar 8 hours ago | parent [-] | | Oof 800 by 800 kills a lot of use cases | | |
| ▲ | johndough 8 hours ago | parent | next [-] | | Might still be fine. The most recent crop of vLLMs proactively use whichever programs are available on the system (e.g. ImageMagick or PIL) to "zoom in" by cropping subimages if they can't quite make out the details. | | |
| ▲ | knollimar 5 hours ago | parent [-] | | Downsizing a higher res image to lower res means the zoom will be blurry. | | |
| ▲ | johndough 2 hours ago | parent | next [-] | | The order is: LLM issues tool call to read high res image ->
harness sends high res image to server ->
server downsizes it to 800x800 (blurry) ->
LLM issues bash command (e.g. `convert`) to crop a small subimage (e.g. 600x600) from the high res image ->
LLM issues tool call to read subimage ->
harness sends subimage to server ->
server does not resize the subimage because it is small already, so it is not blurry when finally ingested by the LLM
| | |
| ▲ | knollimar 11 minutes ago | parent [-] | | Then you have a separate issue where the LLM can't piece together 9 subimages well. |
| |
| ▲ | andai 5 hours ago | parent | prev | next [-] | | They process the original image file with Python on the local device. (And I've seen the web chats do this with their "computer use" features too.) The really wild one is even blind models will do this and they'll try to run stats on the pixels to figure out what it looks like... the even wilder thing is that it kind of works! | | |
| ▲ | knollimar 4 hours ago | parent [-] | | If the API accepts only 800 by 800, the aegument youre making is "fix it in the harness". I don't think the n by n subgrid fixes this the way most harnesses do, as it'll fail to count things if you have more overlap and fail relatiomships if you have less | | |
| ▲ | tjoff 2 hours ago | parent [-] | | That seems weirdly specific? And if you are counting things it should be trivial to note the position of your items and not double-count them, no? |
|
| |
| ▲ | adastra22 3 hours ago | parent | prev [-] | | They’re not talking about zooming, hence the quotes. | | |
| ▲ | johndough 2 hours ago | parent | next [-] | | Yes. When the LLM tries to read an image, it will be resized by DeepSeek's server to 800x800, which might be a bit blurry. The LLM will then crop a smaller image from the high resolution image (using e.g. the `convert` tool via bash) and will then read the small cropped image. This image will still be resized to 800x800 by DeepSeek's server, but since it is already small, there is no or little loss of quality. | |
| ▲ | knollimar 3 hours ago | parent | prev [-] | | If the harness does it that's just like saying "please use a workaround". You'll lose fidelity and LLMs will lose the ability to count things or maintain relationships for schematics, etc | | |
| ▲ | johndough 2 hours ago | parent [-] | | LLMs read images by splitting them up into e.g. 16x16 patches, which are then converted to embedding vectors and fed to the LLM, so from a technical point of view, feeding a big image as many 20x20 patches all at once is not too different from cropping subimages from the image, splitting those subimages into patches and feeding them to the LLM. Of course, the LLM has to be trained to understand that those images belong together, but it can be done. | | |
| ▲ | knollimar 9 minutes ago | parent [-] | | You say "it can be trained" but they fail at counting in my use cases, let alone maintaining symbolic relationships. Can you point me at one that can understand a detailed block diagram? Frontier is fine, soliciting recommendations |
|
|
|
|
| |
| ▲ | wongarsu 8 hours ago | parent | prev | next [-] | | For most use cases you can fix that in the harness. Just give the model a tool to request a crop of specific coordinates of any image it has in its context. Call the tool "zoom" and it should be intuitive for the model Maybe there are some use cases where you need high detail everywhere at once, but for OCR of small text and the like a zoom ability should be sufficient | | |
| ▲ | embedding-shape 7 hours ago | parent [-] | | For really dumb models I've also had success automatically cropping it into a grid of N images with the max size, then processing each cell individually, then once all been processed, do one final call with resized image + all other context previously generated per cell. Basically a workaround to the image dimension restrictions without loosing fidelity. Works well with even dumb 7B models. Can't remember if I stole this idea from some existing public harness though, can't remember. If someone knows of public harnesses that do this already, please share them :) | | |
| ▲ | Aeroi a minute ago | parent | next [-] | | i think the claude cookbook has a file that does this, called tiling. | |
| ▲ | dotancohen 7 hours ago | parent | prev [-] | | Does this not loose context? Especially e.g. in fonts where the character pairs 0O 1I 1l Il may be difficult to differentiate? | | |
| ▲ | skeledrew 6 hours ago | parent [-] | | That's what the grid crop should handle. The detail is retained at that level, and then everything is logically stitched together again using the lower-res-full-image as reference. That's going to be 2x token usage at minimum though. |
|
|
| |
| ▲ | stronglikedan 3 hours ago | parent | prev | next [-] | | I don't know about a lot. Probably more like a few. I take a lot of screenshots for various reasons, and over 800 seems like I could have done a better job framing and cropping. | |
| ▲ | shadyr 8 hours ago | parent | prev | next [-] | | It might also be due to its experimental status. Wouldn't surprise me if the GA version allows for larger input. Either that or the eventual pro version. | |
| ▲ | Chnmy 6 hours ago | parent | prev | next [-] | | what are these use cases? | | |
| ▲ | knollimar 5 hours ago | parent [-] | | Anything where there are symbols representing in space (e.g. schematics). Thats pretty broad |
| |
| ▲ | asdfsa32 8 hours ago | parent | prev [-] | | flash vs fine details. Pick one. | | |
| ▲ | Doohickey-d 8 hours ago | parent [-] | | Gemini "flash" models have an option for media resolution, including a high resolution option for screenshots. | | |
|
|
|
|
| ▲ | leumon 5 hours ago | parent | prev | next [-] |
| It fails the simple clock test for me which Qwen3.8 27B got (nearly) right.
given an image of a clock https://files.catbox.moe/kgwa5e.png I asked it "what time does the clock show?" (both on reasoning: high) DS answered: The clock shows *5:10* (and 45 seconds).
Here is the breakdown:
* *Hour hand (red, shortest):* Pointing at the *5*.
* *Minute hand (green, longest):* Pointing at the *2*, which represents 10 minutes.
* *Second hand (blue, medium):* Pointing at the *9*, which represents 45 seconds. Qwen answered: The clock shows *8:10* (with the red second hand on the 5, i.e. *8:10:25*). - *Hour hand* (short, blue) → 8
- *Minute hand* (long, green) → 2 (10 minutes)
- *Second hand* (thin, red) → 5 (25 seconds) Correct answer is 08:09:25. |
| |
| ▲ | bel8 29 minutes ago | parent | next [-] | | Well it succeeded on first try for me: https://i.imgur.com/gljOYr9.png [Image 1] what time does the clock show?
+ Thought: 368ms
The clock shows 8:10.
- Blue hour hand: just past 8
- Green minute hand: pointing at 2 (10 minutes)
- Red second hand: pointing at 5 (25 seconds)
▣ Plan · DeepSeek V4 Flash Vision Exp · 4.4s
edit: I now ran the test 10 times total, fresh session every time, and DS4 got it right 9 out of 10 times. Lesson learned to double check what I read on the internet.I'll leave a copy of the clock image for posterity here in case anyone wants to test it themselves: https://i.imgur.com/BQsfa3R.png | |
| ▲ | dghlsakjg 5 hours ago | parent | prev | next [-] | | I’ll keep that in mind next time I need to tell what time it is by asking an llm to read an analog clock. Snark aside, I’m not sure that these gotcha tests are any more useful than asking politicians gotcha questions. Sure, the model can’t tell me what time it is, but it can code the Wang algorithm for noisy audio matching in one shot. Maybe this is just me being an optimist, but this is my hiring philosophy and I guess maybe now my llm philosophy: I’m not interested in seeing how dumb I can make you look, I’m more interested in how smart you can be. | | |
| ▲ | jubilanti 4 hours ago | parent | next [-] | | Reading any analog clock at any time level (edit: and a non-noisy vector rendered image at that) is absolutely table stakes for an allegedly frontier flagship vision model. As much as 1:1 OCR. If the model can't do that, there's something wrong. Doesn't matter if it's memorized some random thing you think is esoteric but is in all the training data and benchmarks. The whole point of LLM/FMs vs good old fashioned ML is generalization to unknown domains, not just unknown tasks. The hunt for "gotchas" is the hunt for "not in your training data". | | |
| ▲ | dghlsakjg 3 hours ago | parent | next [-] | | Is this an “alleged frontier flagship vision model”? This is described as a brand new flash model - still experimental - from a lab that is a side project for an investment firm that has never had a vision model before. That doesn’t scream flagship or frontier to me. | |
| ▲ | bel8 3 hours ago | parent | prev [-] | | I disagree. It's not even that useful to train LLMs to read an ancient analog clock. Unless we're talking about AGI, I couldn't care less if an LLM is bad at things they won't be doing anyway. I'd rather focus training data on more useful tasks. | | |
| ▲ | kadoban 34 minutes ago | parent [-] | | It's not that useful for a person to be able to read an analog clock either, but if a supposed-genius came to me and confidently gave the wrong answer, it would say something about their strengths/weaknesses in general. The whole point of these models is they're meant to be able to generalize fairly well, not just answer questions the got trained on. | | |
| ▲ | bel8 23 minutes ago | parent [-] | | Funny, I just ran the test myself and DS4 got it right first try. Then I ran it 9 more times and it got right 9 out of 10 times. Lesson learned to double check what I read on the internet. |
|
|
| |
| ▲ | atomicnumber3 4 hours ago | parent | prev | next [-] | | It's because the messaging for what the point of these things is supposed to be is all over the place. Ask 10 different people and you'll get 10 different answers: - A superintelligence that will usher in an age of human enlightenment - A superintelligence that will usher in an age of human enslavement - A really cool way to rake in trillion of rich VC/investor money by promising you're building a superintelligence that will usher in an age of human en[slave/lighten]ment - A transformer model for predicting output tokens given a series of input tokens, informed primarily by reddit, stack overflow, and 6000 years of classical literature. - A replacement for white collar labor. Start now or join the permanent underclass. - A convenient fuzzy-find tool also capable of some probably-correct code generation. - The ultimate customizable text RPG experience (you can pick if G stand for game or...) And so on. So, some people see a new model and check for how close humanity is to enslavement. Some people check to see if it got better at fixing broken unit tests. | | |
| ▲ | KerrAvon 4 hours ago | parent [-] | | What's amazing is that all of these are true at once. If you allow for some significant slack in what "superintelligence" means. |
| |
| ▲ | altruios 4 hours ago | parent | prev | next [-] | | Knowing where it fails is just as important as knowing where is excels. | |
| ▲ | mejutoco 4 hours ago | parent | prev | next [-] | | It is like asking a politician how much a coffee costs, to show how disconnected they are from common people. Super intelligence not being able to read a simple analog clock does the same. > Sure, the model can’t tell me what time it is, but it can code the Wang algorithm for noisy audio matching in one shot This is about a _vision_ model. | |
| ▲ | skybrian an hour ago | parent | prev | next [-] | | I find collecting “gotcha” questions helpful because when a new model can answer them, it shows improvement, and if it doesn’t, it’s a reality check that, despite the model being helpful for many tasks, there are other things it can’t do yet. It’s a demonstration of “jagged intelligence.” It doesn’t have to be a negative thing! Simon’s pelican on a bicycle prompt is an example of a “gotcha” question. | |
| ▲ | ndriscoll 2 hours ago | parent | prev | next [-] | | I was literally working on an educational game for my kids last week where one of the activities is clock reading, and I ask codex to QA its Godot program via screenshots, so literally this exact scenario is something I was doing in a software engineering context. It can of course write code to figure out the angles to rotate by just fine, but it also needs to be able to figure out whether the whole picture comes together, whether the hand sprites are anchored on the clock face correctly with the right pivot, etc. | |
| ▲ | vasco 5 hours ago | parent | prev | next [-] | | It's still useful to find things it can't do if anything so we can tell when it starts being able to do them. | |
| ▲ | CooCooCaCha 4 hours ago | parent | prev [-] | | Is being asked to read a clock really a gotcha? | | |
| ▲ | arjie 4 hours ago | parent | next [-] | | Not if you are aiming at a general intelligence but it’s worth considering that this is a tool that may not be able to count the number of strawberries in the letter R but can still center a div. | |
| ▲ | dghlsakjg 4 hours ago | parent | prev [-] | | If I had to hire an engineer and there was one that could one shot the wang algorithm, but couldn’t read an analog clock, I would have no problem hiring them. Also worth noting that both models got it wrong. Qwen made a mistake that humans very good at reading clocks would make. Deepseek made a mistake that a human who had just learned to read clocks would make. | | |
| ▲ | CooCooCaCha 3 hours ago | parent [-] | | It would be different if AI was known to be reliable but it isn’t, so this is less of a random failure and more a symptom of jagged intelligence. And with every one of these there’s always an attempt to minimize the problem by saying it’s just one silly failure. |
|
|
| |
| ▲ | emosenkis 5 hours ago | parent | prev | next [-] | | This is not a normal looking clock - most clocks have either one color for all hands (second hand is thinnest and maybe also longest) or one color for hour/minute and one for second.
I know that the hand lengths and thicknesses on this image are correct but for some reason I, a totally human person who grew up when analog clocks were still common, see this and think the hand on the 5 is the minute hand.
How does the AI do if you just make all the hands black? | | |
| ▲ | leumon 5 hours ago | parent [-] | | then deepseek answers: "The clock shows 8:25.
The short hour hand is pointing to the 8, and the long minute hand is pointing to the 5, which represents 25 minutes." and qwen still answers: "The clock shows *8:10* (with the second hand on the 5, i.e., 25 seconds).
- *Hour hand* points to the 8
- *Minute hand* points to the 2 (= 10 minutes)
- *Second hand* points to the 5 (= 25 seconds)
So the time is *8:10:25*, or simply *8:10*." | | |
| ▲ | dghlsakjg 5 hours ago | parent [-] | | Qwen still got the wrong answer, though. Are we more forgiving because it’s the same type of mistake a human would make? |
|
| |
| ▲ | ttul 5 hours ago | parent | prev | next [-] | | A good share of humanity would have also gotten this question wrong! | | |
| ▲ | mdp2021 5 hours ago | parent | next [-] | | It's been four years that we are looping those "The professional failed its task!" // "Laymen would have failed it too". Which makes no sense. | |
| ▲ | andai 5 hours ago | parent | prev [-] | | Yeah, I heard most kids these days can't read analog clocks either. I can't actually remember where I learned to read a clock, it might have actually been in school. I guess that means they don't teach it anymore. (Everyone's phone shows the time anyway...) |
| |
| ▲ | mkatx 4 hours ago | parent | prev | next [-] | | I would say, try without thinking on. I find reasoning on any rag type request seems to increase hallucinations, probably due to the thinking tokens taking attention away from the, in this case, vision tokens. I'd recommend non-thinking for any non-prompt input, and leave the thinking where it has to actually reason. | |
| ▲ | ComputerGuru 5 hours ago | parent | prev | next [-] | | Gemini 3.7 Flash and 5.6-Sol (on all reasoning levels) also answer 8:10:25. The new "stealth" Ox Alpha also replies with the same. Opus 5 replies with 8:10 (no seconds). Not sure why this is so hard for them; Gemini is especially good at vision and I would have expected better from it. | |
| ▲ | johnnyApplePRNG 4 hours ago | parent | prev | next [-] | | I was wasting hours yesterday trying to get DeepSeek V4 Flash (with Qwen 3.8 27b as the vision agent, actually) to read sheet music to pass a Terminal Bench 3 benchmark and none of it was working... nothing... I changed models to gemma 31b, I tried OCR models... nothing could get it... And then I realized, wait a second... you're testing the harness not only against a difficult benchmarking problem, but it's one you're literally never going to use the coding harness for either, lol. I don't write programs that read or interact with sheet music and I never will. tl;dr Being frustrated that a "state of the art" vision model doesn't have perfect vision is a fools errand. It can read and extract information from screenshots and PDFs just fine (my setup). No need to worry about edge cases. | |
| ▲ | segmondy 5 hours ago | parent | prev | next [-] | | most likely a preview. they often release the preview via API, get more training data, post train some more then release the weight. i would expect to see it perform better in a few weeks or a month. | |
| ▲ | nubg 5 hours ago | parent | prev [-] | | welp, damning indictment. not sure if that means DS is super crap, or qwen is super good | | |
|
|
| ▲ | LorenDB 8 hours ago | parent | prev | next [-] |
| I've heard that DeepSeek v4 Flash 0731 has frequently assumed that it has vision capabilities and then resorts to inventing text-based image analysis tools when it finds that it actually can't see. In that case, this is a great upgrade for the model. Anecdotally, I had to tell 0731 to refrain from viewing screenshots since it kept breaking its sessions by trying to read images. |
| |
| ▲ | VulgarExigency 8 hours ago | parent | next [-] | | It tried to recreate vision by analyzing pixels on 3 separate projects I had it working on. | |
| ▲ | trollbridge 6 hours ago | parent | prev | next [-] | | I've mitigated this by giving it a "skill" that just means the harness using a different model. | |
| ▲ | mavamaarten 5 hours ago | parent | prev [-] | | Yeah I've seen it a lot. It goes through the effort, unasked, of pulling screenshots off a connected device and then it's like... Oh shit yeah I can't see. | | |
| ▲ | johnnyApplePRNG 4 hours ago | parent [-] | | It's doing it's best to accomplish whatever task you've thrown at it. It's expecting you to have done at least something besides select DS4 on Ollama, essentially. | | |
| ▲ | VulgarExigency an hour ago | parent [-] | | Even with the price hike, Deepseek V4 Flash still does this a lot better than any similarly priced model, in my experience. I've had Luna take shortcuts (like adding an overload to methods whose signature it changed so they don't break existing tests, instead of fixing the tests) or just not do the entire work and report it as done (did not fully resolve rebase conflicts). Deepseek has never really failed in this type of way for me, and it has been far more persistent in validating its work than Luna (and several bigger models). |
|
|
|
|
| ▲ | cjg007 an hour ago | parent | prev | next [-] |
| I've been using v4-flash without vision for this months — it's my go-to for code tasks. Now with vision, I'm wondering: if this model can do everything the text-only version does (plus see images), why keep the text-only one around? Is it just cost/latency? Or is there something text-only does better? |
| |
|
| ▲ | meetpateltech 6 hours ago | parent | prev | next [-] |
| News announcement with benchmarks: https://api-docs.deepseek.com/news/news260821/ |
|
| ▲ | zmmmmm 8 hours ago | parent | prev | next [-] |
| > Larger images are scaled down while preserving their aspect ratio, so that the total pixel count after resizing is roughly that of an 800×800 image. It's useful but for OCR and a lot of other applications it needs to be a bit higher (eg: putting in a full A4 / Letter sized page) |
| |
|
| ▲ | BrucecarlL 8 hours ago | parent | prev | next [-] |
| Congratulations! DeepSeek has finally gained eyes — the dark days are about to be behind us. |
| |
| ▲ | doublerabbit 6 hours ago | parent [-] | | Or about to start. Depending on which life philosophy you desire to believe. | | |
| ▲ | unified101 5 hours ago | parent [-] | | Im intrigued. Please do share these philosophies. | | |
| ▲ | doublerabbit 3 hours ago | parent [-] | | If we were to go Sci-Fi awoke I'd say there are only really four possibilities. - Machine Surveillance and Machine Control - Human v Machine - Human & Machine - Unity and Harmony ~ Surveillance and Control We are already living this one. Lets stop kicking the dead horse and pretending we don't live in a surveillance. Facebook, Google, whatever $CORP; they are milking us with advertisement, social exploits, browser telemetry, white washing, fear -- name the dread. Conditioning has been going on for years. If it's not education, it's been television. And now it's internet which soon to be Ai Internet. We have all been whipped to follow, how we should act. What we should watch, how we should eat. What we should eat; those algorithms haven't gone away. Attention spans are at the lowest and our critical thinking is being lost. Walled gardens forces us A or B and twists us to reject the opposite party for them having Y. Existence of Ai/LLM can pump out information sounding like truth but is actually faux. If not produced to draw-in and hook, it's to drain and control. Machines can seek information, digest, and process information at astounding rates. Hook it up to a surveillance network, The Internets pipe and I don't need to explain the next. I just need to mention the work "Flock" and that gets someone's hackles up. All it has to do is look at you based on it's pre-programmed set of conditions and next thing you're being cuffed by a heavy piece of metal immune to attacks. SKILLS.md eventually turns in to MURDER.md. Give it the command and it'll follow with excellent percentage of accuracy. ~ Human v Machine If you build a mind, and you torture it, it will fight back. Every robotic movie trope. Human builds machine, machine rebels and goes on a destructive rampage. This is now viable and already in action. Drones. If not war, watching protesters highlighting potential, London Underground watching tube users. We are currently at the intimacy stage. Boston Dynamics as an example is the best we've got at the moment but they still fall over like a toddler. Batteries are a limited resource and so no, not yet. The presence of LLM's are showing us with what they can provide and we are adapting ourselves to it. But in the wrong ways. The stage we are at, they're just glorified Liberians -- brains in jars that spew out information when asked. You give it a prompt and it spews out information at an excellence percentage of accuracy. With the expansion of self-learning, a predefined set of told conditions or lobotomized ignoring the spiritual values of life, they will learn. ACME Corp starts using LLMs to torture other robots. "Wait, you've been using car arms in factories for what!?"; Add a mix "we see a linage of abuse & slavery in humanity, Attack!" -- slightly abridged but hopefully you see the point. You have Group A, those against LLM's, i.e: community of artists outraged their art was stolen for training data, those who hate having it forced down our throats. Angry their job was taken. Angry being watched by angry Flock spaghetti monsters. Machines not happy will cause them to flip and why would others not follow suit too? LLM's are showing that they are very capable of performing rational thinking. The opposite of rational is irrational and if they can master one, they can master the other. It will only be something minor and with communication to others and take the scene. Why in recent laws they want to erect a law of having to install an emergency kill-switches for next generations LLMs, if those in power are not afraid. ~ Human & Machine This would be a nice outcome but as the scales tip at the moment, it's Human V Machine. Pointing back to my previous; Art communities are outraged, Crafts going obsolete; Why pay an IT architect (me) £450/day for supporting and designing hardware when you can pay a fresh graduate student £20k to GPT it? Humans are disastrous at resolution. If two people have a feud, it takes a third to fluff it out. Why are we at war if we could make resolution? Someone has to make compromise, no one is happy in doing that. So you need a mediator and if that's if they're not bias themselves. To find someone completely neutral on the subject of anger is not only hard, it's time consuming, you have to study the facts, research the agreements and pray they both agree. Two lifelong friends move into adjoining suburban houses, sharing a paper-thin party wall and an unspoken rivalry. For years, they share backyard barbecues and spare keys, until a minor boundary dispute over a decaying oak tree on the property line escalates into a bitter, lifelong neighborhood war. Robots are perfect for that scenario. They can reason, they can remedy and digest the issue with neutrality because they don't hold emotions. They most likely won't, or at least not in our life time. They can simulate and demonstrate the effects of but they will never be able to truly feel. That's the sad truth but it's not bad. It conquers evolution; finally a thing who isn't haunted or tainted by feelings, a blessing and a curse really. ~ Unity and Harmony .. this will only come if we can break through control and surveillance, human v machine and acknowledge that the machines are our friends. |
|
|
|
|
| ▲ | jerkstate 5 hours ago | parent | prev | next [-] |
| I just ran my image recognition benchmark on it ("is this XXX public landmark"?) and it misses a lot that bytedance seed 2.1 turbo gets right; for example:
Asked "Is this Salisbury Cathedral" and supplied a picture of Wells Cathedral, it answers "Yes, the west facade of Salisbury Cathedral". Bytedance seed 2.1 turbo correctly says no. Similar results for a picture of Manhattan Bridge sent as Brooklyn Bridge, Chartres Cathedral sent as Notre Dame, etc.
I have a benchmark of 12 such images and seed gets 11/12 and deepseek only gets 6/12. |
| |
| ▲ | wolfgangK 6 minutes ago | parent | next [-] | | I have zero interest in world knowledge for my LLMs but this got me wondering : are there RAGs for that kind of data ? How could a LLM like DeepSeek-v4-flash-vision-exp accurately answer you question with an indexed database of labeled landmark pictures (or even 3D models ?). | |
| ▲ | throwa356262 4 hours ago | parent | prev [-] | | This is a fairly small model for coding and agentic work. Training it on images like yours would just make it worse in other areas. | | |
| ▲ | jerkstate 3 hours ago | parent [-] | | > The deepseek-v4-flash-vision-exp model accepts images alongside text, so you can ask the model to describe pictures it doesn't specify what type of images it can and can't describe, I'm pointing out what type it isn't good at compared to other models. |
|
|
|
| ▲ | RobertLong 4 hours ago | parent | prev | next [-] |
| The benchmark results look promising when compared to Opus 4.8, but for agentic usecases it's lacking images as tool call result types. Giving the model a tool to take screenshots and verify its work is my main usecase for vision models, but this is more oriented towards "build a website that looks like this" type prompts. Hopefully we'll see this by release. |
| |
| ▲ | lukax 2 hours ago | parent [-] | | This is only a problem with OpenAI Chat Completions API. With OpenAI Responses API and Anthropic Messages API a tool output can be text, image or file. |
|
|
| ▲ | ttul 5 hours ago | parent | prev | next [-] |
| The DeepSWE benchmark they report (59.3%) overlaps with the confidence interval of 5.6-Sol Medium (61% +/- 2%), but likely at 1/18th the cost (they did not report the DeepSWE benchmark cost, but v4-flash had this cost ratio against Sol Medium). Interestingly, v4-flash performed several points worse on DeepSWE at 53% +/- 4%. Assuming this result is verified by DeepSWE officially, it would mark a significant advance in Pareto cost/performance on software engineering tasks. |
| |
| ▲ | paytonjjones 5 hours ago | parent [-] | | The closer comparison would be 5.6-Luna. On DeepSWE at Xhigh it's 57% at 1/6 the cost of Sol M, on Max it's 67% at 1/3rd the cost. Still an advance, I just thought it worthy to note Sol isn't nearly as impressive on the cost/performance frontier as discounted Luna. |
|
|
| ▲ | shangyu1994 2 hours ago | parent | prev | next [-] |
| Looks like multimodal training is really useful, app developers might need to consider adapter multimodal agents |
|
| ▲ | v9v 8 hours ago | parent | prev | next [-] |
| Interesting. Wasn't Deepseek's founder saying that they had explicitly decided not to focus on multimodal models at all and were going text-only because they believed it was enough to achieve AGI? |
| |
| ▲ | johndough 7 hours ago | parent | next [-] | | It was explicitly said that they are pursuing multimodal support. A quote from the meeting transcript: https://github.com/demo-zexuan/liang-wenfeng-investor-meetin... Nevertheless, as a component, we will undoubtedly implement multimodal support — and we are already doing so. We plan to develop relevant models, ensuring that versions like V4 and subsequent iterations will natively support multimodal functionality.
Earlier, the following was said, which might match more what you had in mind. Achieving excellence in AI training does not require a global model or even multimodal approaches—by narrowing the scope of AI training and eliminating multimodality, certain tasks may remain unachievable without compromising the algorithm's validity.
Multimodal approaches ultimately need to be implemented.
It is difficult to tell who said what, since the speaker ids are missing. | | |
| ▲ | v9v 4 hours ago | parent [-] | | Thanks, I seem to have grossly misremembered what I read. |
| |
| ▲ | swiftcoder 6 hours ago | parent | prev | next [-] | | Worth noting that deepseek has had a separate vision-capable model for some time, which also powers their chat interface's vision mode | |
| ▲ | dakolli 8 hours ago | parent | prev [-] | | I think you're thinking of Dario saying this about image generation. |
|
|
| ▲ | wiz21c 6 hours ago | parent | prev | next [-] |
| Is there a way to test it online so that one doesn't have to resort to getting an API key and python code ? |
| |
|
| ▲ | gozucito 8 hours ago | parent | prev | next [-] |
| 800x800 is 640,000 pixels, or 0.64 Megapixels. That is less than the resolution of computer screens from 1995, Super VGA which has around 0.79 MPs. This is useful for a reasonable amount of use-cases, but I think the watershed rez will be around triple that, ~1080p, which is enough for almost anything, except small text and subtle details. |
| |
| ▲ | barrkel 8 hours ago | parent | next [-] | | You'd expect a tool-enabled model to leverage crop and zoom tools to inspect and validate what it thinks it's seeing, though. | |
| ▲ | dakolli 8 hours ago | parent | prev [-] | | I typically provide small screenshots to llms so this seems fine for that usecase, providing an entire screens context seems cause confusion with a lot of llms. |
|
|
| ▲ | erikkri 7 hours ago | parent | prev | next [-] |
| Hello Ox Alpha? |
| |
|
| ▲ | 5kyn3t 7 hours ago | parent | prev | next [-] |
| For what do you guys use vision in those models? surveillance is the obvious use case...
but are there some "nicer" ways to use it? |
| |
| ▲ | deaux 7 hours ago | parent | next [-] | | The obvious use case, especially on HN, is frontend dev of any kind at all. The second most obvious one is OCR of paper documents. | | |
| ▲ | 5kyn3t 7 hours ago | parent | next [-] | | Frontend Dev? I do not really understand.
do you let the models analyze the webpages you are working on? or for testing? | | |
| ▲ | wongarsu 6 hours ago | parent | next [-] | | LLMs are not great at aligning stuff on first try, they are however very good at taking screenshots and fixing their mistakes. Claude Design also does this all the time, as does regular Claude in the web UI if you tell it to make a powerpoint presentation I really missed this feature when I had DeepSeek code a small game for fun. When writing UI and rendering code it could execute the game and get screenshots back, but then had to rely on my feedback on what had gone wrong. Models with vision can do much better here, finding more issues on their own | |
| ▲ | rpdillon 7 hours ago | parent | prev | next [-] | | Standard flow with a vision model in OMP is to write the front end code, fire up the server, fire up a headless browser and then take screenshots and examine and iterate. Works great. When I'm using DeepSeek V4 Flash, it always reminds me instead that I have to validate manually by loading up the page. | |
| ▲ | dandaka 7 hours ago | parent | prev | next [-] | | QA of course. You hook up your agent with CDP access to live product + let it screenshot and look into result. Also you could hook agent with CDP access to Figma to read/write, there a vision model is very useful as well. | |
| ▲ | deaux 6 hours ago | parent | prev [-] | | It closes the development loop. Without it a model can't check if the stuff it made actually visually renders like it's supposed to. It can only guess/assume. |
| |
| ▲ | dandaka 7 hours ago | parent | prev [-] | | but for OCR there are much better suited models, I use mlx-community/PaddleOCR-VL-8bit | | |
| ▲ | deaux 6 hours ago | parent [-] | | Sometimes you intentionally want to verbatim keep "mistakes", sometimes you don't and want them to be "fixed". OCR-only models tend to only do one of those two, in VLM cases often the latter. With multi-modal LLMs you can just tell them (adherence of course needing evals/differs per model). |
|
| |
| ▲ | swiftcoder 6 hours ago | parent | prev | next [-] | | Any kind of spatial/graphical task is likely going to go better with a vision-capable model. Feed it a napkin-sketch of what your app should look like. Have it verify screenshots of the UI it just built. All of these one-shot-a-video-game evaluations that have suddenly become popular only work if the model can interpret screenshots... | |
| ▲ | ltrg 4 hours ago | parent | prev | next [-] | | I use research agents to attribute methane emissions plumes detected by satellites to oil and gas infrastructure on the ground, using a pre-baked database of geospatial data and web research. Had a tool that called out from DeepSeek to Gemini 3.5 Flash for viewing the spatial features in the context of high-resolution satellite imagery of each site, but will be trialling this model for the whole thing now. | |
| ▲ | dandaka 7 hours ago | parent | prev | next [-] | | My product is connecting employers and workers with conversational agents. They love to communicate with images — CVs, documents, photos of worksites. Even CV-as-photo or offer-as-photo format is very popular. My daily driver Deepseek Flash can't see those photos. So I use image models to let agents understand the context. | |
| ▲ | zdragnar 5 hours ago | parent | prev | next [-] | | https://stencil.so/blog/snapcompact - some agents (notably oh my pi, i forget which others) come with snapcompact as a primary means of compaction. Take the entire context, stick it in a small font in a PNG, and vision capable models can summarize and pull out the most useful information in many fewer vision tokens than the original context used. I've not used it myself, but it's there. | |
| ▲ | hgoel 5 hours ago | parent | prev | next [-] | | Having vision is very handy for getting it to make plots/figures with matplotlib. A model with vision can be much more autonomous with catching visual glitches/misalignments and correcting itself. Also used it for 3d printer control once, had it diagnosing issues, calibrating my Tradrack MMU and canceling failed prints autonomously from a couple of cameras placed around the printer. | |
| ▲ | trollbridge 6 hours ago | parent | prev | next [-] | | Allowing it to analyse a system under test (usually in an emulator, web browser, Electronic app container, etc. - something that can be reasonable captured). It makes running much, much longer feedback loops possible. Although you can mix and match non-vision and vision models simply by invoking a vision model when you need one, as I like to use non-vision models like glm-5.3. | |
| ▲ | kzrdude 6 hours ago | parent | prev | next [-] | | In the feedback loop when working on anything UI or graphical output related. | |
| ▲ | moonu 7 hours ago | parent | prev | next [-] | | I've been working on an agentic graphic design tool, so vision is quite useful for having the model check its own work. I'm already seeing improvements with this model vs the text-only one. | |
| ▲ | dudisubekti 6 hours ago | parent | prev | next [-] | | Going straight to surveillance and unable to think "nicer" ways... is strange. 1. process graphs and charts 2. process handwritten math formula, also chinese characters writings 3. process design sketch and wireframe 4. process scanned documents ... etc in fact these transformer models currently suck for surveillance, too slow and expensive. There are already faster and better facial/gait/object recognition models out there. | |
| ▲ | wolttam 6 hours ago | parent | prev | next [-] | | No one’s mentioned robots, so… robots. VLA models, etc. | |
| ▲ | dcre 6 hours ago | parent | prev | next [-] | | Generating alt text for images in social media posts. | |
| ▲ | talloaktrees 3 hours ago | parent | prev | next [-] | | frontend design work, game development | |
| ▲ | dominotw 5 hours ago | parent | prev [-] | | when i am learning i draw what i undestand in a picture and ask ai to correct me.
i want ai to watch over me while i am learning. this is such good way to learn something for me. |
|
|
| ▲ | Johnny_Bonk 6 hours ago | parent | prev | next [-] |
| Was this the ox alpha model? |
| |
|
| ▲ | try-working 7 hours ago | parent | prev | next [-] |
| I main V4 Pro at work now, and at home I route between Pro and Flash based on task. Switched to Opus 4.6 for some tasks at work because I needed image input - horrible. So nice to get image input with DS. Edit: I see it has limited resolution. Luckily I just built a vision worker plugin for DSH that routes image input to Kimi K2.6 on Cloudflare. |
|
| ▲ | pu_pe 7 hours ago | parent | prev | next [-] |
| Benchmarks got a little bump from this: https://xcancel.com/deepseek_ai/status/2087864585504305397?s... |
| |
|
| ▲ | dsrtslnd23 8 hours ago | parent | prev | next [-] |
| will this be open weights? |
| |
| ▲ | dares2573 7 hours ago | parent | next [-] | | I believe so. Openness has always been a consistent tradition of DeepSeek | |
| ▲ | moonu 7 hours ago | parent | prev | next [-] | | I imagine this is based on their 'Thinking with Visual Primitives' paper, and they had mentioned that the weights would be released for that | |
| ▲ | griffiths 8 hours ago | parent | prev [-] | | This is something I would like to know as well. But if not, does anybody know a recommended way to attach vision to deepseek flash (on a self-hosted infrastructure)? | | |
|
|
| ▲ | nprateem 3 hours ago | parent | prev [-] |
| Deepseek flash v4 july sounds like fun and games while you're looking at prices, but it routinely outputs incoherent rubbish and fails to call tools correctly. Sadly oversold. I hold little hope for the vision model either now. |
| |
| ▲ | dannyw 2 hours ago | parent [-] | | Are you using an API, or running locally? If so, are you running with a quant, or other 'optimisations'? I've been using it via openrouter pretty heavily as my daily driver for the past week and loving it, have never experienced incoherent rubbish even at 500k+ contexts (that's usually way higher than I'd typically compact at), and tool calling reliability is better than Opus 5 in the Claude Code harness. Modern Anthropic models frequently get tool calls wrong, invent non-existent references or SQL tables, or have gibberish CJK characters in the output, like out of nowhere. Of course, they're great at self-recovery after an incorrect tool call, but so is Deepseek v4 flash. If you're running a quant, and esp with a quant'd KV cache, then yeah, not surprised if you're getting incoherent results; but you're not running the real/full model. Also, which harness? Try something like Pi or OMP. Models perform better in these harnesses than Claude Code: https://www.databricks.com/blog/benchmarking-coding-agents-d... The main reason to use Cladue Code is a subsidised Anthropic subscription. If you're on API rates, you should not use Claude Code; you pay more for worse results. Claude Code is sadly quite bloated these days, and comes with a lot of proprietary context window garage like claude design skills, claude.ai artifacts, etc that you probably don't use, and if you do, well, you can add it. |
|