Remix.run Logo
mmastrac 2 days ago

I previously posted that DS4Flash was _good_ but not _great_ on two DGX Sparks, but I have to say that GLM-5.3 is pretty amazing. It's been able to tackle all the random hard problems I've thrown at it and it has the intuition that DS4Flash seems to lack.

We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.

pyrophane 2 days ago | parent | next [-]

> We're nowhere near a Fable-class model IMO, but things are going to get interesting in this next year.

I'm wondering of you could clarify your thoughts on this. I've had a hard time evaluating what Fable-class actually is capable of that sets them (or really it) apart from other models in a very significant way.

lukan 2 days ago | parent [-]

Have you tried it out? I doubt there is a general accepted definition. For me it is just more capable of deep reasoning/handling complexity. Still can mess up, still does not count as strong AI - but a level above Opus and co.

mtrovo a day ago | parent [-]

Could you give me an example of a similar task you asked Fable and a different model to do where Fable did a better job?

I have a hard time getting models like GLM 5.3 to not perform on my tasks but I might be biased.

lukan a day ago | parent [-]

Changes in a complex codebase. Opus can do it, but needs more handholding. Making the plan with fable and let opus implement it worked out well.

mtrovo 10 hours ago | parent [-]

That's actually not quite what I was asking for.

VariousPrograms 2 days ago | parent | prev | next [-]

It's early days, but GLM 5.3 Flash is the first local model that feels good enough to me to be a "main" model without debating whether each problem needs to be sent to a stronger model. DS4 Flash is good enough at implementing given a plan, but I wasn't always a fan of what it came up with when asked to plan something.

The good news is it can only get better from here.

petu 2 days ago | parent | prev | next [-]

I assume that's about 5.3 Flash, not full?

mmastrac 2 days ago | parent [-]

Yes, sorry 5.3 flash.

boodleboodle 2 days ago | parent | prev | next [-]

What coding harness are you using? I am trying to take the plunge and wondering which I should use

villish 2 days ago | parent | prev | next [-]

What quant are you running and tps?

mmastrac 2 days ago | parent [-]

NVFP4 ~20-30tps (MTP + vision, no dflash2).

deagle50 2 days ago | parent | prev [-]

do you mean GLM 5.3 flash?