Remix.run Logo
▲ wewewedxfgdf 2 hours ago

Within one question of their web interface, it has lost context and asks you to clarify what you are talking about.

I am very often giving the same programming task to multiple LLMs for various reasons - the answers from Google are so bad that I gave up.

I have no interest in benchmarks.

▲mattlondon 2 hours ago | parent | next [-]

So you have no experience of their latest model release then? Just repeating the usual tropes about Google having messed up? Or basing your opinions on their website chatbot?

If you have actual independent benchmarks and evidence about how this new model release is "so far behind" and refutes the stuff from their blog then please do share because I think we'd all love to see that?

▲wewewedxfgdf 2 hours ago | parent [-]

No I am commenting on my real world experience of using Gemini daily. I still ask it questions alongside Claude and OpenAI and Gemini is always the worst of the three.

▲mattlondon 2 hours ago | parent [-]

So you've not used this new release then? So how can you say that they are "so far behind" if you are not using the most recent model for your comparison. This is their first 4.0 model, that you are not using and instead basing all your opinions on on some ancient months-old model from a previous generation?

With respect, I don't find your arguement about them being "so far behind" especially convincing when you are using previous-gen releases and not actually using their current release.

▲fwip 2 hours ago | parent | prev | next [-]

Yeah, they've definitely got some recurring tooling/infrastructure problems around the models.

▲ 2 hours ago | parent | prev [-]
[deleted]