Remix.run Logo
andrewstuart 2 hours ago

This is equally bad as a pelican test.

LLMs should be tested in the same way people should be tested for a job interview (but often aren’t) - with tasks RELEVANT to usage.

So you don’t just randomly pick some random thing to make the LLM randomly do (like many job interviewers do).

You start with clear statements about real world usage scenarios. THEN you come up with tests that give insight to how well the LLM/hob seeker gets the job done.

Please, stop coming up with random tests like it’s Microsoft in 1990 and you’re asking job seekers how the would move Mount Fuji, as a way of assessing their programming skills.

No stupid irrelevant pelicans on bicycles and no stupid renderings of Lord Of The Rings. Unless those are relevant use cases.

Any test that anyone comes up with must clearly state the context and how the outcome is measured.

hansmayer 27 minutes ago | parent [-]

[dead]