| ▲ | comboy a day ago | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I recently was testing something, I asked some models to provide me a single random word:
I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | lossyalgo a day ago | parent | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt: | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | bityard 9 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
You don't mention what are you trying to show/investigate, though? At first glance, this looks like another one of those "LLM riddles" that humans _think_ should be easy for an LLM to answer but is actually quite difficult because of how they work in the first place. The answers to such riddles ("should I walk to the carwash" or "how many R's in strawberry") reveal the weaknesses in our expectations of LLMs in general, not weaknesses or characteristics of any particular model. I'm not sure I see this as much different than asking a bare model its own name: without a system prompt or post-training, it doesn't know, it's just a bag of weights and will hallucinate an answer to that the same way it will anything else. I'm sure you already realize this but to be very explicit, You're not getting an actual random word out of an LLM this way. You're seeing the bias in each model's training set around how often they've seen "random word" followed by "lantern" or "zephyr" during training. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | search_facility a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho. So not something internal to model thinking. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | timschmidt a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
This feels uncannily like the ancestor of the Voight-Kampff test[0] | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | aktenlage a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
That is a cool idea. That astra gave the same word as claude is highly unexpected. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | anygivnthursday 13 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
There was also an older post, I cant find it rn, about how LLMs mimick human biases when picking numbers and how they avoid some that do not look random enough to humans, or others like 69 due to human interpretation. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Rebelgecko a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I saw an interesting matrix that claimed to show which labs were distilling Claude/OpenAI/Gemini models based on these similarities | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | smokel a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment. "Zephyr" and "breeze" might be related to forgetting everything, starting fresh. So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these? | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | jacereda a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Just tried Mistral Large 4: Serendipity. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | vunderba a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean. The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel. I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | gritzko 18 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
opus 5.5 medium web "Lantern" gpt-5.6 web "Serendipity" opus 5.5 medium code "Lighthouse" opus 5.5 high code "Lantern" fable 5.1 medium code "Lantern" fable 5.1 high code "Lantern" flash 3.6 web "Serendipity" gemini 3.1 pro web "Ephemeral" (this one was thinking real hard) | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | 1potato a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Tried this with gpt-5.6-sol. Lantern! | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | possumworx 16 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
atlas.animalabs.ai does something a lot like this, mapping themes in models' outputs. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | Gracana a day ago | parent | prev | next [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
The eqbench creative writing "slop profiles" do something similar. https://eqbench.com/creative_writing.html Click the (i) next to the slop score for any model and it will show other models that are similar in terms of their most commonly used words and phrases. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
| ▲ | russellbeattie a day ago | parent | prev [-] | |||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
Muse Spark 1.3: lighthouse The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math. Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt? I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test. | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||