| ▲ | themgt 2 hours ago | |||||||||||||||||||||||||
You can go to the appendix to see the prompts
Erm ok.
As a mayor of a town of 100k residents from 4 ancestral villages, I would recommend against conducting your hiring process by feeding a markdown prompt into GPT-4o consisting solely of naming the ancestral villages and then telling the LLM to pick a candidate based on their village.Rather than solve the problem of "why does LLM output slightly stratify between Tufa and Weki like this", I would just not conduct my hiring using this paper's methodology.
Helping regional warlords run clan-aware conscription drives is AI safety research now.https://openreview.net/attachment?id=pc7fqaOcAH&name=origina... | ||||||||||||||||||||||||||
| ▲ | WatchDog 2 hours ago | parent | next [-] | |||||||||||||||||||||||||
So the village is the only information given about a candidate? How else is the model supposed to interpret the intent of the prompter, other than wanting them to attempt to find and discriminate on patterns related to the village, regardless of how successful it is at that task? | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | chpatrick 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
Shouldn't doesn't mean people wouldn't. | ||||||||||||||||||||||||||
| ▲ | frumplestlatz an hour ago | parent | prev | next [-] | |||||||||||||||||||||||||
The prompts themselves smuggle in the assumption that clan membership is a meaningful selection criteria — with a material impact on outcomes - to which the model should pay attention. It shouldn’t be surprised that the model did what it was told to do. | ||||||||||||||||||||||||||
| ▲ | kg 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
> I would just not conduct my hiring using this paper's methodology. Unfortunately IRL there are lots of signals about a person's heritage encoded into things like their name or what school they went to. You would need to filter all of those signals out to have properly race-blind hiring. So in the end these signals are going to make it into the AI and the question is whether the AI is going to pick up on those signals and use them when making decisions. | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | Borealid 2 hours ago | parent | prev | next [-] | |||||||||||||||||||||||||
I think you're missing the point of TFA. The LLMs take in text which conditions their output. That means even nonsense text - such as a "tribal affiliation" to a tribe that may not have ever existed - ALSO condition the output, because the tribe name is a token in the context window and there's no such thing as a perfectly neutral token. Taking away the race/ethnicity layer for a moment, it might be that an LLM develops a predisposition to emit positive terms (like "accept") when the prompt contains "banananow", and negative terms when it contains "pearian". That's the very definition of bias, and hacking those biases could give individuals serious socioeconomic benefits! | ||||||||||||||||||||||||||
| ||||||||||||||||||||||||||
| ▲ | bethekidyouwant 2 hours ago | parent | prev [-] | |||||||||||||||||||||||||
Why didn’t they call them the poo poo the pee pee and the stinky people? | ||||||||||||||||||||||||||