The study tested OpenAI’s GPT-4, Meta’s Llama, Google’s Gemini, and Mistral’s Vibe across various tasks, including job applications and resignation letters. When prompts included hedging phrases like "maybe," collective pronouns such as "we," or expressive adjectives like "lovely," the models abandoned their standard, crisp corporate output. Instead, they generated responses that were structurally more complex and emotionally charged, a phenomenon that persisted even when the messages were signed with traditionally male names.
Katherine Van Koevering, the study’s lead author, noted that the disparity was startling enough to question the reliability of AI-generated workplace correspondence. The findings suggest that the models are not merely mimicking the user's tone but are actively reinforcing gendered linguistic tropes. As users increasingly rely on spoken AI interactions where unconscious habits are harder to edit, the researchers argue the burden of correction must shift. They contend that technology companies, rather than end-users, are responsible for auditing and fixing these ingrained biases within the systems themselves.

Comments (0)
No comments yet. Be the first!