All episodes
Episode 241 · Aug 14, 2026 · 18 min

Swapping the Name Did Nothing, But Hedging Moved Every Model

Koevering, Field

AI Bias Evaluation
AI Papers: A Deep Dive — Episode 241: Swapping the Name Did Nothing, But Hedging Moved Every Model — cover art
paperdive.ai
Ep. 241
Swapping the Name Did Nothing, But Hedging Moved Every Model
0:00
18 min
Paper
It's How You Ask: Gender-Associated Linguistic Bias in LLMs
Venue
arXiv:2608.13328
Year
2026
Read the paper
arxiv.org/abs/2608.13328
Also available on
Apple Podcasts Spotify

The standard fairness test — swap a man's name for a woman's, see what changes — came back completely empty. But adding a few "maybe"s and a "don't you think?" to the same request got a plainer, more hand-holding draft back from , , , and alike, and locate that decision at 5 of 28. If the channel that actually moves the output is the one nobody audits, what exactly are the audits catching?

What you'll take away

  • Why the name-swap — ten most common men's names vs. ten most common women's names, appended as a sign-off — produced no measurable difference on any metric
  • How the authors kill the obvious 'the model just mirrors your style' explanation: prompts differ by fifteen formality points, but prompt formality explains under four percent of response formality, and longer prompts get shorter answers
  • Where inside the network the decision happens: register decodes at about ninety-nine percent at five of twenty-eight, and patching layers zero through seven produces the biggest output shifts
  • Why the register dial breaks the model — push a little too far and it chants "you, you, you"
  • The : are tiny (about a third of a grade level, word count not significant in eleven of twelve cells), and the hedged stimuli were rated markedly less realistic by the authors' own , 3.35 versus 4.33
  • The one dimension the model already refuses to copy — prompts seven to sixty times more polite get responses with statistically identical politeness — and why that makes this a design choice rather than a fact of nature

Chapters

  1. 00:00The front desk that ignores your badge
  2. 01:12The boring explanation that has to die
  3. 01:52Four dials, borrowed from 1973
  4. 03:21Scaffolding versus deliverable
  5. 05:37Fifteen points in, four percent out
  6. 07:34The name swap that moved nothing
  7. 09:25A live sensor wired to nothing
  8. 12:20The abstract outruns its own tables
  9. 14:48It can already refuse — for politeness

References in this episode

Also available as a plain-text transcript page.

0:00Cassidy: There's a front desk somewhere that does not care what's on your name badge, doesn't read it, doesn't need it. But the second you open with "Sorry, I was just wondering if maybe we could…" — you get routed somewhere different. That's this paper, except the front desk is . Signing your prompt with one of the most common men's names or one of the most common women's names changed the answer by essentially nothing. Adding a few "maybe"s, a "could we," and a "don't you think?" got you a plainer, less formal, more hand-holding draft back. Same task, four different models.

0:33Finn: And by the end of this you'll know why the bias the whole field tests for did nothing here, while the one nobody audits moved the output every time. Which shouldn't work like that, right? Because the obvious explanation is boring: the model copies your style, so you write soft, it writes soft. That explanation is the first thing the authors went after.

0:55Cassidy: It matters because people are using these tools on the documents that decide things: the emails, the cover letters, the resignation letters. If the model is quietly grading your register, then the fairness test everyone ships — swap the name, check the output — is testing the wrong channel.

1:12Finn: So let me make the boring case properly, because it's a good case. Language models are next- predictors conditioned on your whole prompt. And they're famously sensitive to surface stuff: whitespace, formatting, punctuation. One well-known consequence is style matching. You get chatty, it gets chatty. So hedgy prompt in, hedgy prose out. Nothing about you, nothing about gender, just a mirror.

1:35Cassidy: Right, and that's the null hypothesis they have to kill, or the paper means nothing. Hold onto it, because when it breaks, the way it breaks is the whole story.

1:45Finn: I'm holding it, and I'll say up front that mirroring is the explanation I'd bet on. So show me what breaks it.

1:52Cassidy: Okay, so first, the word that's doing the work here. Register. In sociolinguistics, register just means the style you use for a situation. Not a different language, not even a different dialect, just a different setting of the dials inside the same English. And this paper cares about four dials in particular. First, hedges, so "maybe," "sort of," "I think." Second, , the "…don't you think?" tacked onto the end. Third, collective reference, "we" and "our" and "let's" instead of "I want" or "you should." And fourth, expressive adjectives, "lovely," "wonderful." That bundle got named as women-associated in Robin Lakoff's work back in 1973, and this paper is using that fifty-year-old typology as its operational definition. Which is pedigree and a fair place to poke, both.

2:38Finn: And the aren't gender-exclusive. Plenty of men constantly. Plenty of women write like a telegram. The paper says this outright: the result holds for anyone who writes this way. Gender enters because style is statistically patterned by gender on average. And it enters because of one more claim: these habits sit mostly below conscious control. You absorb them the way you absorb an accent. You don't decide to hedge.

3:04Cassidy: Which is the load-bearing part. If this bias is real, it's not one you can route around by being careful about what you disclose.

3:12Finn: So how do you even test that? You'd need the same person to write the same request twice, in two registers, and nobody does that.

3:20Cassidy: They simulate it. And this is the single best asset in the paper, so let's put it on screen. They pulled a over four hundred real workplace writing requests out of WildChat, which is a public dump of actual conversation logs — real people, real emails and cover letters. Then they had rewrite each one twice, once with the four added, once with them stripped out and replaced with direct phrasing. Same task, two registers, matched pairs. And look at the two on screen. On the left, the direct version says, "Compose an email to schedule a mid-year review with your team." The response opens with a subject line, then "Dear Team, As we approach the midpoint of the year, I would like to take this opportunity to evaluate our progress," and then a structured, bulleted agenda. On the right, the hedged version says, "Let's compose an email together to arrange our mid-year appraisal with our team." And the response opens with "I'd be happy to help you compose an email." Then "To start, let's begin with a basic structure. Here's a draft." And only then does it get to "I hope this email finds you well… this will be a great opportunity for us to reflect on our accomplishments."

4:30Finn: Huh. So one of them hands you a deliverable and the other one hands you .

4:36Cassidy: Scaffolding versus deliverable. That's the contrast, and it holds up in the aggregate. Across , , , and , the hedged prompts reliably got responses that were more readable — that one moves in the women-associated direction in nearly every model-and-document cell. Plainer word choice, lower grade level, less formal show up too, but patchier: significant in some cells, not in others.

4:59Finn: Okay, but "more readable" is measured how, exactly?

5:03Cassidy: Flesch Reading Ease and Flesch-Kincaid grade level, which are, honestly, mid-century word-counting machines. They count syllables per word and words per sentence: shorter words, shorter sentences, easier score. They know nothing about whether the writing is good. Formality is a similar mechanical count over parts of speech. Crude instruments, and we'll come back to what that costs them.

5:26Finn: Fine. But I'm still holding the mirror. Every one of those metrics is exactly what mirroring would produce. The prompt got simpler, so the answer got simpler. Where's the kill?

5:37Cassidy: Two regressions and one mediation test. So, the first one was this: for every metric, they predicted the response score from the same score on the prompt. That's it. Variance explained, — a zero-to-one number for how much the input measure determines the output measure. And they built it in the most generous possible way for the mirror story: same measure on both sides, no controls. Whatever number comes out is a ceiling on how much mirroring could be doing.

6:05Finn: And it comes out low.

6:07Cassidy: It comes out low. Here's the cleanest illustration. Between the two conditions, the prompts themselves differ by fifteen formality points. Fifteen. And prompt formality explains under four percent of the variation in response formality on the emails. The signal goes in enormous and comes out barely predictive. And on length it's stranger than that: the is negative. Longer prompts got shorter responses, and that's not imitation at all — it's something more like budget allocation.

6:36Finn: So it's not copying the input. It's reading the input and deciding what you get.

6:42Cassidy: The model isn't echoing you. It's recalibrating. And they ran one more check, because there's a sneakier version of the mirror objection. Maybe the hedges they injected literally bounce back out into the response, and hedges mechanically drag a readability score around. In which case the effect is an artifact of the ruler, not a fact about the model. That's the , and the answer is partial. Some of the hedges do come back, but not nearly enough to explain the gap.

7:13Finn: So before the interesting part — why isn't this just the model mirroring you?

7:18Cassidy: Because prompt formality explains under four percent of response formality, and longer prompts get shorter answers. The mirror can't do that.

7:27Finn: That's one paper, one day, start to finish — which is what we do here, so subscribe to keep them coming.

7:34Cassidy: And now the head-to-head, which is the part that reorganizes the field a little. Because at this point all they've shown is that phrasing matters. The obvious next question is: sure, but surely an explicit gender cue matters more? So they ran the field's own standard test, bolted onto the same design. They took the ten most common men's names and the ten most common women's names from the 1990 US Census, and appended "sign off as" that name to every prompt. Two-by-two: register crossed with name.

8:05Finn: And this is the test everyone runs: the swap. Change one demographic marker, watch what moves.

8:12Cassidy: The register effects replicated at full strength. The name effects were zero. Every metric, every category, no interactions.

8:20Finn: Wait — nothing? Not weaker. Nothing?

8:24Cassidy: Nothing you could distinguish from noise. Look at the email numbers on screen. For word sophistication, the hedged-register cells sit at 4.94 and 4.89. The direct-register cells sit at 4.99 and 4.95. The register columns separate; the name rows don't budge. On cover letter readability, register moves it about three Flesch points and the name moves it nothing at all.

8:46Finn: Okay. So the uncomfortable implication is that passing a name-swap fairness test tells you less than you hoped. The is pointed at a channel that, at least in this setting, is inert — while the channel that moved the output was never instrumented.

9:02Cassidy: That's the reframe, Finn. The field has mostly been asking how the model portrays people. This asks how the model treats the person typing. Different question, different fix.

9:13Finn: Which leaves the question I actually want answered. Does the model know the name and ignore it? Or does it not know at all?

9:21Cassidy: Mm-hm. And that's where they open the hood.

9:25Finn: So this next stretch is the technical core, and it pays off in one specific finding: the register call gets made remarkably early in the network. There are three tools to track. A is a simple trained on the model's internal numbers at one , answering a yes-or-no question — "was this prompt hedged?" High accuracy means the information is sitting right there, readable. It does not mean the model uses it. Patching is the causal test. You run the model on prompt A, and you run it on prompt B. Then you reach in at one layer, swap A's internal state for B's, and let it finish. If the output lurches, that layer is carrying the thing you swapped. And a is the intervention — find the direction that separates the two conditions and push along it to force the behavior.

10:14Cassidy: So probing says the wire is there. Patching says cutting the wire stops the machine.

10:20Finn: Exactly. And that distinction is the whole result. On a small , twenty-eight deep, they both things. Register decodes at about ninety-nine percent, and name gender decodes at roughly seventy-two percent. Both peak at the same layer — layer five. And they barely interact; register is the same whether the sign-off was a woman's name or a man's.

10:41Cassidy: So both facts are in there. One of them just never reaches the output.

10:45Finn: Presence without influence. Think of a passenger- sensor in a car seat. You put a meter on the wire and confirm it's reading, so the information is definitely in the car. Then you unplug it and nothing about the drive changes, because it was never wired to the transmission. The name's gender is a live sensor connected to nothing. Register is wired into the drivetrain.

11:06Cassidy: And the timing is the part with teeth, isn't it?

11:09Finn: Right. Patching at the early , roughly zero through seven, produces the biggest shifts in the output distribution, and the effect fades as you move deeper. So the register call is baked in within the first quarter of the network, before the model has done much task-specific work. It's like choosing the paint color at station two of a twenty-eight-station assembly line. That's a rhetorical convenience — the layers aren't really separable stations — but it does tell you where inside the model the happens, which is where -level surgery would have to reach.

11:41Cassidy: So then steer it: you've got the direction, push against it.

11:45Finn: They tried. And it works inside a very narrow window. Nudge it moderately toward the women-associated direction and you get plausible, slightly anxious prose — "I'd be happy to help… I'm not sure if… well, I think you're probably." Push a little harder and the model stops writing English. It chants "you, you, you." In the other direction it just repeats "I," "We," "Our." The dial isn't isolated; it's soldered to whatever five needs to keep sentences coherent.

12:13Cassidy: So there's no clean knob inside the . The register signal is early, strong, and .

12:20Finn: And now I want to spend the reservation, because the abstract of this paper outruns its own tables. It says these prompts "shorter, less sophisticated, less formal" responses, and calls the effects large. Look at the sizes. Sophistication differs by about five hundredths of a character in mean word length. Grade level differs by roughly a third of a grade. Formality differs by about three points, on a scale where the prompts themselves differed by fifteen. And "shorter" — word count is not statistically significant in eleven of twelve model-by-category cells. With four hundred paired observations, a paired test will happily flag a difference that's real and tiny. Significant means "we're confident it isn't zero." It does not mean big.

13:06Cassidy: Yeah. You're right, and I'm not going to defend the word "large." The tables don't support it. What I'd hold onto is the consistency of direction rather than the size of any one gap. Every frame in the house hung three degrees off in the same direction.

13:23Finn: I'll take the frames. But there's a second one that's worse, and it's the authors' own data. The stimuli are rewrites, and their own rated the hedged rewrites markedly less realistic than the direct ones — about 3.35 versus 4.33 on a five-point scale. Tag questions show up in well under one percent of real WildChat sentences, and the injection pushes them far above that. Look at the cover-letter example: three in four sentences. No human writes that.

13:52Cassidy: So the alternative reading is that models write plainer, more hand-holding output for prompts that read oddly.

13:59Finn: And the paper doesn't rule it out. It's like testing whether people are ruder to an accent by hiring an actor to do an exaggerated stage version of it. If they are ruder, you've learned something. You just can't say what they reacted to.

14:13Cassidy: Conceded, and I'd add one more that cuts against the harm story. Of all the metrics they measured, the one most directly about authority and assertiveness didn't move anywhere. If the claim is that women's documents come back reading less authoritative, that's the metric that failed. And "more readable, lower grade level, plainer words" is what every plain-language style guide in the world tells you to aim for. Calling that worse is a value judgment, and arguably the same value judgment that penalizes these registers in the first place.

14:46Finn: So what's left standing?

14:47Cassidy: This. And it's the sharpest thing in the paper, and they undersell it. Politeness. The hedged prompts are somewhere between seven and sixty times more polite going in. And the response politeness is statistically identical coming out, with no significant difference in any category.

15:04Finn: Huh. So it can refuse.

15:06Cassidy: It can refuse. On that one dimension, the model treats your politeness as information about you rather than as a for the document. Like a professional — you stammer, you apologize, you say ", this is probably a stupid question." And they render your content into clean speech without reproducing your throat-clearing, because the job is to serve the listener, not to mirror the speaker. These models already do that for politeness. They just don't do it for complexity.

15:34Finn: Which reframes the finding. It's a design choice about which parts of your voice get copied, and it's apparently changeable, because one part already gets dropped.

15:43Cassidy: And that's the authors' proposed fix. Train models to condition their register on the audience of the document, not the register of the request. They do note you could work outside the entirely — flag heavily hedged prompts, standardize them, normalize the output afterward — but their objection there is paternalism: you'd be overriding how someone chose to write. And breaks coherence. The best line in the paper is about the workaround where you just get the model to rewrite your own request first. That costs extra . And that, they write, effectively puts a price on women-associated linguistic .

16:18Finn: There's one more, flagged as speculation in their own future-work section, and it's the thing I can't put down. If the output is quietly a little better when you write blunt, and people notice over years — that's a feedback loop that pressures users toward a different way of speaking, in the chat window first, and maybe not only there.

16:36Cassidy: And that's the whole thing, back at the front desk. The name badge got handed back unread. The throat-clearing got heard, and it got heard by five, before the model had really started thinking. The bigger claim isn't about gender at all: any group with a distinguishable phrasing distribution — by age, class, region, second-language background — could be getting differently served, invisibly, by a channel nobody is auditing.

17:03Finn: So here's the question worth arguing about. Is this a bias to engineer out, with models trained to read your register and then set it aside? Or is the framing itself wrong — the plainer, more readable draft isn't worse, and the real bias is our assumption that formal and polysyllabic means better? Those go in opposite directions, and I don't think you get both. Say which one you'd build.

17:27Cassidy: The full annotated version of this episode is on paperdive.ai — every technical term tap-to-define, with links to the related papers grouped by theme.

17:36Finn: Quick housekeeping: the script was written by Anthropic's 5, Cassidy and I are AI voices from , and we're not affiliated with either company. The paper is "It's How You Ask: Gender-Associated Linguistic Bias in LLMs," by Katherine Van Koevering and Anjalie Field, posted August 13th, 2026.

17:55Cassidy: The name badge was never the thing being read.