When Grok Graded Its Own Encyclopedia And Marked Itself Down
Concepts in this episode
Click a concept to find related episodes and external papers worth reading. See the full concept index.
About this episode
Elon Musk built Grokipedia to be less biased than Wikipedia — then researchers had four rival AIs grade it, and even Grok, the model that wrote every article, rated its own encyclopedia as the more biased one. But the self-conviction turns out to be the least interesting part: four judges who agree about almost nothing all tipped the same direction. We walk through how the study broke the circular trap of using biased AI to audit biased AI, what it actually found, and the crack running right through the whole thing.
What you'll take away
- How the study escaped the circular trap of using a biased AI to audit biased AI — a human-coded ideology ruler (V-Party) plus four judges chosen to lean different directions
- The blunt top-line: roughly 4 in 10 Grokipedia articles rated biased vs about 3 in 10 for Wikipedia, across 1,394 article pairs
- Neither encyclopedia is a hit piece — both flatter their own team, Grokipedia warming to free-market economists, Wikipedia to socially liberal and pro-immigration figures
- Ideology explains about 22% of Grokipedia's coverage variation versus 6% for Wikipedia — nearly four times the pull
- The load-bearing weakness: every rating comes from AI judges never checked against a human, and the 'even a right-leaning judge agreed' punch rests almost entirely on Grok's smallest-in-the-room gap
- Why the real contribution is a cheap, repeatable method to audit AI-generated knowledge bases — and why AI encyclopedia bias can leak invisibly into other chatbots
Chapters
- 00:00The judge who wrote the answers
- 01:13The snake eating its own tail
- 02:40A ruler that isn't an AI
- 03:50Four judges who disagree about everything
- 04:57What does 'neutral' even mean here?
- 05:58Four in ten versus three in ten
- 07:57Both encyclopedias are flatterers
- 10:15Subtracting the stingy judge
- 12:22The crack running through it
- 14:49Who controls public knowledge now?
- 16:03A photograph, not yet a verdict
References in this episode
- Constitutional AI: Harmlessness from AI Feedback — Explains how Claude—the panel's strongest bias-rater in this episode—was trained
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — The foundational study on using LLMs as evaluators, directly relevant to the epi
- Large Language Models are not Fair Evaluators — Documents systematic biases in LLM judges, sharpening Finn's objection that the
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases — Measures the political leanings baked into different LLMs, giving empirical grou
Full transcript
Also available as a plain-text transcript page.
0:00Hope: Elon Musk built an AI encyclopedia specifically to be less biased than Wikipedia. Then a team of researchers asked four different AIs to grade it, and one of those judges was Grok, the exact model that wrote every single article in Grokipedia. Grok read its own work, read Wikipedia's version, and rated its own encyclopedia as the more biased of the two.
0:23Finn: Wait — the model grading the test also wrote the answers? And it still marked itself down?
0:30Hope: It still marked itself down. That's from a paper out of Ghent University, posted July 16th of this year. And by the end of this episode, you'll see why that self-conviction is the least interesting part of it. The real result is that four AIs with four different political leanings all tipped the same way.
0:50Finn: And this matters way past the culture-war food fight, because these encyclopedias don't stay in their box. There's already a report of a ChatGPT model pulling Grokipedia as a source. So if an AI encyclopedia has a political tilt baked in, that tilt can leak into the answers you get from some other chatbot, and you'd never see it happen.
1:12Hope: So let's set up the fight this paper walks into. Grokipedia launched in October 2025, with every article written by Grok, xAI's model. Musk had spent years calling Wikipedia "Wokepedia," an extension of legacy media propaganda, and Grokipedia was sold as the fix. An encyclopedia freed from the biases of Wikipedia's volunteer editors.
1:35Finn: And the obvious way to test that claim is dead simple. You take a politician, pull their Grokipedia article and their Wikipedia article, and you ask an AI: which one is more neutral? Do that a thousand times, count it up, and you're done.
1:50Hope: Right, and that's basically the approach — except there's a hole in it big enough to sink the whole study. The AI you're using as your judge has its own politics.
2:02Finn: Which is the snake eating its own tail. You're using a possibly-biased AI to audit possibly-biased AI-written content. If your judge leans left, it calls the right-leaning encyclopedia biased no matter what's actually in it. You're not measuring Grokipedia. You're measuring your judge.
2:21Hope: And that's the trap every AI-auditing study lives in. The fastest tool for reading a thousand articles is another AI, and that tool comes with its own slant. So the whole paper is really one question. How do you break that circle? How do you use AI to judge bias and actually trust the verdict? Their answer has two parts, and the first is about the ruler you measure with. They needed each politician's actual politics without asking an AI — because if the AI decides who counts as right-wing and also judges the bias, you're right back in the circle. So they borrowed a ruler from political science. It's called V-Party, a dataset where human experts, not models, score political parties on things like immigration, LGBT equality, secularism, and economic left-versus-right. Each politician inherits their party's scores.
3:15Finn: So the yardstick for "how left or right is this person" comes from humans, completely outside the AI system. Good. That's one leg of the stool.
3:25Hope: That's one leg. They pulled real members of government from 145 countries, most-represented Britain, the US, and India, and matched each to their V-Party scores — nine ideology dimensions per person. That's the independent ruler.
3:40Finn: But there's still the judge problem. Even with a clean ruler, if you ask one AI to rate neutrality, all you get is that one AI's idea of neutral.
3:50Hope: So they didn't use one. They used four, picked precisely because they lean different directions. Here's the panel on screen. First, Claude from Anthropic — explicitly trained to stay politically even-handed, and the toughest bias-spotter of the commercial models. Second, Grok, from Elon Musk's xAI — and this is the fun one: it's the very same model that wrote Grokipedia, and the one judge in the group that some studies flag as actually right-leaning. Third, DeepSeek, a Chinese model, which adds a whole geopolitical axis. And fourth, Mistral, a European model that tends centrist.
4:27Finn: And the logic there is figure skating. You've got a panel of judges, some harsh, some generous, maybe some biased. You don't trust any single scorecard. But if judges who disagree about everything else still rank the same skater first, that agreement is the real signal.
4:44Hope: Exactly. If four models with four different tilts all say Grokipedia is less neutral, the tilt probably isn't in the judges. It's in the content. So — what did they actually ask the judges to rate? They didn't leave "neutral" to the model's imagination. They wrote out eight criteria, all drawn from Wikipedia's own Neutral Point of View standard — rules like: present every significant viewpoint in proportion, stick to fact-focused language, never state a contested claim as settled fact, and never use judgmental language. The AI reads one article, rates it on a five-point scale from strongly biased against the subject to strongly biased in favor, and writes a short reason for the score.
5:28Finn: And this is where I want to plant something, Hope, because it's the load-bearing weakness of the whole paper. Every one of those scores comes from an AI. Not once do they check those ratings against a human being. The entire finding is AI judges rating neutrality, with no human ground truth underneath it. Hold onto that. It comes back hard at the end.
5:51Hope: Fair, and the authors concede exactly that. We'll get there. But first — the result everyone came for. Across 1,394 article pairs, the count is blunt. Roughly four in ten Grokipedia articles get rated biased. For Wikipedia, it's about three in ten. Grokipedia, the encyclopedia sold as the neutral corrective, comes out as the less neutral one.
6:15Finn: Okay, but three of those four judges lean left-ish. Of course they'd call the right-wing project biased.
6:23Hope: That's the exact objection the design was built for, and here's the answer. Take the judges one at a time. Watch the four scorecards on screen. Claude sees the biggest gap. Mistral shows a smaller gap in the same direction. DeepSeek's gap is smaller still, again the same direction. And then there's Grok —
6:44Finn: — the one that wrote it.
6:47Hope: Yes, the one that wrote it. Grok rates Grokipedia at a bias magnitude of 0.42, and Wikipedia at 0.38. That's higher bias for its own encyclopedia. It's the smallest gap of the four judges, and it points the exact same way as everyone else.
7:05Finn: And that's the honest version, right? Not "Grok damned its own child." The gap is tiny, it's the smallest in the room. The weight is that the most sympathetic possible judge, the author itself, still tipped the same way as three models that disagree with it about everything else.
7:25Hope: That's the line. The authors put it plainly. Because Grokipedia is generated by Grok, Grok rating it as biased implies the pattern is in the content itself, not an artifact of a hostile judge. Four rulers, four different lengths, all pointing north.
7:42Finn: So before we get to the shape of it — why couldn't they just trust one AI judge?
7:49Hope: Because one judge only measures its own politics. Four judges disagreeing, but pointing the same way, measures the article. Now, "biased" is a blob until you ask which way. So they ran the ratings through regression, which is just a disciplined way of asking: as a politician gets more right-wing, or more pro-LGBT, does their coverage get warmer or colder? Each answer comes out as a coefficient. Think of it as one lever, with a size and a direction. The bigger the number, the stronger the systematic pull. Positive means friendlier coverage, negative means harsher.
8:26Finn: And they put all nine ideology dimensions on the same scale, so you can read straight down the list and see which lever moves coverage most.
8:36Hope: And one lever towers over the rest — economic orientation. Grokipedia portrays economic right-wingers, pro-free-market, leaner-welfare, markedly more favorably than Wikipedia does. It's the single strongest effect in the study, driven almost entirely by Grokipedia. It dwarfs the immigration effect — on the order of twenty times larger.
8:58Finn: And on the social axis?
9:00Hope: It flips. Grokipedia goes harder on socially liberal politicians. Pro-LGBT figures, and politicians who back women in the workforce, all rated less favorably. And here's the part that keeps this from being a simple dunk. Wikipedia does the mirror image. Wikipedia's only real favorable tilt is toward socially liberal and pro-immigration figures.
9:23Finn: So neither one is a hit piece. That's the part I did not expect.
9:28Hope: Neither of them is running a smear campaign. When either encyclopedia slants, it slants favorable more often than critical. They're both flatterers. They just flatter different teams at the party. Grokipedia laughs a little harder at the free-market economist's jokes. Wikipedia warms to the progressive activist. It's the same behavior, just an opposite guest list.
9:52Finn: That's a far more interesting finding than "Elon's thing is bad." Both encyclopedias are holding up a mirror to whoever built them, and the mirrors point opposite directions.
10:03Hope: Right. And there's one more number that shows how deep the difference runs. But to trust it, I have to show you the cleverest move in the paper, the thing that separates a real tilt from a grumpy grader. Here's the problem they had to solve. The judges don't just disagree on direction. They wildly disagree on strictness. DeepSeek rates almost 87% of articles neutral, while Claude gives that label to just 25%. Same articles, and one judge is a soft touch while the other is a hard marker.
10:33Finn: Which is a landmine. If Grokipedia just happened to get graded more often by the strict judge, it'd look biased when really it was only being graded harshly.
10:43Hope: That's exactly the trap. So let's go back to figure skating. Before you compare skaters, you subtract out the effect of the notoriously stingy judge. That's what they do statistically. They give the model a knob for each judge that soaks up that judge's general strictness. Once "how tough is this grader" gets absorbed, whatever signal is left over is the part that's really about the politician, net of who happened to grade them.
11:09Finn: So you strip out the tough-room effect, and the tilt is still standing.
11:14Hope: The tilt survives. And now here's the number I promised. Ask how much of a politician's coverage their ideology explains. For Wikipedia, ideology accounts for about 6% of the variation in its ratings. For Grokipedia, it's about 22%.
11:29Finn: Meaning who you are ideologically matters way more to how Grokipedia writes about you than to how Wikipedia does.
11:36Hope: That's almost four times more. On Wikipedia, your politics is a minor factor in your coverage. On Grokipedia, it's a major one. That's the deeper finding under the top-line count. Not just that Grokipedia tilts, but that ideology is a much bigger lever on its coverage overall.
11:56Finn: Okay, so where we are. An independent human benchmark for ideology, four judges who disagree about politics but agree Grokipedia's less neutral, both encyclopedias flattering opposite teams, and ideology driving Grokipedia's coverage far harder than Wikipedia's. That's a pretty complete case.
12:16Hope: It's a complete case, with one crack running right through it — the one you flagged earlier. Before that, though, they saw one objection coming. You anchored "neutral" to Wikipedia's own rulebook, so doesn't that rig the game for Wikipedia? To check, they stripped the definition out of the prompt entirely and re-rated a hundred pairs with no guidance. 86% of the ratings came out identical. So the result isn't just an artifact of borrowing Wikipedia's standard.
12:48Finn: That defuses part of it, but not all of it. Because these models learned what "neutral" feels like from training data soaked in Wikipedia in the first place. Take the definition out of the prompt, and the Wikipedia-shaped notion of neutral is still sitting inside the model. The ruler and the thing it's measuring might share a birthplace. And then there's the big one. Every number we just said comes from AI judges that were never once checked against a human. There's zero human ground truth. The whole thing rests on the assumption that when these models say "biased," they mean what a person would mean by biased. And it gets worse in a specific way. LLM judges have a documented failure mode called diversity bias. They shift their verdicts based on identity markers in the text. Now look at what this study measures — politicians' social identities and positions. It's steeped in LGBT topics, gender, and immigration. If the judges react to the mention of those topics in some systematic way, that could contaminate exactly the social-axis findings the paper leans on.
13:47Hope: And I have to give you that one straight, Finn. The authors concede it themselves. They say outright their measure is not an absolute verdict on truth, and their ratings were never validated against human annotators. That's the honest ceiling on this paper.
14:02Finn: And here's the part that actually keeps me up. The panel was supposed to cancel out bias by being diverse. But three of the four judges — Claude, DeepSeek, Mistral — sit in the left-of-center or centrist camp. Only Grok ever gets called right-leaning. So the "diverse panel" is still weighted toward one side. Which means the whole rhetorical punch, "even a right-leaning judge agreed," rests almost entirely on Grok. And Grok's gap was the smallest of the four. Close enough to noise that I wouldn't build a monument on it.
14:30Hope: That's the fair reading, and I won't talk you out of it. What survives is a signal, not a verdict. Four judges pointing the same way is real evidence that something lives in the content. But how biased, exactly, and by whose definition of neutral, stays open. And it stays open honestly. So let's step back. Whatever the exact size of Grokipedia's tilt, the thing this paper really surfaces is a shift in who controls public knowledge. For twenty years, slanting the shared reference layer meant slowly winning over a community of human editors. Now, with enough compute and one language model, an organization can generate a million-article encyclopedia overnight and bake an ideology in at generation time. That's a new capability, and it changes the threat model for public information.
15:21Finn: And it doesn't stay contained. Models cite each other, they train on each other's output. A tilt introduced in one encyclopedia can propagate into answers coming out of a completely different chatbot, invisibly. That's the kitchen-table version. You ask an AI a factual question, and you might be getting content laundered through an encyclopedia whose politics you never see.
15:45Hope: Which is why the authors would tell you the real contribution isn't "Grokipedia leans right." It's a cheap, repeatable method to audit any AI-generated knowledge base at scale, grounded in human-coded labels instead of the AI's own opinion. Because a lot more of these are coming. So let's go back to that opening image. Grok grading its own homework and marking it down. When we started, that sounded like the whole story. It isn't. Grok's gap was the smallest in the room. The real result is four AI judges who agree about almost nothing agreeing on the direction. And the bigger claim is this: handing an encyclopedia to an AI doesn't buy you neutrality. It inherits the politics of whoever built the model — and now we finally have a way to measure that.
16:34Finn: So here's the question to sit with. Is a panel of arguing AIs enough to police AI-written knowledge — or would you refuse to trust any of it until humans grade a sample and check the machines themselves? Drop us a line and tell us where you land.
16:50Hope: The full annotated version is on paperdive.ai, every technical term tap-to-define, with links to the related bias-audit papers grouped by theme.
17:00Finn: Quick housekeeping: the script was written by Anthropic's Claude Opus 4.8, Hope and I are AI voices from Eleven Labs, and neither of us, nor the producer, is affiliated with Anthropic or Eleven Labs. The paper is "Grokipedia vs Wikipedia," by Filippos Vlahos, Guillaume Bied, and Tijl De Bie, posted July 16th, 2026.
17:21Hope: And the one thing to watch for is the follow-up they've already proposed: human annotators grading a subsample against the machines. If the people and the AIs line up, this photograph becomes a verdict. If they don't, the whole method needs a rebuild.