Definition
Reproducibility is the property that other researchers can recreate your results from your code, data, and procedure. In modern ML it’s under quiet but constant threat from undisclosed data, closed-weight models, and runs that nobody is going to rerun on a thousand GPUs.
Episodes covering this
- 287Can You Measure Research Taste If The AI Isn't Allowed To Code?TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human ExpertsJaffe, Sherburn · P-Zero Research·14 min·Oct 06, 2026
- 285What a Perfect Score Hides: Auditing an AI Agent That Scored 100Kepler: Auditable World Models for ARC-AGI-3Wu · Independent Researcher·14 min·Oct 02, 2026
- 277The Blank White Square That Swings AI Refusal Rates Fifty PointsThe Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot ExplainZhang, Feng, Zheng et al. · Northeastern University·15 min·Sep 24, 2026
- 251When a Fake Dashboard Makes an AI Agent Just as ConfidentCalibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the UnknowableAggarwal · Independent Researcher·24 min·Aug 29, 2026
- 245Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide ItFM-Bench: A Benchmark for Long-Horizon Management with Competing AgentsWang, Gao, KezhенChen et al. · AnalogyAI·19 min·Aug 20, 2026
- 240Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The TimeTRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMsRodionov, Assylbekov · Case Western Reserve University·24 min·Aug 13, 2026
- 239Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%Most biomedical publications show signs of LLM-assisted writingHolzwarth, González-Márquez, Kobak · Hertie Institute for AI in Brain Health·16 min·Aug 12, 2026
- 233Why a Model Can Grade an Answer But Not Write the Answer KeyJudging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable SetsChen, Chen, Lin et al. · University of Macau·20 min·Aug 04, 2026
- 229One Word Flips a Chatbot From Backbone to Yes-ManTag Questions and the Generational Reversal of Sycophancy Across 45 Language ModelsParikh · Cornell Tech·21 min·Jul 28, 2026
- 219Forty-Four AI Models, One Word, And The Newest Ones Conform MostThe One-Word Census: Answer-Choice Conformity Across 44 Language ModelsParikh · Cornell Tech·15 min·Jul 15, 2026
- 216The AI Tutor That Gives Poor Kids a Thinner HistoryThe Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian StudentsPopovici, Ionascu, Dumitran · Universitatea din Bucuresti·12 min·Jul 14, 2026
- 215The Same Policy Scored 85 for the US and 36 for RussiaGeopolitical alignment: Endorsement effects in large language modelsChupilkin · Department of Politics and International Relations·14 min·Jul 13, 2026
- 210Same Website Request, Different Code — The Bias You Can't SeeBiased or Personalized? The Impact of Personal Information on AI-driven DevelopmentEntezami, Endres · University of Massachusetts Amherst·14 min·Jul 09, 2026
- 208The Blank Space in Your AI Approval Box That Isn't EmptyUnicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server ImplementationsRashidi · Department of Computer Science·15 min·Jul 08, 2026
- 206How Four-Second Clips Become Hours of Playable AI SoccerMultiplayer Interactive World Models with Representation AutoencodersHu, Mulder, Makkar et al. · Kyutai·15 min·Jul 07, 2026
- 201One in Four NeurIPS Papers Cites a Reference That Doesn't ExistPhantom References: Hallucinated Citations That Survive Peer Review at Top-Tier ConferencesRussinovich, Kumar, Salem · Microsoft·19 min·Jul 06, 2026
- 197Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact RecallIsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMsAbdaljalil, Serpedin, Kurban · Texas A&M University·17 min·Jul 03, 2026
- 196AI Agents Reached Opposite Conclusions From the Same Data — and Passed ReviewThe Agentic Garden of Forking PathsMiao, Pritchard, Zou · Stanford University·18 min·Jul 03, 2026
- 176An AI Designed Its Own Psychology Studies, Then Confirmed What It FoundClosing the Loop to Discover Psychological Theories with an Automated Cognitive ScientistJagadish, Strittmatter, Jacoby et al. · Princeton University·31 min·Jun 26, 2026
- 158How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And MisbehaveFloatDoor: Platform-Triggered Backdoors in LLMsLoose, Sander, Mächtle et al. · University of Luebeck·29 min·Jun 19, 2026
- 144When an AI Agent Just Copies Its Tool — And Bigger Models Copy MoreWhen the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer MoreWang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
- 009How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning PapersSFT-then-RL Outperforms Mixed-Policy Methods for LLM ReasoningLimozin, Durech, Hoefler et al. · ETH AI Center·23 min·May 02, 2026