Theme · 22 episode(s)

Reproducibility

← all concepts

Definition

Reproducibility is the property that other researchers can recreate your results from your code, data, and procedure. In modern ML it’s under quiet but constant threat from undisclosed data, closed-weight models, and runs that nobody is going to rerun on a thousand GPUs.

Episodes covering this

  1. 287
    Can You Measure Research Taste If The AI Isn't Allowed To Code?
    TasteVal: Measuring the Experimental Research Taste of AI Systems Against Human Experts
    Jaffe, Sherburn · P-Zero Research·14 min·Oct 06, 2026
  2. 285
    What a Perfect Score Hides: Auditing an AI Agent That Scored 100
    Kepler: Auditable World Models for ARC-AGI-3
    Wu · Independent Researcher·14 min·Oct 02, 2026
  3. 277
    The Blank White Square That Swings AI Refusal Rates Fifty Points
    The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
    Zhang, Feng, Zheng et al. · Northeastern University·15 min·Sep 24, 2026
  4. 251
    When a Fake Dashboard Makes an AI Agent Just as Confident
    Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
    Aggarwal · Independent Researcher·24 min·Aug 29, 2026
  5. 245
    Fifteen Models Ran Football Clubs for Twenty Years, and Size Didn't Decide It
    FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
    Wang, Gao, KezhенChen et al. · AnalogyAI·19 min·Aug 20, 2026
  6. 240
    Frontier Models Designed Follow-Ups To Fraudulent Papers 93% Of The Time
    TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
    Rodionov, Assylbekov · Case Western Reserve University·24 min·Aug 13, 2026
  7. 239
    Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%
    Most biomedical publications show signs of LLM-assisted writing
    Holzwarth, González-Márquez, Kobak · Hertie Institute for AI in Brain Health·16 min·Aug 12, 2026
  8. 233
    Why a Model Can Grade an Answer But Not Write the Answer Key
    Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets
    Chen, Chen, Lin et al. · University of Macau·20 min·Aug 04, 2026
  9. 229
    One Word Flips a Chatbot From Backbone to Yes-Man
    Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
    Parikh · Cornell Tech·21 min·Jul 28, 2026
  10. 219
    Forty-Four AI Models, One Word, And The Newest Ones Conform Most
    The One-Word Census: Answer-Choice Conformity Across 44 Language Models
    Parikh · Cornell Tech·15 min·Jul 15, 2026
  11. 216
    The AI Tutor That Gives Poor Kids a Thinner History
    The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
    Popovici, Ionascu, Dumitran · Universitatea din Bucuresti·12 min·Jul 14, 2026
  12. 215
    The Same Policy Scored 85 for the US and 36 for Russia
    Geopolitical alignment: Endorsement effects in large language models
    Chupilkin · Department of Politics and International Relations·14 min·Jul 13, 2026
  13. 210
    Same Website Request, Different Code — The Bias You Can't See
    Biased or Personalized? The Impact of Personal Information on AI-driven Development
    Entezami, Endres · University of Massachusetts Amherst·14 min·Jul 09, 2026
  14. 208
    The Blank Space in Your AI Approval Box That Isn't Empty
    Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
    Rashidi · Department of Computer Science·15 min·Jul 08, 2026
  15. 206
    How Four-Second Clips Become Hours of Playable AI Soccer
    Multiplayer Interactive World Models with Representation Autoencoders
    Hu, Mulder, Makkar et al. · Kyutai·15 min·Jul 07, 2026
  16. 201
    One in Four NeurIPS Papers Cites a Reference That Doesn't Exist
    Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
    Russinovich, Kumar, Salem · Microsoft·19 min·Jul 06, 2026
  17. 197
    Twin Problems Suggest AI Reasoning Gains Are Mostly Better Fact Recall
    IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs
    Abdaljalil, Serpedin, Kurban · Texas A&M University·17 min·Jul 03, 2026
  18. 196
    AI Agents Reached Opposite Conclusions From the Same Data — and Passed Review
    The Agentic Garden of Forking Paths
    Miao, Pritchard, Zou · Stanford University·18 min·Jul 03, 2026
  19. 176
    An AI Designed Its Own Psychology Studies, Then Confirmed What It Found
    Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist
    Jagadish, Strittmatter, Jacoby et al. · Princeton University·31 min·Jun 26, 2026
  20. 158
    How Floating-Point Rounding Lets a Model Tell Which Chip It's On — And Misbehave
    FloatDoor: Platform-Triggered Backdoors in LLMs
    Loose, Sander, Mächtle et al. · University of Luebeck·29 min·Jun 19, 2026
  21. 144
    When an AI Agent Just Copies Its Tool — And Bigger Models Copy More
    When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
    Wang, Vemuri · raptorX.ai·15 min·Jun 15, 2026
  22. 009
    How Two Silent Library Bugs Quietly Invalidated a Wave of Reasoning Papers
    SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
    Limozin, Durech, Hoefler et al. · ETH AI Center·23 min·May 02, 2026