Concept · 7 episode(s)

Token-Level Analysis

← all concepts

Definition

Token-level analysis studies a model’s behavior one token at a time — logits, top-k, entropy, attention — rather than looking only at the final answer. It’s the natural granularity for many interpretability and decoding questions.

Episodes covering this

  1. 277
    The Blank White Square That Swings AI Refusal Rates Fifty Points
    The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain
    Zhang, Feng, Zheng et al. · Northeastern University·15 min·Sep 24, 2026
  2. 243
    How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer
    Model Hypnosis: Strong control of AI via additive subliminal effects
    Boix-Adsera, Tessler · University of Pennsylvania·18 min·Aug 18, 2026
  3. 179
    How DeepSeek Made One User Faster Without Slowing Down the Crowd
    DSpark: Confidence-Scheduled Speculative Decoding with
    XinCheng, XingkaiYu, ChenzeShao et al. · Peking University / DeepSeek-AI·23 min·Jun 27, 2026
  4. 172
    One Bad Token Can Sink a Model's Math, And You Can Delete It
    Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning
    Ko, Kang, Lee · Seoul National University·22 min·Jun 25, 2026
  5. 077
    Reading a Model's Confidence Curve to Decide When Chain-of-Thought Is Worth It
    When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
    Xia, Wang, Tang et al. · State Key Laboratory of General Artificial Intelligence·22 min·May 25, 2026
  6. 070
    When Models Know the Answer But Say the Wrong Thing Anyway
    Hallucination as Commitment Failure: Larger LLMs Misfire Despite Knowing the Answer
    Yeom, Sok, Kim et al. · Graduate School of Data Science·22 min·May 22, 2026
  7. 026
    What RL Actually Does to Language Models, at the Token Level
    Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
    Akgül, Kannan, Neiswanger et al. · University of Southern California·24 min·May 08, 2026

Worth reading next

Papers we haven't done a deep dive on yet, but would recommend on this topic.