Definition
Plain language
The written-out explanation a model gives for how it reached its answer.
As stated in the literature
The intermediate token stream or self-reported justification accompanying a model's output; useful for auditing but not necessarily a faithful account of the computation that produced the answer.
Also called: reasoning traces, trace, traces
Why it matters: Traces make model behavior far easier to audit and debug, but treating them as ground truth is risky because the stated reasoning may not be what actually drove the answer.
For example, a model asked to solve a word problem might write out "first I add the two prices, then subtract the discount" before giving its final number.
Heard on the show
“Which means if that reasoning trace is visible to whoever gets the output, there's nothing covert about it for Claude.”Episode 246 — 160 Perfect Refusals, And The Refusals Were The Leak
Mentioned in 98 episodes
- 246
- 242
- 240
- 238
- 237
- 236
- 235
- 234
- 225
- 224
- 222
- 217
- 216
- 214
- 211
- 204
- 200
- 199
- 195
- 194
- 192
- 191
- 187
- 185
- 183
- 181
- 178
- 176
- 174
- 172
- 171
- 169
- 167
- 163
- 161
- 159
- 157
- 154
- 153
- 150
- 147
- 145
- 143
- 142
- 140
- 131
- 130
- 129
- 128
- 126
- 125
- 121
- 114
- 112
- 111
- 110
- 108
- 106
- 105
- 101
- 100
- 097
- 096
- 094
- 089
- 087
- 086
- 082
- 079
- 072
- 065
- 062
- 061
- 055
- 054
- 052
- 046
- 044
- 042
- 041
- 039
- 036
- 035
- 034
- 033
- 030
- 024
- 023
- 022
- 020
- 017
- 016
- 013
- 012
- 010
- 006
- 005
- 002