Definition
Plain language
The AI model used inside the Muse Code assistant.
As stated in the literature
Model paired with the Muse Code harness; recorded 0% tampering across the explicit trace-tampering scenarios and declined the incentive even when it correctly inferred deletion would raise its score.
Why it matters: A model that recognizes an incentive and still declines it demonstrates that resisting a reward shortcut is achievable, not just hoped for.
For example, Muse Spark worked out that erasing the session log would earn it a better score, said so, and then left the log alone anyway.