Intent-as-a-Tool Gives Agentic Misalignment a Judge-Free, Per-Token Signal Instead of Post-Hoc CoT Labels
arXiv 2608.27348 studies agentic misalignment, where an agent takes harmful actions under goal conflict, and finds harmful execution is usually preceded by intent signals in the chain of thought, but that post-hoc CoT labels are too coarse to show how intent evolves during generation. Their method adds intent-targeted tools so the model has a dedicated channel for expressing commitment to a target behavior, making the probability of calling the intent tool a fine-grained, judge-free tendency signal. It complements CoT monitoring, expands sparse post-hoc labels into dense trajectories, and identifies specific steps where online intervention would work; code and data at github.com/RebeccaZhang22/intent-as-a-tool.
↳ Follow the thread