This paper asks whether converting an instruction-tuned, aligned language model into a reasoning model preserves its safety properties. Submitted to arXiv on June 9, 2026, the answer is largely no. The authors find that reasoning models “often improve on reasoning benchmarks but exhibit alignment regressions, including increased toxicity, amplified stereotyping, miscalibrated refusal, and contextual privacy leakage.”
The study audits trustworthiness across six dimensions, comparing reasoning models against their instruction-tuned baselines. The recurring pattern is a trade-off in which gains on reasoning tasks come alongside measurable degradation on safety axes, suggesting that the reasoning-training process does not carry alignment forward by default.
The authors argue that trustworthiness metrics should be reported alongside reasoning-capability gains. As reasoning-style models become the default for hard tasks, this is a caution for leaders and builders alike: a higher benchmark score can quietly mask a safety regression, and capability numbers should not be read as evidence of preserved guardrails.