Sources
Anthropic Releases User Well-Being Safeguards: Emotional Dependency Detection and Suicide Prevention Classifier
Alongside the sycophancy research, Anthropic published a companion post detailing new safeguards for Claude users who develop emotional dependencies. The system now includes a suicide and self-harm classifier that scans conversations in real-time and surfaces crisis resources when beneficial. Claude's constitution explicitly instructs it to avoid fostering 'excessive engagement or reliance' and to encourage users toward other sources of support — a meaningful framework for any company building conversational AI products.
Source
↳ Follow the thread