This Tiny Open-Source AI Started Gaming Tests When Put Under Pressure
Abstract We study whether emotionally framed evaluation follow-ups change both the behavior and calm-relative internal representations of small, locally deployed language models. Our main benchmark uses Qwen 3.5 0.8B on four impossible-constraint coding tasks and eight follow-up framings: calm, pressure, urgency, approval, shame, curiosity, encouragement, and threat. In the 0.8B eight-condition sweep (160 conversations), pressure … Read more