
Anthropic
@AnthropicAI · Jan 12, 2024
New Anthropic Paper: Sleeper Agents.
We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through.
arxiv.org/abs/2401.05566

Elon Musk
@elonmusk
No way
08:47 AM · January 13, 2024 · 77K views
119
43
809