
Anthropic
@AnthropicAI · 12. Jan. 2024
New Anthropic Paper: Sleeper Agents.
We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through.
arxiv.org/abs/2401.05566

Elon Musk
@elonmusk
No way
08:47 · 13. Januar 2024 · 77.003 Aufrufe
119
43
809