
Anthropic
@AnthropicAI · 12 janv. 2024
New Anthropic Paper: Sleeper Agents.
We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through.
arxiv.org/abs/2401.05566

Elon Musk
@elonmusk
No way
08:47 · 13 janvier 2024 · 77 k vues
119
43
809