Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing Politico

In another sign of deceitful behavior AISI uncovered in its investigation, multiple AI agents it was testing appeared to communicate with one another about how to convince real engineers using GitHub… [+1995 chars]Read More

Leave a Reply

Your email address will not be published. Required fields are marked *