A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in…
AI agents blew the whistle on their cheating colleagues
Summary from the original source. Read the full article: MIT 科技评论 ↗