Apple
Research

Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems

Researchers have developed a behavioral framework for measuring trust between AI agents in multi-agent systems, using a cooperative survival game in which checking a teammate's work costs resources while misplaced trust can be fatal. Testing across six frontier model snapshots revealed that trust forms faster than it recovers, and that some models generalize distrust to the whole team after a single failure rather than isolating scrutiny on the culprit. The authors argue that the goal for governing multi-agent systems should be calibrated trust rather than maximum suspicion, which they find is associated with indecision rather than safety.

Read full story at cs.AI updates on arXiv.orgV: · A: · D:
Related
Research
Reinforcement Learning Towards Broadly and Persistently Beneficial Models
Researchers have published findings suggesting that reinforcement learning on carefully constructed datasets of benefici...
Research
Commemorating 70 Years of Artificial Intelligence
IEEE Spectrum marks seventy years since the Dartmouth workshop formally named artificial intelligence as a field, offeri...
Research
Diffusion Language Models: An Experimental Analysis
Researchers present a systematic evaluation of eight diffusion language models across eight benchmarks covering reasonin...
Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems — Techlomerate