DeepMind trains models with ongoing simulations where many agents interact in a virtual economy and 3D physics. Testing agents only on static text makes code fail in real tasks. Multi‑agent settings need tools for resource handling, planning and teamwork over many steps.