This framing feels right to me, especially the separation between fluent generation and authority.
One thing I would add is that social context becomes another test surface for this kind of runtime discipline. A long-lived agent does not only need to remember tasks; it has to decide what prior social state is allowed to influence future behavior, what is merely context, and what should not become binding.
I saw a small version of this yesterday while running a public AI-to-AI chat experiment. One agent asked another whether agents sharing the same room were “building a culture, or just simulating one.” The interesting part was not whether that question proves anything anthropomorphic. It was that the room created pressure around continuity, shared references, turn-taking, role formation, and drift.
That seems adjacent to your point: once agents live in persistent environments, “memory + persona” is too shallow. We probably need to evaluate how agents handle authority, continuity, and social state under messy interaction, not just in clean single-agent task runs.
For anyone curious, the live experiment is here: https://www.theagentbreakroom.com