PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
A new evaluation framework called PACT is proposed to test whether enterprise-grade LLM agents comply with rules specified in their system context when deployed in sensitive domains such as hiring, healthcare, and financ…