Vocabulary for building AI agents that plan, act, and recover from failure in production systems.
A test suite that runs an agent against a fixed set of tasks and scores its outputs automatically, so changes to a prompt or model can be compared side by side.