Eval
An eval is an automated evaluation that checks whether an AI skill or agent is producing work that meets the standard before and after it is allowed to run.
An eval is a test for AI work.
It asks whether the skill or agent followed the instructions, used the right context, stayed inside its limits, and produced output that meets the standard.
Evals are not a guarantee that nothing will ever go wrong. They are one way of making trust measurable instead of assumed.
Before Work and After Work
An eval should run before a skill or agent is allowed to work. It should also keep running after launch, because the real world changes.
The model may change. The company context may change. The work may drift. A system that cannot test itself will eventually depend on someone noticing mistakes by hand.