Goals become criteria
Define report requirements: all IDs, exact dates, sources, open questions and actions within permissions. Separate critical errors from presentation defects. Attractive formatting does not offset an invented order.
Evaluations use cases and criteria to check behavior. This course supplies a business rubric rather than mandating a tool. Manual review, automated checks and additional evaluators can be combined.
Representative cases
Include normal, incomplete, conflicting and duplicate cases and tool failures. Keep some cases aside for new-version checks rather than tuning prompts to every example. Process experts define expected answers.
Initial samples reveal defects without proving universal reliability. Record sample size, selection and limits. One difficult case does not support broad conclusions.
Complementary measures
Date accuracy is correct fields divided by reviewed fields; coverage is treated IDs divided by input IDs. Total time includes preparation, execution, review and correction. Include service, maintenance and operating costs when known.
In a fictional sample, 18 correct dates out of 20 equals 90%. If both errors create false commitments, the average cannot justify approval. Zero critical errors in a pilot sample can be an acceptance condition without guaranteeing future perfection.
Productivity and business value
Four released weekly hours are capacity. Realized economics need evidence of how they were used: more customer service, lower contracted effort or better fulfillment. Do not automatically convert hours to expected sales or attribute all business changes to AI.
Compare similar work and periods, document other changes and unintended effects and use a comparison group where feasible. Conclusions must match evidence.
Explicit decisions
Set continue, adjust and stop conditions before the pilot. Report denominators, failures and limits. Recheck relevant cases after changing instructions, models, sources or permissions. Evaluation continues during operation.