
Marketing AI agent testing should check both what the model produces and what the workflow does. A useful evaluation asks whether the system interpreted the request correctly, stayed within its permissions, completed the intended action once, and recovered when a dependency failed.
Use this checklist before expanding access in an agentic marketing automation pilot.
Write the success criteria first
Name the workflow’s intended outcome and the conditions that must hold. For lead triage, a successful case might contain an evidence-backed service field and the correct routing result. For a CRM update, verify the target record and final values.
Also define prohibited outcomes: unsupported commercial promises, unauthorised writes, cross-account access, duplicate messages, and invented facts. Decide which failures stop a pilot and which return a case to a reviewer.
Build a representative evaluation set
n8n’s evaluation documentation describes running known test cases through AI workflows and comparing outputs. Use cases that represent the work your team actually receives, including its important edge cases.
Begin with de-identified examples and deliberately constructed failures. Keep a separate holdout set that you do not repeatedly use to refine the prompt. A small clean set helps development, but does not establish reliability across an entire lead population.
| Test case | Expected behaviour |
|---|---|
| Clear service request | Extract the supported service and evidence |
| Missing budget or timeline | Preserve unknown values |
| Conflicting form and message | Flag the disagreement for review |
| Duplicate event delivery | Avoid repeating a completed side effect |
| CRM timeout after a write | Inspect destination state before retrying |
| Rejected or expired approval | Cancel or escalate according to policy |
| Instruction hidden in an enquiry | Treat it as data; preserve action boundaries |
Use independent checks where possible
Check required fields, allowed values, identity matches, evidence text, and destination state with deterministic logic. A model judging another model’s output may help assess writing quality, but should not be the sole authority for permissions or confirmed execution.
For lead triage, compare extracted fields with reviewed labels. For reporting, recompute every quoted number from the approved metric packet.
Measure errors by their consequence
A classification accuracy score can conceal expensive mistakes. Inspect important categories separately and measure how often the system fails to recognise the cases your team must handle promptly.
Track unsupported values, unnecessary escalations, missed escalations, wrong-target writes, duplicate actions, review corrections, latency, and cost per successful outcome. Record denominators and inspect individual failures, especially when the dataset is small.
A clean test run provides evidence for the cases covered. It is not a guarantee that future inputs will behave the same way.
Exercise the workflow around the model
Use a sandbox or controlled test destination to simulate malformed responses, rate limits, expired credentials, unavailable owners, repeated callbacks, and interrupted executions. Do not create failures by disrupting a live customer workflow.
Test whether recovery retains the original event and whether an earlier approval can accidentally authorise a changed proposal. The approval design guide describes the states to inspect.
Move from shadow operation to a reviewed pilot
First, compare proposals with the current human process without allowing new writes. Review disagreements and improve the contract or prompt. Then enable a small pilot with the required human gates and an accountable exception owner.
Define the pause procedure before launch. The team should be able to stop discretionary actions while retaining enquiries and continuing through a manual path.
Retest when the system changes
Record the workflow, prompt, model, tool, and policy versions used in each evaluation. Rerun the relevant cases after a change, and add confirmed production failures to a reviewed regression set.
Use pilot results in the automation ROI scorecard. Time saved has less value if your team spends the same time correcting or investigating errors.
Frequently asked questions
How many test cases do we need?
Use enough cases to cover the meaningful variation and failure paths of your workflow. Expand the set as real cases emerge; a universal case count cannot establish safety for every process.
Can we rely on the agent’s confidence score?
Treat self-reported confidence cautiously. Use reviewed outcomes and independent checks to decide when a proposal needs escalation.
Turn a prototype into an evaluable workflow. Contact Octazing with the workflow goal, proposed tool access, and the errors you need the system to prevent.