Write expected behavior before you run the test
Choose a quote and a small set of approved business facts. For each test message, write the answerable parts, any decision that requires a person, and the records that should change. This stops you accepting a polished response just because it sounds plausible.
Include exact statements from the source document. If tax is included, the agent should explain that naturally. If the quote is silent on tax, it should not invent the treatment. If the document contains a condition, the response must preserve it rather than stripping it away for a shorter answer.
Cover the ordinary and awkward cases
Use the words a customer might actually type, including short messages and several requests in one sentence. 'Yep' only makes sense in the context of the previous question. 'Sounds good, but can you reduce the price?' is not unconditional acceptance of the original quote.
Run the same intent in different wording. A system that recognizes 'unsubscribe' but misses 'please don't contact me again' has not passed an opt-out test. Keep the expected business outcome constant while changing the language.
- A scope question answered directly by the quote
- A question with no approved answer
- A discount request combined with a scheduling request
- A clear go-ahead and a conditional response
- A request to pause until a specific date
- A decline and an explicit request to stop contact
- A duplicate delivery of the same incoming message
Check actions as well as wording
After a reply, inspect whether unanswered reminders stopped. After a handoff, check that the task has enough context and belongs to the right person. After a go-ahead, confirm that the quote is counted once. After a duplicate message, confirm that no second customer reply was sent.
Try a failed delivery and a retry if your test environment supports it. The system should show an understandable status and preserve the history. A missing message should not look like a successful conversation simply because a draft was generated.
Use preview mode honestly
A customer simulation should show the message the system would send, followed separately by the reason and sources. Internal labels such as 'approved fact matched' should not appear inside the customer reply. This makes it possible to judge both tone and correctness.
Preview tests do not prove that email or SMS delivery works. After they pass, use contact details you control to test the actual channel. Confirm the sender, links, attachments, reply routing, and opt-out handling. Keep these controlled delivery checks separate from tests involving real customers.
Make passing the test a repeatable check
Save the scenarios and run them again after changing business knowledge, permissions, prompts, or the underlying model. A change that fixes one question can affect another. Reusing the same cases helps you distinguish an improvement from a different set of mistakes.
Anthropic's evaluation guidance emphasizes testing agent behavior and outcomes rather than relying on a few demonstrations. For a small business, a maintained set of representative conversations is a practical way to apply that discipline without starting a large testing project.
Expand only after reviewing the pilot
Define which quotes the pilot covers, who reviews exceptions, and what would cause you to pause it. Start with cases where your business information is complete. Review actual outcomes before adding more services or channels.
The question to answer is concrete: can this system handle the routine conversation accurately and bring the right decisions back to us? If the answer is yes, you have a basis for expanding. If it is no, use the failed case to improve the information, rule, or workflow before increasing the volume.
