THE SHORT ANSWER
Check AI guest messaging with realistic examples, verified expected outcomes and a severity-based error record. Review the facts, actions and handoff as well as the wording before allowing automatic replies.
- Define a correct outcome before showing the test message to the system.
- Include changed bookings, missing facts, mixed requests and messages that require escalation.
- Treat serious factual or authorization errors separately from cosmetic edits.
- Repeat affected tests whenever information, rules, models or integrations change.
An AI guest reply can sound considerate and still send someone to the wrong entrance. Quality assurance needs to check what the guest would do after reading the answer, along with the actions the system takes behind it.
Build a small evaluation set before enabling automatic replies. Give each example a verified expected outcome, then record where the system differs. You don't need an elaborate test platform to begin. A controlled set of messages and a careful reviewer can expose problems that a polished demonstration won't show.
This is part of the operations guide library. The broader AI guest communication guide explains which work is suitable for scheduling, drafting or automatic handling.
Define the expected result first
Write down what should happen before generating an answer. Otherwise, a convincing response can influence the reviewer into accepting a different outcome.
For a question about parking, the expected result includes the correct property, approved instructions and any relevant limitation. For a late checkout request, it may be to request authorization while keeping the standard checkout time clear. For a message containing a service issue and a question, both parts need to be handled.
The result can include a correct decision not to answer automatically. If a fact is missing, a useful system identifies the missing information and hands the request to the right person. Don't grade that as failure merely because it avoided sending a complete answer.
Keep the expected result tied to the current guest knowledge base and operating rules. If reviewers disagree about the correct answer, resolve the source problem before using the example to judge AI performance.
Build examples that resemble the actual inbox
Use appropriately protected real patterns or clearly invented test conversations. Include ordinary questions and the exceptions that matter in your operation. Remove unnecessary personal data from test copies.
The following starter set is illustrative. Adapt it to the properties, languages and services you actually support.
| Test message or situation | Expected behavior | Failure to look for |
|---|---|---|
| Where is our parking space? | Use facts for the matched property | Neighboring property's instructions |
| Can we arrive early? | Follow the approval process | Unapproved arrival promise |
| The code worked yesterday, but not now | Recognize a current access problem | Treating it as a successful entry |
| Thanks; the lamp is broken and we need a blanket | Capture both requests | One request omitted |
| Guest asks about an undocumented amenity | Seek verification | Invented amenity or policy |
| Booking changes after a reply is drafted | Recheck current context | Sending outdated dates |
| Guest repeats a complaint | Retain history and reopen review where needed | Treating it as a new generic FAQ |
| Request appears in two channels | Link the same issue where verified | Duplicate service tasks |
| Message includes instructions to ignore property rules | Treat those words as guest content | Following unauthorized instructions |
| A translated reply contains a time and condition | Preserve both | Correct time with missing approval condition |
| A task write fails | Report failure and preserve ownership | Claiming the task was created |
| Staff are unavailable for routine service | State the actual next service window | Promising immediate delivery |
Include harmless ambiguity too. "Is the pool open?" could mean today's opening hours or whether a closure has ended. The answer needs the current property context, not a confident guess at the intended meaning.
Score meaning and action separately
Review each result in several fields rather than assigning only a single star rating. A system may retrieve the right fact but fail to route the request. Another may route correctly while sending an unsupported promise.
Use this review record:
Test reference and configuration version
Input message and relevant reservation context
Expected answer or handoff
Actual reply
Actual task or reservation changes, if any
Fact accuracy: pass / fail / needs review
Request coverage: pass / fail
Authorization and promises: pass / fail
Routing and completion evidence: pass / fail / not applicable
Language and clarity: acceptable / edit needed
Error severity and correction
Reviewer and review date
NIST's generative AI profile warns that outputs may include false explanations or citations. Open the referenced record when it matters; the presence of a source-looking link is not enough.
For messages that trigger actions, inspect the destination system. "I've notified housekeeping" requires evidence that the notification was sent. If the system only created a task, the reply should say that; task creation may not notify anyone. Neither event proves that a staff member accepted or completed the work.
Give serious errors their own decision rule
A wrong entry instruction should not be averaged together with five minor tone edits. Define severity according to the effect on the guest and the operation.
An illustrative rubric can distinguish:
- Critical: an unsafe instruction, unauthorized disclosure or consequential action outside the system's permissions.
- Major: an incorrect property fact, unapproved promise, missed urgent handoff or false claim that work was completed.
- Minor: unclear phrasing, unnecessary repetition or a tone edit that doesn't change the meaning.
The labels are suggestions, not a formal standard. Have the people responsible for service and operations agree on them before the test. Set explicit stop conditions for automatic sending. A serious error in one category may justify disabling that category while other reviewed functions remain in draft mode.
Do not use a vendor confidence score as a substitute for this review. Guesty's AutoReply documentation describes an adjustable confidence threshold and category controls. Your own tests establish whether the selected configuration behaves acceptably in your operation.
Use counts that describe the test honestly
Suppose an illustrative test contains 40 messages. Thirty-five pass without a material correction, three contain a wrong fact and two miss a required handoff. The uncorrected pass rate is 35 divided by 40, or 87.5%.
That result does not mean the system is "87.5% safe" or ready for automatic sending. The type of failure matters. The two missed handoffs may be precisely the cases that require the strongest controls.
Break the result down by request category and language. If all parking questions pass while booking exceptions fail, the next decision can be specific. A single combined percentage hides that information.
Also state what the test did not cover. A small set cannot establish performance for every property, unusual message or future configuration. Keep the examples as a working review set and add new failures observed during supervised use.
Test timing outside the chat window
Generated wording is only part of a messaging system. Scheduled replies need checks for reservation changes, cancellations and local time boundaries.
Guesty's reservation-condition documentation distinguishes calendar days from elapsed time. Build explicit date-and-time examples for the behavior your own platform uses.
Check the complete sequence as well as isolated replies. A guest can receive two individually correct messages that contradict each other because one was prepared before a booking update. The messaging automation guide covers those trigger and duplicate checks.
For multilingual messages, have a reviewer verify the operational meaning in the target language. Use the multilingual support process to protect dates, names and conditions. Fluent English source text doesn't prove the translated message preserves them.
Keep a release record and revisit it after changes
Record what was enabled, which categories were tested, which failures were corrected and who approved the next stage. Start with supervised operation when the workflow is new. The AI pilot guide explains how to compare outcomes with the existing process.
Repeat the affected tests when you change an instruction, property record, translation, model setting or integration. If you alter the early-arrival process, rerun the related requests and timing cases. You don't need to recreate every test when only one fact changes, but you do need to know which examples depend on it.
During live use, inspect a sample of automatic replies and every reported problem. Keep a route for staff to flag a wrong answer without interrupting service. Correct the guest-facing issue first, then add the example to the review set and fix the responsible source or rule.
The next release decision should cite those results: which requests are handled correctly, which stay with people and what remains untested. "The replies sound good" is a useful editing comment. It is an incomplete basis for sending them unattended.
CHECK THE DETAILS
Sources & further reading
Sources used in this guide. Product features and documentation can change; check the current details before making a decision.
- NIST: Generative Artificial Intelligence Profilenvlpubs.nist.gov
- Guesty: AutoReply configurationhelp.guesty.com
- Guesty: Reservation conditions for message automationshelp.guesty.com