All three policies came back yes, each with a high probability — even though nobody in the thread used the words “delivery delay” or “damaged item.” include_evidence was on, so each match also cites the message IDs that support it: late arrival in m1, cracked screen in m3, refund in m5. reference_id comes back unchanged — it is just your own label for the call.
A support tool can now tag the ticket, start a return, and skip a person reading the thread. Across a week of chats, the same probabilities average into a rate. That is the useful part — one conversation becomes a decision, and many conversations become a number you can watch.
These policies are not on the measured roster, so they are labeled pooled_unseen. The number is still a probability. The tier just tells you which scorecard produced it.