In this illustrative response, all three policies have a high probability — even though nobody in the thread used the words “delivery delay” or “damaged item.” Full mode always returns evidence, so each result also lists the messages that support it: late arrival in m1, cracked screen in m3 and m4, refund in m5. No contradicting message is selected in this example, so every contradicting_message_ids list is empty.
A support tool can now tag the ticket, start a return, and skip a person reading the thread. Across a week of chats, averaging comparable probabilities gives an estimated rate. Check calibration on that traffic first — a stable model and policy definition make changes easier to interpret.
None of these policies was registered anywhere. v2 has no roster: a definition written a minute ago is scored exactly like every other. There is no verdict in the response either — the probability is advice, and the threshold for acting on it is yours.