# The text we compare with the first message changed the cluster match from 21% to 92%. How do we find which change mattered?

> Source: <https://discuss.huggingface.co/t/the-text-we-compare-with-the-first-message-changed-the-cluster-match-from-21-to-92-how-do-we-find-which-change-mattered/190123#post_1>
> Published: 2026-10-08 15:21:23+00:00

I am the founder of Operant Labs, so I have a bias. This is a research run, not a product feature. I want critique of the method.

The task: a new agent conversation starts. Match its first user message to one of 4 known task clusters, or to “none”. We have 92 cases from one demo workload.

Run 1 used text-embedding-3-small. The cosine floor was 0.5 for this run. It compared the message with the centroid of the goal summaries in each cluster. It found the right cluster for 21.1% of the cluster members.

Run 2 used all-MiniLM-L6-v2 and bge-small-en-v1.5 with sentence-transformers. It compared the message with the text “label: description” of each cluster. For each conversation, we tuned the threshold on the other 53 conversations. The two models found 92.1% and 86.8% of the cluster members. On all 92 cases, they got 92.4% and 90.2% right.

We changed two things at the same time: the model and the text. So this run does not show which change matters more.

My questions:

Next, we want to run text-embedding-3-small on “label: description”. We also want to run the two small models on the centroids. Is that the correct next step?

Each message got the query instruction of bge-small-en-v1.5 in front of it. Is that the right use when the other side is a short label and description?

The cluster descriptions came from the same conversations. How do you make a fair set of cases for this task?

The full table and the method: [https://operantlabs.com/results/open-models-cold-start/](https://operantlabs.com/results/open-models-cold-start/)
