The text we compare with the first message changed the cluster match from 21% to 92%. How do we find which change mattered? Operant Labs founder reports that swapping both the embedding model and the comparison text raised first-message task-cluster matching from 21.1% to 92.1% on 92 demo-workload cases, so the run cannot isolate which change drove the gain. Run 1 used text-embedding-3-small with a 0.5 cosine floor against cluster goal-summary centroids; Run 2 used all-MiniLM-L6-v2 and bge-small-en-v1.5 via sentence-transformers against each cluster's "label: description" text, with thresholds tuned on the other 53 conversations, reaching 92.1% and 86.8% of cluster members and 92.4% and 90.2% across all 92 cases. The founder asks for critique of the method and plans to cross the two variables by running text-embedding-3-small on "label: description" and the two small models on the centroids. I am the founder of Operant Labs, so I have a bias. This is a research run, not a product feature. I want critique of the method. The task: a new agent conversation starts. Match its first user message to one of 4 known task clusters, or to “none”. We have 92 cases from one demo workload. Run 1 used text-embedding-3-small. The cosine floor was 0.5 for this run. It compared the message with the centroid of the goal summaries in each cluster. It found the right cluster for 21.1% of the cluster members. Run 2 used all-MiniLM-L6-v2 and bge-small-en-v1.5 with sentence-transformers. It compared the message with the text “label: description” of each cluster. For each conversation, we tuned the threshold on the other 53 conversations. The two models found 92.1% and 86.8% of the cluster members. On all 92 cases, they got 92.4% and 90.2% right. We changed two things at the same time: the model and the text. So this run does not show which change matters more. My questions: Next, we want to run text-embedding-3-small on “label: description”. We also want to run the two small models on the centroids. Is that the correct next step? Each message got the query instruction of bge-small-en-v1.5 in front of it. Is that the right use when the other side is a short label and description? The cluster descriptions came from the same conversations. How do you make a fair set of cases for this task? The full table and the method: https://operantlabs.com/results/open-models-cold-start/ https://operantlabs.com/results/open-models-cold-start/