I am the founder of Operant Labs, so I have a bias. This is a research run, not a product feature. I want critique of the method.
The task: a new agent conversation starts. Match its first user message to one of 4 known task clusters, or to “none”. We have 92 cases from one demo workload.
Run 1 used text-embedding-3-small. The cosine floor was 0.5 for this run. It compared the message with the centroid of the goal summaries in each cluster. It found the right cluster for 21.1% of the cluster members.
Run 2 used all-MiniLM-L6-v2 and bge-small-en-v1.5 with sentence-transformers. It compared the message with the text “label: description” of each cluster. For each conversation, we tuned the threshold on the other 53 conversations. The two models found 92.1% and 86.8% of the cluster members. On all 92 cases, they got 92.4% and 90.2% right.
We changed two things at the same time: the model and the text. So this run does not show which change matters more.
My questions:
Next, we want to run text-embedding-3-small on “label: description”. We also want to run the two small models on the centroids. Is that the correct next step?
Each message got the query instruction of bge-small-en-v1.5 in front of it. Is that the right use when the other side is a short label and description?
The cluster descriptions came from the same conversations. How do you make a fair set of cases for this task?
The full table and the method: https://operantlabs.com/results/open-models-cold-start/