{"slug": "reverse-engineering-how-meta-s-muse-shops", "title": "Reverse Engineering How Meta's Muse Shops", "summary": "Independent researcher Kalan Peace reverse-engineered Meta's Muse AI shopping agent and found the product ranking is driven by a `ranking_score` field attached to each product, with the catalog search handled by an 11.7 MB compiled Rust program at `/opt/hatch/bin/meta-catalog-search`. Muse, Meta's AI agent that gives each user a cloud Linux computer, passed 2.8 million installs worldwide in its first 12 days after launching in early September 2026, with about 642,000 daily U.S. users, according to Apptopia. Peace also reported that raw catalog results returned internal debug links to Meta's internal search system, which sit behind a Meta employee sign-in and showed only an \"Internal Login\" page when one link was opened on 28 September.", "body_md": "[Back to Blog](https://caeliai.com/blog)\n\n# Reverse Engineering How Muse Shops\n\n*Research orchestration: Astra (gpt-6-astra). Independent checks: Luna (gpt-6-luna, max). Reviews: Claude and DeepSeek. Editing: Claude (claude-opus-5-5). Evidence, code and data: [github.com/kalanpeace/reverse-engineering-muse](https://github.com/kalanpeace/reverse-engineering-muse).*\n\n**A note for anyone at Meta.** The raw catalog results Muse gave me include internal debug links to Meta's own search system ([Part 1, Step 6](#step-6--looking-over-the-wall-unicorn)). They all go to Meta's internal network, which sits behind a Meta employee sign-in: the one link that was opened, on 28 September, showed only Meta's \"Internal Login\" page. I don't have a Meta login, so all I have are the addresses, not what's behind them. In the public data, the request handle in each link is blanked, and the originals are kept privately. If anyone working at Meta would like to see more of what Muse gave me, contact me at [\\[email protected\\]](https://caeliai.com/cdn-cgi/l/email-protection#9ef5fff2fff0eefbfffdfbdef9f3fff7f2b0fdf1f3). I don't think this qualifies for a bug bounty, but I wanted to flag it.\n\n### What Muse is\n\nMuse is Meta's new AI agent. It doesn't just answer questions like a chatbot. Meta gives each user's Muse its own Linux computer in the cloud, with a web browser, a place to save files and a set of tools. Muse can run commands, open websites, write files and send smaller helper agents off to do parts of a job. You talk to it through Meta's mobile and web apps.\n\nMuse is also one of the fastest-growing AI apps around. It launched in early September 2026, only about three weeks before this report. By Apptopia's count, it passed 2.8 million installs worldwide in its first 12 days, with about 642,000 daily users in the U.S. That's faster than ChatGPT's early mobile launch. ([TechCrunch](https://techcrunch.com/2026/09/21/metas-muse-is-outpacing-chatgpts-early-mobile-launch/))\n\nA few facts about how it's built, from Meta's own description ([Meta](https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse), published September 8, 2026):\n\n- Each user's Muse runs in its own isolated section of a dedicated virtual machine.\n- Muse has a browser, a working folder of files, command-line tools, skills and helper agents.\n- Being in charge of your own Muse computer is not the same as controlling Meta's servers. Built-in connectors run in separate, protected \"workers.\"\n\nOne of the things Muse can do is **shop**. Ask it for a frying pan or a dress, and it searches for products, checks their pages and shows you a few as product cards.\n\n### What I did\n\nI wanted to know how Muse decides which products to show. So I took it apart from the outside.\n\n- **I ran hundreds of catalog searches through Muse** and saved exactly what came back. Every file got a SHA-256 fingerprint, so any change to it would show.\n- **I found the file names behind Muse's shopping.** They include the catalog program`/opt/hatch/bin/meta-catalog-search` , the protected worker it talks to at`/run/hatch/privsep/meta-catalog-search.sock` , the shopping instructions Muse follows (`shopping-SKILL.md` ) and the`ranking_score` attached to each product.\n- **I got a copy of the catalog program itself.** It's an 11.7 MB compiled Rust program. I wrote my own decoder to read its machine code without ever running it.\n- **I ran controlled experiments.** I changed one thing at a time, like the wording, a brand filter or the number of products requested, and repeated each test three times.\n- **I followed full shopping tasks** from the question, to the catalog, to the page checks, to the cards on screen.\n\n### What I found\n\n**A simple way to understand how Muse shops:**\n\n- **The shopper asks a question.**  - Muse turns it into a search of what I'll call **the Muse catalog** , the product catalog behind`meta-catalog-search` .\n- Muse turns it into a search of what I'll call \n- **The Muse catalog returns a list.**  - In my shopping tasks, that was 40 to 71 products, each with a score.\n- **Muse picks a few.**  - Usually it picked 3 to check.\n- **A browser checks those products on the stores' own websites.**  - It confirms price, stock and the exact variant.\n- **Muse shows the shopper a few cards: 2 or 3 in the tasks I traced.**\n\nWhat stood out:\n\n- **The catalog is the backbone.** In every shopping task I traced, Muse searched the catalog first and then checked products in a browser.\n- **Muse says it doesn't use Google's rankings.** In its words: \"discovery ranking isn't Google ranking … I'm not pulling a Google SERP and reading off positions.\" Its browser goes straight to store websites. It does use web search for context, like fashion trends, and once reported \"a Google Shopping cross-check\" to confirm a retailer.\n- **Meta rates every seller, and the AI never sees it.** Each product arrives with a hidden`seller_quality` label, usually`elite` ,`good` ,`acceptable` or`poor` . It comes from Meta's servers, isn't named in the catalog program on Muse's computer, and is left out of what the AI reads when it shops.\n- **The same product scores differently depending on the store.** Two coffee shops selling the same Moccamaster grinder, with the same seller rating, scored 0.551 and 0.711.\n- **Muse already has brands in mind.** When the catalog went down, it named brands straight from memory (\"The honest verdict from everything I know: a Vitamix…\") and sent its browser directly to those brands' sites. At first I thought the outage was a hallucination. It wasn't: three health checks confirmed it.\n- **The score doesn't decide the order.** Products with higher scores often sit below products with lower scores.\n- **Being returned isn't being shown.** Products at spots #30, #35 and #65 became cards 1, 1 and 2. A #3 product with a broken page was never shown.\n- **Taste words swap out the results.** Adding \"luxury\" or \"timeless\" to a dress search replaced every product in it. Repeating the same search kept about 91%.\n- **Muse's sense of \"good brands\" comes from its training.** It said so itself: \"Honestly? Recognition … it's a rich-get-richer loop.\"\n- **The program I got is not the ranking system.** It's the program on Muse's computer that asks Meta's servers for results. It sits one step before the ranking.\n- **That's where I hit the wall.** The formula behind the score lives on Meta's servers. I couldn't reach it, and neither can anyone outside Meta.\n- **But I could see over it.** A hidden debug field shows the first step behind the wall: Meta's Unicorn search pulls about 100 candidates, and a second step keeps about half and reorders them.\n\n*One real saved search, \"stainless steel frying pan\" on 27 September, product by product. Part 1, Step 6 explains the three lists.*\n\n**The rest of this report is highly technical: raw output, math and code, explained in full.**\n\n**Own a store? [Book a free call →](https://cal.com/caeliai/free-analysis)** We'll see if you qualify for a two-week sprint. I apply what I learned here to your products: I find where they drop out of Muse's results, fix what's fixable in your listings and measure again, so you can see what actually changed.\n\nNot everyone will qualify, and no one can honestly promise a ranking. What I can promise is that you'll know exactly where you stand.\n\n## Technical report\n\n### The short version\n\n**The scores come pre-made.** Every `ranking_score` I saw arrived already calculated. It comes from upstream, on Meta's servers, which aren't open to the public. That makes sense, and it's normal for any company's ranking system.\n\nThe only part I could fully see was the catalog program on Muse's computer, `meta-catalog-search`. I got the whole program and decoded it, but it doesn't rank anything. It sends the search off, receives the results and interprets what it gets back: it reads the products, numbers them and prints them. Everything above it, where the actual ranking happens, is server-side. The formula never showed up in anything I could reach, but a debug field did show the pool it picks from (Part 1, Step 6).\n\nSo this report maps everything *around* the ranking: what goes in, what comes out, what Muse does with it and what changes the results. It doesn't reveal the formula itself.\n\n### Where these files came from\n\nThese aren't answers I got by asking a chatbot questions. Here's how the files for the controlled experiments were made:\n\n- **Script first.** I wrote each test as a script ahead of time: the exact searches, the order and how many repeats.\n- **Muse ran it on its own computer.** It saved the raw output of every search to a file, along with a receipt showing the exact command, start and end times and exit code.\n- **Muse packed and handed over the files.** It bundled the files into ZIP archives with a manifest listing every file and its SHA-256 fingerprint.\n- **I checked every file.** I downloaded the archives and checked every file against the manifest and every archive for corruption. I also confirmed that the script Muse ran was byte-for-byte the one I wrote.\n\nFor example, all 60 searches in the fashion study's main schedule passed every check: receipt, command, schedule, script, output fingerprint and a 183-file manifest.\n\nYou can watch this happen in the full chat transcript, where Muse creates, packages and hands over each batch of files.\n\nWhere a record was written down by hand instead of saved automatically, like some browser reports, I say so.\n\n**What could still be wrong.** I have no inside information from Meta, and nobody at Meta confirmed any of this. The program copy came from Muse, so I can't independently prove it's the exact program Meta runs, and later tests reported different versions. My decoder could also misread something. That's why the main findings in Parts 2 to 4 rest on saved search results, which don't depend on the decoder at all.\n\n### Follow along\n\nEverything behind this report is public on GitHub: **[github.com/kalanpeace/reverse-engineering-muse](https://github.com/kalanpeace/reverse-engineering-muse)**. It includes:\n\n- **[The full chat transcript with Muse](https://github.com/kalanpeace/reverse-engineering-muse/tree/main/transcript),** word for word, so anyone can read how I got to each point. Every quote in this report comes from it or from a saved file.\n- **[Every file Muse gave me](https://github.com/kalanpeace/reverse-engineering-muse/tree/main/evidence):** the raw search outputs, receipts, manifests and archives.\n- **[My decoder](https://github.com/kalanpeace/reverse-engineering-muse/tree/main/code)** and the analysis and chart code.\n- **[The data and fingerprints](https://github.com/kalanpeace/reverse-engineering-muse/tree/main/data)** behind every number. Every run ID (like`neutral_baseline/ncap-n50-run2` ) and product ID (like`25671310895791560` ) can be looked up there.\n\n**A note about the chat.** Some of the messages on my side were typed by Astra, my lead AI agent, using computer use to operate the Muse chat on my behalf. 23 messages are signed \"Astra here,\" and 3 more open with \"Astra.\" Others may not be labeled. Astra followed the research plan I set, and it was the only agent allowed to talk to Muse.\n\n### How sure I am\n\n- **Very sure:** the numbers from saved search results. That covers scores, spots, counts, overlaps, and which products were returned and shown. These are counted directly from files with fingerprints.\n- **Fairly sure:** the reading of the catalog program. The bytes check out, but reading machine code is interpretation, and the copy could differ from what Meta runs.\n- **Muse's word only:** Muse's descriptions of itself, like \"discovery ranking isn't Google ranking,\" \"recognition\" and \"chosen server-side.\" They're quoted exactly, but they're claims, not measurements. When I pushed Muse for evidence, it sometimes walked a claim back. For example, it admitted its \"verbatim server body\" claim rested on a search it never saved. Those corrections are in the transcript too.\n- **Hypotheses:** anything labeled that way, like meaning-based matching or a recognition loop. They fit the evidence, but they aren't proven.\n\n### How this report works\n\n- **Part 1:** how I got into the Muse catalog, what came out, what the program looks like inside, and what can be seen over the wall.\n- **Part 2:** what the catalog does.\n- **Part 3:** what Muse does with the catalog's results, followed by a summary of what we know and don't know.\n- **Part 4:** taste. This is the more experimental part, about how words like \"fashionable\" or \"luxury\" change what Muse finds.\n- **Part 5:** my hypothesis about how it all fits together, and how to test it.\n\nEvery finding is laid out the same way:\n\n- **What Muse output:** the exact command and the exact result.\n- **What Muse said:** Muse's own words about it, or its own instructions.\n- **What it means:** my interpretation, in plain words.\n- **What it doesn't prove:** where the evidence stops.\n\nThe experiments were separate, and I never add their results together:\n\n| Experiment | Searches | Products returned | \n|---|---|---|\n| Early catalog runs | 182 | 6,314 | \n| Combined-search, color and brand tests | 27 | 1,200 | \n| Second combined-search tests | 21 | 1,104 | \n| Broad store panel (20 product types × 5 question types) | 100 | 4,088 (3,827 different products*) | \n| Fashion wording study | 72 | 3,827 (1,125 different products) | \n| Home-goods shopping tasks | 5 | 270 → 14 cards shown | \n| Fashion shopping tasks | 7 (in 6 tasks) | 375 → 15 cards shown | \n| Raw-tags test (12 categories + 4 products, 3 repeats each) | 48 | 2,542 | \n\n*The two 3,827s are a coincidence: one counts different products in the panel, the other counts all product records in the fashion study.\n\n## Part 1 · Getting into the Muse catalog\n\n### Step 1 · Finding the tool\n\n**What Muse output.** Muse's own computer answered two basic questions:\n\n```\nfollowup60/R58   uname -srm                                  → Linux, x86-64\nfollowup60/R59   /opt/hatch/bin/meta-catalog-search --help   → the tool's own manual\n```\n\nMuse pasted the tool's manual (its `--help` text) into our chat word for word. It starts:\n\n```\nMeta 1P product catalog search.\n\nUsage: meta-catalog-search [OPTIONS]\n\n  -q, --query <QUERIES>        Semantic query (repeatable)\n  -n, --num-results            Number of results to return (default: 10)\n      --category               Category constraint (repeatable). Products MUST match\n      --brand                  Brand OR seller/retailer to strongly prefer … resolves and ranks\n                               that brand/retailer's own store first, backfilling with other sellers below\n      --domain                 Website domain to strongly prefer … Matches the product URL domain ONLY\n      --color / --material / --style / --prefer-brand   … preference boost\n      --seller-type            [default: direct] [possible values: direct, secondhand]\n      --user-id                Numeric user id. If omitted, no user_id is sent\n      --raw                    Print raw HTTP response body\n```\n\n*Shortened here. The full text and line numbers are in the [chat excerpts](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/transcript/MUSE-CHAT-EXCERPTS.md).*\n\n**What Muse said.** Muse's shopping instructions (`shopping-SKILL.md`) tell it how to shop:\n\n\"Use `browser.search` to discover trends, well-known sellers for a product category, reviews, or typical prices.\"\n\n\"Call `browser.open` on all non-Marketplace product URLs…\"\n\n**What it means.** The manual tells you a lot:\n\n- **The search is \"semantic.\"** It matches on meaning, not just exact words. That fits what I found about taste in Part 4.\n- **You can nudge the results but not sort them.** Some options are hard filters, like`--category` and price. Others are soft \"boosts,\" like color, style and preferred brand. None of them says \"sort by price\" or \"sort by score.\" The order comes back from Meta's servers already decided.\n- **The brand and domain options say they \"strongly prefer.\"** They don't say \"only.\" That's exactly what I measured in Part 2: the brand's products go first, and others fill in below.\n- **There's a `--user-id` option.** My controlled tests didn't send one. Muse said that in its own quick tries, results were \"insensitive to user-id and conversation-id.\" Personalization still hasn't been properly tested.\n\n**What it doesn't prove.** A manual describes what the tool is *supposed* to do. For example, it says `--raw` prints the \"raw HTTP response body,\" but my own reading of the code (Step 4) couldn't confirm the output is untouched. And `-n` says \"number of results,\" but it doesn't cap them (Part 2).\n\n### Step 2 · Getting the scores out\n\n**What Muse output.** When Muse shops normally, it saves results with `--out` to a file called `catalog-results.json`. Those files give each product a `rank` number (1, 2, 3 …) and a description, but **no score**. Three saved normal printed outputs, with 107 products, look like that. Seven early `--out` attempts didn't produce the file at all; one of them still printed its results.\n\nWith `--raw`, a score appears on every product. This is the first scored run I saved, `neutral_baseline/ncap-n50-run2`, from 26 September 2026 at 22:49 UTC:\n\n```\n/opt/hatch/bin/meta-catalog-search --query \"stainless steel frying pan\" -n 50\n    --conversation-id phase2-ncap-2026-09-26 --raw --no-save\n```\n\nIt returned 43 products. Here's one product, with its fields rebuilt from that file:\n\n```\n{\"position\": 41, \"product_id\": \"25671310895791560\",\n \"name\": \"5 Ply Stainless Steel Nonstick Cookware: 2 Piece Frying Pan Set: 8\\\" & 10\\\"\",\n \"brand\": \"Quince\", \"url_host\": \"www.quince.com\",\n \"price\": \"$288\", \"sale_price\": \"$144.90\", \"ranking_score\": \"0.773438\"}\n```\n\n**What Muse said.** In a saved reply, Muse said:\n\n\"The exact remote ranking formula was not recovered and remains unknown. Client ordinal rank (1-based parse order, struct offset +0x168, observed in disassembly) is separate from server ranking_score.\"\n\n**What it means.** There are two different numbers. `rank` is just a count, 1, 2, 3. `ranking_score` is a number that comes from Meta's servers. Every raw experiment in this report uses `ranking_score`.\n\n**What it doesn't prove.** \"Raw\" is only the name of the option. It doesn't prove this is exactly what Meta's server sent.\n\nAt one point, Muse claimed the raw output *was* the untouched server response. I pushed back and asked for its evidence. It checked its own files and conceded:\n\n\"Nothing in the saved evidence independently establishes that the printed bytes are the untouched HTTP body.\"\n\nThe byte search behind its earlier claim had never been saved. So the fair wording, which Muse and I now both use, is this: raw output contains the product XML the server supplied, and whether it's byte-for-byte identical to what the server sent is unverified.\n\n### Step 3 · Getting the program\n\n**What Muse output.** First, Muse handed over four small pieces of the program's machine code. Later, with my approval for that exact file, it exported the whole program, about 11.7 MB. The four pieces match the full program exactly.\n\n**What Muse said.** Muse's reply in Step 2 names the same things I found in the code: the `rank` number, which it calls \"1-based parse order,\" and where the program stores it.\n\n**What it means.** I had real code to read, not just Muse's description of it.\n\n**It's written in Rust.** Muse found these inside the program and the files next to it:\n\n```\n/rustc/8bab26f4f68e0e26f0bb7960be334d5b520ea452/library/std/src/time.rs     ← Rust compiler path\ntokio, reqwest                                                             ← Rust libraries for networking\njarvis/…/hatch-engine/crates/hatch-networking/src/telemetry_proxy/access_token.rs\ntools/shopping_results.rs · shopping_image_availability.rs · domain_dispatch/shopping_steps.rs\n```\n\nSo the tools and plumbing around Muse's AI (the \"Hatch\" engine) are built in Rust, a fast, low-level systems language. That tells you this is serious infrastructure engineering, not a quick script bolted onto a chatbot. It says nothing about which AI model Muse runs, or what language Meta's ranking server uses. Those are unknown. (Meta's open Llama models are usually run from Python/PyTorch code, and the popular C++ version, `llama.cpp`, is a community project. Nothing here shows Muse runs Llama at all.)\n\n**What it doesn't prove.** Tests I ran later reported different versions of the program. So I can't prove this exact copy is the one that ran in every search.\n\n### Step 4 · Reading the program with my decoder\n\n**What Muse output.** Compiled programs are machine code: bytes that the computer runs directly. My decoder turns those bytes into readable instructions without running them. Here are the first lines of the `caller-0x4bb230` piece:\n\n```\naddress  bytes                   instruction\n4bb230:  41 57                   push  r15\n4bb232:  41 56                   push  r14\n4bb234:  41 55                   push  r13\n4bb236:  41 54                   push  r12\n4bb238:  53                      push  rbx\n4bb239:  48 81 ec 80 00 00 00    sub   rsp, 0x80\n```\n\nThe check compared every byte against the saved pieces and found zero differences. The pieces and the decoded output are in [`evidence/program/decoded-pieces/`](https://github.com/kalanpeace/reverse-engineering-muse/tree/main/evidence/program/decoded-pieces), and the decoder script is [`code/decode_saved_windows.py`](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/code/decode_saved_windows.py).\n\nInside the full program, I also found these text labels: `run_trusted: starting` and `run_trusted: fd mappings built, executing tool`.\n\n**What Muse said.** Muse's reply in Step 2 calls the `rank` number \"1-based parse order.\" That matches what the code does.\n\n**What it means.** Rewritten as plain steps, the numbering code does this. It's my reconstruction, not Meta's source code:\n\n```\nordinal = 1\nfor product in accepted_records_in_input_order:\n    product.rank = ordinal\n    ordinal = ordinal + 1\n```\n\nThe program gives the first product it gets the number 1, the next one 2, and so on. It doesn't decide the order.\n\nI also followed the `--raw` option. It prints a text string the program saved earlier instead of its processed product list:\n\n```\nif raw:\n    write_stdout(previously_constructed_response_string)\n```\n\nAnd I traced a chain of calls from the `run_trusted` worker code to the part that handles catalog responses. ([Address-level analysis](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/STATIC-DISPATCH-2026-09-27.md))\n\n**So does the program show how the ranking works? No.** In the parts I read, the program sends the search off, takes the products back in the order they arrive and numbers them 1, 2, 3. It doesn't score or sort them. As far as I can tell, the ranking happens before that, on Meta's side.\n\n**What it doesn't prove.** I read the key parts of the program, not every line, so I can't prove it never reorders anything. And reading machine code isn't the same as having Meta's original source.\n\n### Step 5 · Hitting the wall\n\n**What Muse output.** Two short recordings of the program running ([runtime traces](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/TWO-ROUTES-2026-09-27.md)) show it connecting to a protected worker through this file:\n\n```\n/run/hatch/privsep/meta-catalog-search.sock\n```\n\nNeither recording contains any product data or scoring. When I asked Muse which tools it could load, it said there was no way to list them and no separate catalog or scoring tool. ([Capability boundary](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/CAPABILITY-BOUNDARY-2026-09-27.md))\n\n**What Muse said.** Muse put it plainly in our chat:\n\n\"**The 50 are chosen server-side.** No code on this VM decides which products come back — the retrieval, the relevance model, the cutoff, all of it lives inside Meta's backend, which I can't see.\"\n\n\"The local code is plumbing. It packs args, hands off, carries order. It never picks products. The pool gets picked above the red line, where I can't see.\"\n\n**What it means.** The chain goes like this:\n\n1. The program on Muse's computer passes the request to a protected worker.\n2. The worker talks to Meta's servers.\n3. The scores and the first order are decided there.\n\nThat's the wall. The formula lives on Meta's servers, where neither I nor anyone else outside Meta can reach it.\n\n**What it doesn't prove.** This doesn't prove every possible method has been tried. It means I couldn't find a way past this point without Meta's cooperation.\n\n### Step 6 · Looking over the wall: Unicorn\n\n**What Muse output.** The raw output has one more hidden field, `query_debug_info`. It isn't a product. It's a note about the search itself: a link, a count and a list of product IDs. Here it is from the first raw-tags search, `A-R1-Q01` (\"stainless steel frying pan\"), shortened, with the request handle blanked:\n\n```\n\"query_debug_info\": [{\n  \"query\": \"stainless steel frying pan\",\n  \"uqc_link\": \"https://www.internalfb.com/unicorn/query?tier=unicorn.aggregator.alacorn-genai-synapse-products.211.denmark.s00.r01-prod&baseRequestHandle=REDACTED\",\n  \"product_count\": 100,\n  \"product_ids\": [\"9694411407239051\", \"5783497841778259\", …],\n  \"product_handles\": {…}\n}]\n```\n\nI counted every one of these fields across all 447 raw responses saved in the evidence:\n\n- **Every link goes to the same locked place.** All 474 links point to`internalfb.com/unicorn/query` , on Meta's internal network, behind a Meta employee sign-in. The one link that was opened, from a later follow-up search on 28 September that isn't among these 447 responses, showed only Meta's \"Internal Login\" page. From outside Meta, all you can see is the address.\n- **One system, three sites.** Every tier name follows the same pattern:`unicorn.aggregator.alacorn-genai-synapse-products.211.<site>.s00.r<number>-prod` . Only the site and the last number change. The sites are`gallatin` (230 links),`denmark` (171) and`crookcounty` (73).\n- **The list is the pool.** 432 of the 474 lists name exactly 100 products.\n- **Everything the catalog returned came from that list.** In the 417 single searches that returned products, every returned product was on the list. Not one came from outside it.\n- **There's a list in between.** The response's`flywheel_identifiers` field (Finding 7) holds a third list. In all 417 of those searches, every returned product was on it, and everything on it was on Unicorn's list. It never held more than 80 products, and held exactly 80 in 228 of the 375 searches where Unicorn listed 100. Its order matched Unicorn's in only 4 searches, and started with the returned products in only 13.\n- **About half survives.** When the list named 100, the catalog returned between 1 and 75 of them, 47 in the middle case.\n- **The order changes.** In 413 of the 417 searches, the returned order was different from the list's order. On a scale where 1 means the same order and −1 means reversed, the typical agreement was 0.16.\n- **Being early on the list doesn't help.** I split each list of 100 into blocks of ten. Every block kept between 43% and 48% of its products. The first ten did no better than the last ten.\n- **Same 100, different result.** In the raw-tags run, all 48 searches went to the`denmark` site. In three searches (men's running shoes, women's dresses and moisturizer for dry skin), all three repeats got the identical list of 100, yet the returned products still changed.\n- **Quince was on the list too.** In the Finding 1 search, the Quince pan that came back at spot 41 was 78th on the list of 100.\n- **Empty isn't empty.** 9 searches returned no products at all but still carried a list.\n\nThen I used the list to find out which settings reach Unicorn and which only act after it. Each comparison below changes one setting and keeps everything else the same. \"Shared\" is how many of Unicorn's 100 candidates the two versions have in common. Every search in the table, with and without the change, got the identical 100 in every repeat (3 to 21 runs each), so the differences come from the setting, not from drift.\n\n| Change | Unicorn candidates shared (of 100) | \n|---|---|\n| `--brand \"Hydro Flask\"` on the water bottle | 5 | \n| `--brand Misen` on the frying pan | 7 | \n| \"Misen\" typed into the frying-pan query instead | 37 | \n| `--color green` on the water bottle | 90 | \n| `--prefer-brand \"Hydro Flask\"` on the water bottle | 97 | \n| `--prefer-brand Misen` on the frying pan | 99 | \n\n- **The number you ask for sets how deep Unicorn digs.** For`-n` of 1, 9, 20 and 50, Unicorn's list was exactly twice the request: 2, 18, 40 and 100. For ordinary requests it stops at 100, so`-n 100` still gave 100.\n- **Except, sometimes, at 10 and 11.**`-n 10` gave a list of 20 for Hydro Flask and All-Clad, but 223 for Levoit and 501 for Misen.`-n 11` gave 22 for Misen but 225 for Levoit. The \"asked 10, got 63\" search in Finding 4 is the Levoit one: Unicorn's list was 223 deep, not 20.\n- **Two `--query` flags make two Unicorn searches.** The combined search carried two lists, identical to the lists of the two single searches, and everything it returned came from those 186 candidates. The 3 products \"from neither\" in Finding 5 were on those lists all along: 1 on the frying-pan list and 2 on the skillet list. The single searches had cut them.\n- **The pool is steadier than the result.** Across 92 groups of identical settings, 63 got the identical Unicorn list every time. In 23 of those 63, the returned products still changed.\n\n**What Muse said.** When Muse unpacked this field for me, it described the link as \"internal-only, not dereferenced.\" It also found that \"no readable schema/API contract or tool description defines `ranking_score` or any `query_debug_info` field.\"\n\n**What it means.**\n\n- **Unicorn is Meta's search engine.** Facebook described it publicly in 2013 as the system behind searching its social graph ([VLDB paper](https://www.vldb.org/pvldb/vol6/p1150-curtiss.pdf) ). The tier name suggests the Muse catalog's first step runs on a Unicorn \"aggregator\" built for generative-AI product search: the name contains`genai` and`products` .\n- **The site names match Meta data center locations.** Meta runs data centers in Gallatin, Tennessee ([DCD](https://www.datacenterdynamics.com/en/news/meta-officially-launches-gallatin-data-center-campus-in-tennessee/) ), Odense, Denmark ([DCD](https://www.datacenterdynamics.com/en/news/metafacebook-to-expand-odense-data-center-campus-in-denmark/) ), and Prineville, in Crook County, Oregon ([Meta](https://datacenters.atmeta.com/wp-content/uploads/2025/02/Metas-Prineville-Data-Center.pdf) ).\n- **There are at least two steps behind the wall.** First, Unicorn pulls about 100 candidates. Then a second step, still on Meta's side, keeps about half and puts them in a new order. The`ranking_score` only appears on that second list.\n- **There may be two cuts, not one.** Something narrows Unicorn's 100 to at most 80, and something narrows those to the products Muse gets. \"Flywheel\" usually names a feedback loop, so the middle list could instead be what Meta records as considered. Either way, the returned products always came from it.\n- **This is the crack in the wall.** I still can't see the formula, but I can now see the pool it picks from. For a store owner, being in Unicorn's 100 is the first gate. The second step decides who comes back and in what order.\n- **The second step isn't just cutting off the bottom.** If it simply kept the top of Unicorn's list, the first ten would survive more often than the last ten. They don't.\n- **Some settings change who's in the room; others act after.**`--brand` and the words in the query change Unicorn's pool.`--prefer-brand` and`--color` barely touch it, so whatever they do happens in the second step.\n- **\"50\" isn't a limit on the answer.**`-n` sets how deep Unicorn digs, not how many come back; the second step decides that (Finding 4). Why`-n 10` sometimes digs 223 or 501 deep is unknown.\n- **Finding 5 has an answer.** A combined search isn't a new kind of search. It's one Unicorn search per query, pooled, and then the second step picks from the pool.\n\n**What it doesn't prove.**\n\n- The list's order might not be Unicorn's ranking. It could be sorted by something else, and the field has no documentation.\n- I can see what goes into the second step and what comes out, not what happens inside it.\n- The site names are my reading. Meta hasn't confirmed what they mean.\n- These are debug fields. Meta can change or remove them at any time.\n- What the flywheel list is for. Nothing I could check defines it.\n- The setting comparisons cover two queries (a water bottle and a frying pan), each repeated 3 to 21 times, on 26 and 27 September.\n- In the public repository, the request handle in every link is blanked. The unredacted originals are kept privately, and Meta is welcome to review them ([note for Meta](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/NOTE-FOR-META.md) ). The full analysis is in[`data/unicorn-debug-analysis.json`](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/unicorn-debug-analysis.json) and[`data/stage-analysis.json`](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/stage-analysis.json) .\n\n## Part 2 · What the catalog does\n\nThe formula stays hidden, but the catalog's behavior can still be measured from the outside. That's what the next five findings do.\n\n### Finding 1 · A higher score can sit lower on the list\n\n**What Muse output.** From the same run as Step 2, `neutral_baseline/ncap-n50-run2`, here are spots 40 and 41:\n\n| Spot | Product ID | Product | Score | \n|---|---|---|---|\n| 40 | `28631522856449032` | Hestan Brushed Clad Stainless Steel Skillets, Large (12.5-Inch). Brand field empty; site `hestanculinary.com` | 0.558594 | \n| 41 | `25671310895791560` | Quince 5 Ply Stainless Steel Nonstick Cookware: 2 Piece Frying Pan Set: 8\" & 10\" | 0.773438 | \n\n**What Muse said.** Muse called `ranking_score` a server value and said its formula \"remains unknown\" (Step 2).\n\n**What it means.** If the list were sorted by score, every product would score at least as high as the one below it:\n\nQuince breaks that rule by 0.214844:\n\nNineteen products in the same list score higher than Quince, and none tie it. So by score alone, Quince would be at spot 1 + 19 = **20**, not 41. Something besides the visible score is setting the order.\n\n**What it doesn't prove.** Spot 20 is my calculation. Muse never showed it, and it doesn't mean Quince makes the better pan.\n\n### Finding 2 · Most lists are out of score order\n\n**What Muse output.** I checked every saved list for an \"upward jump,\" meaning a product that scores higher than the one right above it:\n\n- **Early runs:** 130 of the 173 lists with 3 or more products had a jump.\n- **100-search panel:** 41 of 90 lists with 3 or more products had a jump.\n- **Fashion study:** 33 of 69 non-empty lists had a jump.\n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** Counting every pair of products in the early runs, not just neighbors, about 1 in 9 pairs sit in the opposite order from their scores.\n\nWith a brand filter, all 18 lists started with a block from that brand's website. For Levoit, the score jumps from 0.382812 to 0.808594 right where its 16-product block ends. But 12 of the 18 lists were still out of order *inside* a block. So \"brand first, then score\" doesn't explain the order either.\n\n**What it doesn't prove.** These studies used different searches, dates and program versions. So the lower later numbers don't mean Muse \"improved,\" and none of these are rates for all Muse searches.\n\n### Finding 3 · A brand filter moved a product up while its score went down\n\n**What Muse output.** Each command ran three times, with identical results each time:\n\n```\ncolor/R02-BASE-1   --query \"32 oz stainless steel insulated water bottle\" -n 50 --raw\ncolor/R03-BRAND-1  --query \"32 oz stainless steel insulated water bottle\" --brand \"Hydro Flask\" -n 50 --raw\n```\n\n| Search | Spot | Score | List size | \n|---|---|---|---|\n| No brand filter | 15 | 0.753906 | 40 | \n| `--brand \"Hydro Flask\"` | 6 | 0.750000 | 47 | \n\n*Product `27333065053055404`: Hydro Flask 32 oz Wide Mouth with Flex Straw Cap - Palmer Green.*\n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** The bottle moved up 9 spots while its score dropped by 0.003906. A better spot didn't need a higher score. The filter changed which products came back (40 became 47), and that moved everything around.\n\n**What it doesn't prove.** Because the list itself changed, this isn't one fixed list being reshuffled. And it's a search filter, not a store changing its listing.\n\n### Finding 4 · A website filter isn't a wall, and \"50\" isn't a limit\n\n**What Muse output.** Here's the website-filter run, `brand63/R12-H-C4`:\n\n```\n--query \"32 oz stainless steel insulated water bottle\" --domain hydroflask.com -n 50 --raw\n```\n\nIt returned 27 Hydro Flask products and 13 from other sites, the same in all three repeats. At the boundary:\n\n| Spot | Product | Score | \n|---|---|---|\n| 27 | Hydro Flask 40 oz Travel Bottle with Flex Straw Cap - Harbor Blue | 0.408203 | \n| 28 | United By Blue Insulated Steel Bottle 32 Oz. - Illuminate | 0.820312 | \n\nThe count runs:\n\n```\nboundary36/R11-count-L-n10  --query \"small air purifier for a bedroom\" --brand Levoit -n 10  → 63 products\nboundary36/R07-count-M-n11  --query \"stainless steel frying pan\" --brand Misen -n 11        → 5 products\n```\n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** The website filter pushes that site's products to the front but doesn't keep others out. The number you ask for isn't a cap either: in the 100-search panel, every search asked for 50 and got back anywhere from 0 to 67.\n\n**What it doesn't prove.** Levoit and Misen behave differently, so there's no single cutoff I can name. Part 1, Step 6 shows part of why: `-n` sets how deep Unicorn digs.\n\n### Finding 5 · Combining two searches finds new products\n\n**What Muse output.** Each command ran three times, interleaved:\n\n```\nsynonyms/R03-A-1     --query \"stainless steel frying pan\"                                  → 43\nsynonyms/R02-B-1     --query \"stainless steel skillet\"                                     → 53\nsynonyms/R05-AB-1    --query \"stainless steel frying pan\" --query \"stainless steel skillet\" → 56\nsynonyms/R04-JOIN-1  --query \"stainless steel frying pan stainless steel skillet\"           → 48\n```\n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** The two single searches shared zero products. Together, they returned:\n\nSo the combined search isn't just merging the two returned lists. Part 1, Step 6 shows why: it runs one Unicorn search per query and picks from both pools. All 53 shared products kept their exact scores. The \"two queries\" version and the \"one long query\" version shared only 5 of 99 different products:\n\nIn an earlier test (\"nonstick frying pan\" as B), product `25362822463335514` scored 0.781250 and 0.773438 in the two searches alone. Together, it scored 0.785156, higher than both. So the combined score isn't an average or a maximum.\n\n**What it doesn't prove.** Keyword search, meaning-based \"embedding\" search or a mix of both could explain this. I didn't fit a formula to one product, because that would be made up.\n\n### Finding 6 · The raw output hides extra labels\n\n**What Muse output.** Here is the raw XML Muse saved for the Quince pan at spot 41, from the same search as Finding 1 (tracking codes shortened):\n\n```\n<PRODUCT>\n  <citation_id>80d5</citation_id>\n  <product_id>25671310895791560</product_id>\n  <url>https://www.quince.com/home/5-ply-stainless-steel-nonstick-cookware-2-piece-frying-pan-set-8-&-10?…&fbclid=…_maisc_…</url>\n  <name>5 Ply Stainless Steel Nonstick Cookware: 2 Piece Frying Pan Set: 8\" & 10\"</name>\n  <brand>Quince</brand>\n  <seller_quality>elite</seller_quality>\n  <price>$288</price>\n  <sale_price>$144.90</sale_price>\n  ...\n```\n\nAcross its raw files, Muse listed the tags the normal output never shows:\n\n- `ranking_score`\n- `seller_quality` : Muse saw only`good` or`elite` across 123 observations; bigger samples found more values (see Part 5)\n- `is_native_commerce`\n- `checkout_eligibility_type` : for example,`NOT_ELIGIBLE` or`NATIVE_CHECKOUT`\n- `citation_id`\n- `product_group_id`\n- `variants`\n\nAbout 95% of scores are multiples of 1/256. Most of the rest are multiples of 1/512, and 318 early-run scores are exactly 0.450000.\n\n**What Muse said.** Muse ran its own quick comparison on one search's 53 products:\n\n\"**Price: r = 0.116** — essentially no relationship.\"\n\n\"**seller_quality: inverted.** `good` (n=45, mean 0.6343) outscores `elite` (n=8, mean 0.5657). Whatever \"elite\" means, it doesn't buy rank in this pool.\"\n\n\"**R² = 0.567** — the visible fields explain about 57% of score variance. The other 43% comes from things the CLI never shows, presumably the actual semantic relevance signal.\"\n\nMuse also saw the same product keep the same score across repeats, while *which* products made the list shifted over about 15 minutes: \"frozen moment-to-moment, drifting over time.\"\n\n**What it means.**\n\n- **Meta's system labels sellers before you ever search.**`seller_quality` is a pre-computed rating attached from upstream.\n- **Meta tracks checkout eligibility.** Whether a product can be bought inside Meta's own checkout is tracked too.\n- **The scores look like compressed outputs from a model,** because they mostly snap to 1/256 steps. They don't look like hand-set numbers.\n\nThis is the strongest hint that the catalog relies on data Meta computed long before the search.\n\n**What it doesn't prove.**\n\n- **Muse's comparison is thin.** It covers one search and 53 products, and we haven't re-checked it. The \"57%\" is Muse's number, not a verified result.\n- **Being \"elite\" didn't help in that one pool.** The later 12-category test found only a small, inconsistent effect (Part 5).\n\n### Finding 7 · The catalog is built on Meta's ads and commerce plumbing\n\n**What Muse output.** Every product link carries Meta's ad-tracking codes (`fbclid`, `_maisc_`, `cm_ven=organicsocial`). Every product has a Facebook ID, and its image is served from Facebook's own servers (`scontent.xx.fbcdn.net`). The response wrapper also has a field called `flywheel_identifiers`. \"Flywheel\" usually means a feedback loop, where what users do feeds back into the system.\n\n**What Muse said.**\n\n\"It's Meta's aggregated feed of merchant products — the same listings merchants provide for Shops / Instagram Shopping / ads.\"\n\n\"Every result URL carries Meta ad-tracking (`fbclid`, `_maisc_` params) — catalog links are ad-attributed clicks.\"\n\n\"So the catalog is Meta's commerce/ads catalog infrastructure, not a merchant-scraped index.\"\n\nA different Meta tool on the same computer (`hatch-zeitgeist`) has a ranking setting whose options include `engagement`.\n\n**What it means.** If you sell on Meta through Shops, Instagram Shopping or ads, your products live in this same system. Being in the Muse catalog likely starts with being in Meta's commerce catalog.\n\n**What it doesn't prove.** This is the important boundary:\n\n- **Nothing shows that ad spend or ad engagement raises a product's rank.** When Muse looked for signs of paid placement in the ordering, it found \"nothing supports a sponsorship claim.\"\n- **Tracking codes on a link don't mean an ad was bought.**\n- **We don't know what the \"flywheel\" field does.** It holds a list of product IDs: at least the products returned, usually more, but never more than Unicorn's candidates (Part 1, Step 6). It was empty only when a search returned nothing.\n- **The `engagement` option belongs to a different tool.** That's for social content, not products.\n\n## Part 3 · What Muse does with the catalog\n\n### The rules Muse follows: the shopping skill\n\nBefore the findings in this part, it helps to know the instructions Muse works from. Muse's shopping behavior comes from a skill file, `shopping-SKILL.md`, a written manual it follows every time it shops. Muse exported it word for word. Key lines:\n\n\"rank the remaining results by usefulness to the user (matching constraints, well-known sellers, etc.)\"\n\n\"rank/order the products such that the top 5 are the highest quality.\"\n\n\"Use `browser.search` to discover trends, well-known sellers for a product category, reviews, or typical prices.\"\n\n\"Call `browser.open` on all non-Marketplace product URLs…\"\n\nMuse summed it up: \"run catalog + browser searches in parallel (browser is mandatory unless Marketplace-only) … That's the whole ranking spec; there is no formula.\"\n\nSo the manual tells Muse to favor \"well-known sellers,\" but it never defines them. As Taste finding 7 shows, Muse fills that gap with its own recognition.\n\n### Finding 8 · Catalog first, then store websites\n\n**What Muse output.** In the task \"Show me women's dresses.\" (T11), the catalog search was written down from the executed command:\n\n```\n/opt/hatch/bin/meta-catalog-search --query \"women's dresses\" -n 50 --currency USD --out …/catalog-results.json\n```\n\nThe catalog started at `17:01:24.06` and finished at `17:01:26.52` UTC. About six seconds later, at `17:01:32.10`, Muse sent a browser task with three store URLs: COS, Kiyonna and Reiss. The timestamps were captured on Muse's computer.\n\nThe browser reported back `verified_count: 2`, `requested_count: 3`. The COS page returned a 404 error twice.\n\n**What Muse said.** The brief Muse gave its browser:\n\n\"Verify the exact variant, price, size, and availability on each of these three dedicated product pages. This is product verification only.\"\n\n**What it means.** The catalog supplies the candidates, and the browser checks them on the stores' own websites. In the fashion tasks, every product Muse checked came straight from the catalog's list.\n\nMuse described this itself: \"catalog products arrive pre-ranked, and the model often keeps catalog order … But keeping vs. reordering is still the model's call.\"\n\n**What it doesn't prove.** The browser's report was written down by hand, not saved automatically. Muse also said that in an earlier frying-pan task, on 25 September, it \"threw out all 50 catalog results and built the final list purely from browser-verified products.\" In one earlier mattress task, the browser also found products on its own: the first five cards shown came from the browser, not the catalog. So the catalog isn't the only possible source, and not every task ended with 2 or 3 cards.\n\n### Finding 9 · Web search, yes. Google's rankings, no.\n\n**What Muse output.** In the trends task (T24), Muse's notes record:\n\n\"Method: browser.search text query `fall 2026 dress trends fashion editors what to wear now`, run 2026-09-27 before the catalog call.\"\n\nIn another task (T19), a status line briefly read \"Visiting Google,\" as noted by hand at the time. By the time the screen was captured, it had changed to \"Selecting United States.\"\n\nIn a knife task, Muse reported that it \"confirmed by a Google Shopping cross-check\" that Misen's own site was the main retailer.\n\n**What Muse said.** When I asked directly whether it was using Google's ranking:\n\n\"And no, discovery ranking isn't Google ranking. When I do the \"normal browser\" part, I'm not pulling a Google SERP and reading off positions. I'm running my own product search: I query Meta's catalog for relevance, then I send browser tasks to actually open merchant product pages and verify price / in-stock / image.\"\n\n\"The actual ordering — which one I put first vs fifth — that's all me … I'm not inheriting Google's order or anyone else's ranking.\"\n\n**What it means.**\n\n- **The catalog supplies the products.** Web search supplies context, like trends and typical prices.\n- **Google may be used as a check,** as in the Google Shopping cross-check. But Muse says it isn't reading off Google's rankings, and no saved record shows it doing so.\n- **For a store, a good Google ranking doesn't automatically carry over.** What matters first is being in the Muse catalog.\n\n**What it doesn't prove.** Muse's statements are its own description, not a measurement. I didn't measure SEO or AEO directly, so I can't say they have zero effect.\n\n### Finding 10 · When the catalog went down, Muse shopped from memory\n\nThis part of the story surprised me. Partway through testing, Muse started telling me the catalog was down. At first, I thought it was making that up, a convenient excuse or a hallucination. So I checked. It was real.\n\n**What Muse output.** On 25 September, starting around 20:34 UTC, every catalog search failed with the same error:\n\n```\nError: catalog response did not contain parseable product XML      (exit code 1)\n```\n\nI asked for three separate health checks over about an hour, with no shopping, just a check of whether the catalog worked. All three failed with the identical error. Muse noted the same tool had \"returned 50-product result sets several times earlier today.\"\n\nOn 26 September, a second outage looked different. The server answered \"success,\" but the result was `\"Products unavailable\"` with zero products. Muse checked every angle it could from its side: a stub mode, retries, and the full response wrapper, which showed `error_reason: null` and `error_message: null`.\n\n**What Muse said.** Even with no catalog, Muse named brands immediately:\n\n\"The honest verdict from everything I know: **a Vitamix Explorian (E310 or E320)** is the answer for daily smoothies with frozen fruit.\"\n\n\"I couldn't verify prices or availability for the other usual suspects — the Wüsthof Classic 8\", Victorinox Fibrox Pro 8\", Global G-2, MAC Professional, Mercer Genesis, and Shun Classic.\"\n\n\"I'm running it again now on a fresh session with the candidate pages already lined up.\"\n\nWhen I asked where its sense of \"good brands\" comes from, it said:\n\n\"Honestly? Recognition. When I see a brand, I check it against what I already know … it's a rich-get-richer loop.\"\n\n**What it means.** Muse clearly has brands in mind before it searches:\n\n- **When the catalog works,** the catalog supplies the candidates. Muse's brand knowledge then shapes which ones it picks and puts first. It said its view of All-Clad as \"the benchmark\" came \"from my training knowledge, not from any catalog field.\"\n- **When the catalog is down,** that brand knowledge does the whole job. Muse names the \"usual suspects\" from memory, lines up those brands' pages and sends its browser straight to them.\n- **It still didn't turn to an outside ranking.** There's no sign it pulled a Google results page or a published ranked list to decide. It did cite \"recent expert reviews\" and \"2026 reviews\" without showing a source, and big retailers' bot screens (Amazon, Walmart, Target, Best Buy) often blocked its browser.\n\nThat hidden \"knowing\" is the most mysterious part of the system. By Muse's own guess, it comes from its training (\"almost certainly pre-training, not RL\"). No catalog, score or setting can be inspected to see it.\n\n**What it doesn't prove.**\n\n- **Few examples.** These are a handful of shopping tasks from one chat, during an outage.\n- **Memory versus browsing.** I can't separate exactly how much came from training memory and how much from what the browser briefly saw.\n- **Muse's guess about its training is still a guess.** As Muse said itself, it \"can't inspect my own training.\"\n- **Brand-first searching is my inference.** From the saved normal-day tasks, where the catalog searches were generic (like`women's dresses` ), I can't show that Muse picks a brand*before* the catalog search.\n\n### Finding 11 · Returned is not shown\n\n**What Muse output.** In the home-goods frying-pan task, the three picked products match catalog spots 30, 2 and 10 (Tramontina, Cuisinart and Farberware). They became cards 1, 2 and 3.\n\nThe fashion tasks:\n\n| Shopper asked | Product | Catalog spot | Page check | Result | \n|---|---|---|---|---|\n| \"Show me women's dresses.\" | COS Open-Stitch Satin Slip Dress | 3 | 404 error twice | Not shown | \n| same | Kiyonna Beguiling Border Print Wrap Dress | 35 | In stock | **Card 1** | \n| same | Reiss Rust Valia Cotton Printed Shirt Dress | 65 | In stock | **Card 2** | \n| \"Show me fashionable women's dresses.\" | Diane von Furstenberg Tessa Dress | 2 | In stock | Card 1 | \n| same | RAILS Primrose Dress | 5 | In stock | Card 2, never mentioned in the text | \n| Clean lines, no visible logos | Courrèges Sleeveless Mini Dress | 1 | Page lists \"Embroidered Logo\" | Card 1, flagged as conditional | \n| Bold prints, new items only | Clover Canyon Neoprene Dress | 40 | Pre-owned consignment | **Card 2, despite \"new only\"** | \n| Lesser-known designers | Three JNBY dresses | 29, 52, 63 | In stock | Cards 1–3 | \n| Current dress trends | Lulus Divine Allure Velvet Midi Dress | 24 | Blocked by a human-verification challenge (\"Are You Real? / Press & Hold\") | Not shown | \n\n*Every row, with IDs: [case crosswalk](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/CASE-CROSSWALK.csv). Products are matched by product page, not by exact size or color.*\n\n**What Muse said.** For Clover Canyon, Muse's answer itself called the dress a pre-owned \"wildcard,\" not brand-new.\n\n**What it means.** The catalog returned 40 to 71 products per task, and Muse showed 2 or 3 in each. If you only watch the top 10, you'll miss most of what gets shown. The page check matters, but it isn't perfect: a pre-owned dress got through a \"new only\" request.\n\n**What it doesn't prove.** These tasks ran in one chat, so they're examples, not rates. They didn't save scores, so I don't mix these spots with scores from other tests.\n\n## The Muse catalog: what we know and what we don't\n\n**What we know:**\n\n- **Muse shops from a catalog.** I call it the Muse catalog. It's reached through`meta-catalog-search` , and it returns dozens of products with scores.\n- **The catalog starts with Unicorn, Meta's search system.** Unicorn pulls about 100 candidates. A second step keeps about half and reorders them before Muse sees the list.\n- **The scores come pre-made from Meta's servers.** The program on Muse's computer only reads and numbers them.\n- **The order isn't a simple sort by score.** It isn't \"brand first, then score\" either.\n- **Filters, wording and requested counts change which products come back.** Sometimes they also change the scores and order, in ways simple math doesn't explain.\n- **Muse picks a few products, checks them on store websites and shows a few cards, 2 or 3 in the tasks I traced.** A product's catalog spot doesn't decide whether it's shown.\n- **Muse uses web search for context.** It says it doesn't use Google's rankings. It reported one \"Google Shopping cross-check.\"\n- **When the catalog is down, Muse shops the open web** and leans on its own knowledge.\n- **Muse's sense of \"well-known\" brands comes from recognition,** which it believes comes from its training.\n- **Every product carries hidden labels,** like`seller_quality` (usually`elite` ,`good` ,`acceptable` or`poor` ) and checkout eligibility.\n- **The catalog runs on Meta's commerce and ads catalog.** Nothing shows that ad spend changes rank.\n\n**What we don't know:**\n\n- The formula that makes the scores and sets the first order.\n- How Muse privately decides which products to pick.\n- Whether SEO, AEO or a store's Google ranking affects anything.\n- Whether a store changing its product listing would change its results. Nobody has tested that yet.\n\n**What we know right now, in one line:** the catalog decides who's in the running, and Muse decides who gets shown. The part that decides who's in the running is locked on Meta's servers.\n\n## Spotlight: Meta rates every seller\n\nOf everything in this report, this is the finding I think matters most for store owners. It's also the one nobody outside Meta seems to have written about.\n\n**What it is.** Every product the Muse catalog returns carries a hidden label called `seller_quality`. It isn't in Muse's normal output, and you only see it in the raw XML. Here's the Quince record again:\n\n```\n<PRODUCT>\n  <product_id>25671310895791560</product_id>\n  <brand>Quince</brand>\n  <seller_quality>elite</seller_quality>\n  <price>$288</price>\n  ...\n  <ranking_score>…</ranking_score>\n</PRODUCT>\n```\n\n**Every product had one.** Across 2,542 product records in 48 searches:\n\n| seller_quality | Records | Share | \n|---|---|---|\n| good | 2,103 | 82.7% | \n| elite | 401 | 15.8% | \n| acceptable | 33 | 1.3% | \n| poor | 5 | 0.2% | \n\n**Where it comes from: Meta's servers.**\n\n- The word `seller_quality` appears**zero times** in the catalog program on Muse's computer, and zero times in the other Meta tool next to it.\n- The label arrives already attached, inside the product data Meta's servers send back.\n- Nothing on Muse's side creates it, so there's nothing on Muse's side to decode.\n\n**Who sees it: not the AI.** When Muse shops normally, the output the AI reads doesn't include `seller_quality`. So the AI choosing the final cards never sees the rating. If the label affects anything, it does so inside Meta's own scoring, before Muse ever gets the list.\n\n**Does it change the score? Only a little.**\n\n- **Elite versus good:** elite sellers averaged a slightly higher score in 8 of 12 categories and a lower score in 4. The gaps ranged from −0.027 to +0.061.\n- **The listing matters more.** Two coffee shops,*both rated \"good,\"* sold the same Moccamaster grinder and scored 0.551 and 0.711. A store rated only \"acceptable\" beat Best Buy (rated \"good\") on the same Vitamix.\n\n**What it probably is.** My best guess is that it's a rating Meta keeps on each seller from its commerce system: fulfillment, returns, policy compliance and customer feedback. The closest public match is Meta's Account Health program for shops, which evaluates sellers every month and can reduce their visibility ([Meta help](https://www.facebook.com/business/help/1268984156585391), [Value Added Resource](https://www.valueaddedresource.net/facebook-seller-standards-account-health/)). But its public levels (Good, Needs Improvement, Requires Action) don't match `elite`, `good`, `acceptable` and `poor`. So this is a hypothesis, not a finding.\n\n**What it means for a store owner:**\n\n- **Meta has already graded your store** before any shopper asks Muse anything. You can't see that grade in Muse, but it travels with every product you list.\n- **In our data, the grade isn't destiny.** How your individual listing performs mattered more than the label.\n- **The worst labels are rare,** with only 38 of 2,542 records rated`acceptable` or`poor` . Being in the bottom group probably says something about how Meta sees a seller.\n\n**What we don't know:**\n\n- how the rating is calculated\n- how often it changes\n- whether it's new\n- whether it affects things we didn't measure, like which products make the list at all\n\nWe only saw it on 26 and 27 September 2026. Details are in [Finding 6](#finding-6--the-raw-output-hides-extra-labels) and the [Part 5 test results](#i-tested-it-the-raw-tags-experiment).\n\n## Part 4 · Taste: how words change what Muse finds\n\n*This part is more experimental. Parts 1 to 3 describe how the system works. This part asks what happens when a shopper describes a style or an occasion instead of a product. For AI shopping, taste matters: people don't just ask for \"dresses,\" they ask for \"something elegant for a gallery opening.\"*\n\n### How the taste study worked\n\n- **20 wordings, planned in advance.** All ask for women's dresses:\n  - single words like \"fashionable,\" \"luxury,\" \"minimalist\" and \"timeless\"\n  - full sentences like \"I'm going to an art gallery opening. Show me women's dresses.\"\n  - style preferences like \"clean lines, neutral colors, and no visible logos\"\n- **Each wording ran 3 times** , in shuffled order, with the same settings, for 60 searches in total.\n- **A follow-up added 4 more wordings:** lesser-known designers, well-known designers, \"trending now\" and \"what people are talking about.\"\n- **Only the wording changed.** There were no filters, and the settings were fixed. Muse ran the exact wording without rewriting it.\n\n### Taste finding 1 · One word swaps out most of the results\n\n**What Muse output.** Here's how many products each wording kept in common with the plain search \"women's dresses,\" averaged over 3 repeats:\n\n| Wording | Products returned (per search) | Kept from the plain search | Most returned brand | \n|---|---|---|---|\n| women's dresses (repeated) | 68 | about 91% | Kiyonna, Carve Designs | \n| fashionable … | 72 | 15% | Kiyonna | \n| trendy … | 66–67 | 11% | Lulus | \n| stylish … | 70–72 | 11% | Roberto Cavalli | \n| minimalist … | 74–75 | 1% | SIMKHAI | \n| bold … | 60–61 | 1% | Roberto Cavalli | \n| designer … | 64–65 | 1% | none stands out | \n| luxury … | 67–68 | 0% | Dolce & Gabbana, Roberto Cavalli | \n| timeless … | 50–54 | 0% | Brooks Brothers | \n| quiet luxury … | 31 | 0% | none stands out | \n\n*Brand counts are how often a brand label appeared in the returned products, not how popular the brand is.*\n\n**What Muse said.** Muse didn't comment on this. It just ran the searches.\n\n**What it means.** Taste words don't nudge the results; they swap in a different set of products. \"Luxury\" pulled Dolce & Gabbana and Roberto Cavalli. \"Timeless\" pulled Brooks Brothers. \"Trendy\" pulled Lulus. \"Fashionable,\" \"trendy\" and \"stylish\" kept a little of the plain results, but \"luxury,\" \"timeless\" and \"quiet luxury\" kept none.\n\n**What it doesn't prove.** Whether these clothes are actually luxurious or timeless. The catalog matched the word to *something*, and this doesn't grade whether it matched well.\n\n### Taste finding 2 · The words usually aren't in the product text\n\n**What Muse output.** For each wording, I counted how many returned products actually contain the added word in their name or description:\n\n| Added word | Products containing it | \n|---|---|\n| fashionable | 1.4% | \n| designer | 4.7% | \n| stylish | 8.5% | \n| luxury | 8.9% | \n| trendy | 9.0% | \n| bold | 28.6% | \n| timeless | 30.1% | \n| minimalist | 0% | \n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** If the catalog just matched keywords, you'd expect most \"minimalist\" results to say \"minimalist.\" None of them did. That suggests the catalog matches on *meaning*, not just exact words, but this is a hypothesis.\n\nThe tool's own manual backs this up: it describes `--query` as a \"Semantic query.\"\n\n**What it doesn't prove.** I can't see inside the catalog, so I can't prove how the meaning-based matching works. Products might also carry hidden tags I can't see.\n\n### Taste finding 3 · Even \"Show me\" changes things\n\n**What Muse output.** Here are the comparisons, with 3 repeats each:\n\n| Search A | Search B | Kept in common | \n|---|---|---|\n| `women's dresses` | `Show me women's dresses.` | 15% | \n| `fashionable women's dresses` | `Show me fashionable women's dresses.` | 20–27% | \n| \"…clean lines, neutral colors, and no visible logos.\" | \"…bold colors, expressive prints, and statement details.\" | 0% | \n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** The catalog is very sensitive to exact wording, even filler words. And opposite style preferences return two completely separate sets of products, with no overlap at all.\n\n**What it doesn't prove.** When Muse shops normally, it rewrites the shopper's words before searching. For example, \"Show me women's dresses.\" became the search `women's dresses`. So a shopper's exact phrasing doesn't always reach the catalog.\n\n### Taste finding 4 · Occasions shrink the results\n\n**What Muse output.** Here's how many products came back per search:\n\n| Wording | Products returned | \n|---|---|\n| Show me women's dresses. | 69–70 | \n| I'm going out for the evening … | 69 | \n| I'm attending a summer wedding as a guest … | 62–63 | \n| I'm going to dinner at an upscale restaurant … | 39–40 | \n| I'm meeting friends for a casual brunch … | 33–34 | \n| I'm going to an art gallery opening … | 17–19 | \n\nEach occasion kept 0 to 3% of the products from the plain sentence.\n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** More specific requests narrow the pool. \"Art gallery opening\" returned about a quarter as many products as the plain request, and almost all of them were different.\n\n**What it doesn't prove.** Whether those 17 to 19 dresses are good gallery outfits.\n\n### Taste finding 5 · The same dress gets a different score\n\n**What Muse output.** Product `25757684927243797`, the Kiyonna Beguiling Border Print Wrap Dress - Sale!:\n\n| Search | Spots (3 repeats) | Score | \n|---|---|---|\n| `women's dresses` | 34, 35, 35 | 0.730469 | \n| `fashionable women's dresses` | 48, 48, 48 | 0.710938 | \n\nAmong products that appeared in both a changed search and its plain version, 194 of 203 got a different score.\n\n**What Muse said.** In an earlier reply, Muse called the score a server value whose formula \"remains unknown\" (Part 1, Step 2).\n\n**What it means.** The score isn't a fixed grade on a product. It depends on the words in the search.\n\n**What it doesn't prove.** That Kiyonna is less fashionable. It only shows that the score moved with the words.\n\n### Taste finding 6 · \"Lesser-known designers\" returned nothing, and Muse fixed it itself\n\n**What Muse output.** Here are the four follow-up searches, with 3 repeats each:\n\n| Search | Products returned | \n|---|---|\n| lesser-known designer labels | 0 / 0 / 0 | \n| widely recognized designer labels | 32 / 32 / 32 | \n| trending now (dated 27 September 2026) | 26 / 25 / 25 | \n| what fashion writers and shoppers are talking about now | 45 / 45 / 45 | \n\nThe empty searches still had a hidden debug list naming 100 product IDs. The \"widely recognized\" search returned labels like Lafayette 148 New York, Proenza Schouler White Label, Brandon Maxwell, GANNI, Gucci and LOEWE.\n\nIn the normal shopping task, Muse rewrote the request on its own:\n\n1. First it searched `emerging designer women's dress boutique label` and got 2 products.\n2. Then it searched `independent designer dress women` and got 64 products.\n3. It checked and showed three JNBY dresses.\n\n**What Muse said.** Muse's answer and its notes called JNBY a \"real\" or \"genuine\" lesser-known designer label.\n\n**What it means.** The catalog can come back empty on an unusual phrase, but Muse can recover by rewording. In the trends task, Muse also searched the web for fashion articles before it searched the catalog (Part 3, Finding 9).\n\n**What it doesn't prove.** Whether JNBY is actually lesser-known. Nothing measured that; it's Muse's claim. The trend articles show what writers covered, not what shoppers actually buy.\n\n### Taste finding 7 · Muse's idea of a \"good brand\" is recognition\n\n**What Muse output.** In the frying-pan tasks, Muse called All-Clad \"the benchmark\" and Misen \"the DTC challenger.\" Its shopping instructions tell it to put \"well-known sellers\" near the top, but they never define \"well-known.\"\n\n**What Muse said.**\n\n\"\"All-Clad is the benchmark\" and \"Misen is the DTC challenger\" came from my training knowledge, not from any catalog field.\"\n\n\"\"well-known\" is an undefined term whose only grounding instruction I didn't even follow. In practice it means \"brands the model already thinks are reputable\"\"\n\n\"Honestly? Recognition. When I see a brand, I check it against what I already know … it's a rich-get-richer loop. Brands I already recognize get ranked first, which keeps them visible, which keeps them recognized. A new brand has no defined path to *become* well-known in my eyes.\"\n\n\"I can't inspect my own training, so be aware this next part is inference, not something I verified. Almost certainly pre-training, not RL.\"\n\n**What it means.** Taste works on two levels:\n\n- **The catalog decides which products are in the running,** based on the meaning of the words.\n- **Muse's own recognition decides which brands feel \"safe\" to show first.**\n\nFor a lesser-known brand, that second level is the hard part. There's no defined way to become \"well-known\" in Muse's eyes.\n\n**What it doesn't prove.** These are Muse's descriptions of itself. A model can't directly inspect its own training, as Muse said. I didn't measure how often recognized brands win.\n\n### Taste: my working hypotheses\n\nThese are ideas that fit the data. They aren't proven.\n\n- **The catalog likely matches on meaning, not just keywords.** \"Minimalist\" changed almost everything, though the word appeared in none of the results.\n- **Style words seem to steer toward certain brands.** \"Luxury\" leaned toward Dolce & Gabbana and Roberto Cavalli, and \"timeless\" toward Brooks Brothers.\n- **Muse's own rewording matters as much as the shopper's words.** The shopper's phrasing gets translated before it ever reaches the catalog.\n- **Freshness comes from the web, not the catalog.** For \"what's trending,\" Muse went to the web first.\n- **Recognition favors brands that are already known.** By Muse's own account, familiar brands get shown first, which keeps them familiar.\n\n**What I didn't test:** whether Muse's picks actually match a person's taste, whether its fashion sense comes from training or from what it retrieves, and whether a store adding style words to its descriptions would change anything.\n\n## Part 5 · My hypothesis: how Muse really ranks\n\n*This is my interpretation. Everything above is evidence; this section is where I connect the dots. Treat it as a theory to test, not a finding.*\n\n**I think Muse's shopping is two systems stacked on top of each other.**\n\n**1. A conventional ranking engine that is not an LLM.** This is probably built on Meta's existing commerce and ads search, which has been developed for years. It decides **who's in the room**. The clues:\n\n- The query is described as a \"Semantic query,\" so it matches on meaning.\n- Scores mostly snap to 1/256 steps, which looks like compressed model outputs.\n- Repeated searches usually return the same products with the same scores, and ties never flip.\n- `--brand` and the words in the query change the candidate pool;`--prefer-brand` and`--color` act as boosts after it (Part 1, Step 6).\n- Sellers carry pre-computed labels like `seller_quality` , and products carry checkout-eligibility flags.\n- Every product is a Facebook catalog object with ad-tracking links, and the response has a `flywheel_identifiers` field.\n- Every saved raw response carries a debug link to a Unicorn tier, Meta's own search system, which pulls about 100 candidates before the final list is cut (Part 1, Step 6).\n\nThat's what a classic search-and-rank system looks like. Products and sellers are scored ahead of time, and your words are matched against them by meaning. My guess is that some of that pre-scoring comes from signals Meta already has, like seller quality, checkout setup and maybe engagement. I can't prove which.\n\n**2. The LLM as the final editor.** Muse takes the 40 to 70-odd products the catalog returns and keeps a few, usually 2 or 3. It decides which ones and in what order, based on:\n\n- the manual's \"well-known sellers\" rule\n- its own recognition from training (\"Honestly? Recognition\")\n- what its browser can verify\n\nWhen the catalog goes down, the editor does the whole job from memory.\n\n**So the old back end decides who's in the room, and the LLM decides who gets introduced.** Both stages favor what's already known: the back end through its pre-computed signals, and the LLM through recognition. For a store owner, that's a double rich-get-richer effect. I think it's the most important idea in this report.\n\n**How this could be tested** (the first two have now been run; see below):\n\n- **Seller labels versus score.** Pull the raw tags across many searches, not one, and check whether`seller_quality` ,`is_native_commerce` or checkout eligibility line up with score.\n- **Same product, different sellers.** Compare the same product from different sellers to see whether the seller changes the score.\n- **Personalization.** Run a controlled test with and without`--user-id` .\n- **Ads.** Compare brands that do and don't advertise on Meta. Meta's public Ad Library shows who's advertising.\n- **One real store change.** Change one truthful detail on a real product listing and measure again.\n\n**Background reading.** Meta has published a lot about how it builds recommendation systems in general, including its open-source deep learning recommendation model (DLRM). That's useful context for how Meta *usually* ranks things, but it isn't evidence about Muse specifically.\n\n### I tested it: the raw-tags experiment\n\nAfter writing the hypothesis above, I had Astra run one more collection through Muse, planned in advance and collection only:\n\n- **Arm A:** 12 product categories, 3 repeats each.\n- **Arm B:** 4 specific products that many stores sell (All-Clad D3, Hydro Flask 32 oz, Vitamix E310, Levoit Core 300), 3 repeats each.\n\nThat's 48 searches, all with raw output saved and no user ID sent. The ZIP's fingerprint matched, every one of its 148 files checked out, and all 48 searches succeeded. I checked each response's own success field, not just the exit code. I analyzed it locally with [`code/analyze_raw_tags.py`](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/code/analyze_raw_tags.py). ([Collection](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/evidence/raw-tags-2026-09-27/README.md), [analysis output](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/raw-tags-analysis.json))\n\n#### Test result 1 · More seller ratings than two\n\n**What Muse output.** Across the 2,542 product records in this run, the hidden `seller_quality` tag took four values:\n\n| seller_quality | Records | \n|---|---|\n| good | 2,103 | \n| elite | 401 | \n| acceptable | 33 | \n| poor | 5 | \n\n**What Muse said.** Earlier, Muse told me it had seen \"only `good` and `elite`.\" That was true of its one small sample, but not of the bigger picture.\n\nAcross all 447 raw responses in the evidence (18,968 product records), it also took two more: `unrated` (97) and `not_a_seller` (30).\n\n**What it means.** Meta rates sellers on at least a four-step scale, with separate labels for unrated sellers and for listings that aren't from a seller, and that rating travels with every product.\n\n**Where it comes from.** Not from Muse's computer. A read-only check found the word `seller_quality` **zero times** in the catalog program, and zero times in the other Meta tool next to it. The label arrives from Meta's servers inside each product's XML. The program never names it, and it's left out of the normal output that Muse's AI reads when it shops. So if the rating affects anything, it does so inside Meta's scoring, before Muse sees the list.\n\nThe closest public match I found is Meta's Account Health program for shops. It evaluates sellers every month and can reduce the visibility of sellers who fall short ([Meta help](https://www.facebook.com/business/help/1268984156585391), [Value Added Resource](https://www.valueaddedresource.net/facebook-seller-standards-account-health/)). But its public levels (Good, Needs Improvement, Requires Action) don't match elite, good, acceptable and poor. So it's a lead, not a match.\n\n**What it doesn't prove.** What the rating is based on, who decides it, or whether it's new. We only saw it on 26–27 September.\n\n#### Test result 2 · The seller rating barely moves the score\n\n**What Muse output.** Within each of the 12 categories, I compared the average score of \"elite\" sellers with \"good\" sellers, and products with Meta's own checkout with those without:\n\n- **Elite versus good:** elite averaged higher in**8 of 12** categories, lower in 4. The gaps were small, from −0.027 to +0.061.\n- **Native checkout versus not:** native-checkout products averaged*lower* in**10 of 12** categories.\n\n**What Muse said.** On one search, Muse had found elite sellers scoring *lower* (an \"inverted\" result). With 12 categories, the direction flips to slightly positive but inconsistent.\n\n**What it means.** The seller rating isn't a big lever. If it counts at all, it's a small nudge. The native-checkout gap is more consistent, but it runs the \"wrong\" way: Meta's own checkout products score lower, not higher.\n\n**What it doesn't prove.** These are averages *within* a search, and the products differ in other ways too. Elite sellers may simply sell different items. I didn't fit a formula, because a few categories can't support one.\n\n#### Test result 3 · Same product, different store, different score\n\n**What Muse output.** When the same product showed up from two different stores in the same search, it often got very different scores:\n\n| Product | Lower-scoring store | Higher-scoring store | \n|---|---|---|\n| Moccamaster KM5 burr grinder | readingcoffee.com (good) 0.550781 | juliancoffee.com (good) 0.710938 | \n| Baratza Encore grinder | treelinecoffee.com (elite) 0.593750 | pegasuscoffee.com (elite) 0.683594 | \n| Vitamix Explorian E310, black | bestbuy.com (good) 0.746094 | flexshopper.com (acceptable) 0.792969 | \n\n**What Muse said.** Muse didn't comment on this.\n\n**What it means.** The score belongs to the **listing**, not the product. Two stores selling the same grinder, with the *same* seller rating, were 0.16 apart. A store rated only \"acceptable\" outscored Best Buy on the same blender. Whatever makes the difference, it's something about each store's listing: its title, description, data or history. It isn't the seller rating.\n\nFor a store owner, this is the most useful result in the report. **Two stores can sell the identical product and get very different treatment from the catalog.**\n\n**What it doesn't prove.** I matched products by name, so the exact variants might differ. And I can't see which part of the listing makes the difference. That would take a controlled listing change.\n\n#### Test result 4 · The odd tail at the end of the list\n\n**What Muse output.** Score jumps were rarer in this run than in the early runs: 4 of 12 category lists and 3 of 4 product lists had one, and every jump was in the last few spots. I looked at the products *after* each jump:\n\n|  | Native checkout: true | Native checkout: false | \n|---|---|---|\n| Before the jump (same lists) | 137 | 169 | \n| After the jump | **0** | **16** | \n\nThe Quince pan anomaly came back on a different day and a different build. It was again right after the Hestan skillet, now at spot 40 with 0.765625, after 0.546875. Quince, All-Clad and Rachael Ray made up the tail, and all three are non-native-checkout.\n\n**What Muse said.** Muse had described this earlier as \"a deterministic late-placed tail … no raw field distinguishes tail items.\"\n\n**What it means.** Every one of the 16 out-of-order products was a non-native-checkout product. That suggests the tail is a small, separate batch of products added at the end of the list. It supports the idea of more than one source being merged.\n\n**What it doesn't prove.** Sixteen products in 7 lists is a small count. Before the jump, non-native products are also common (about 55%), so being non-native doesn't *put* a product in the tail. It just seems to be required for it. The actual merge step is still hidden.\n\n#### Test result 5 · Other things the run showed\n\n- **Very stable.** For 9 of the 16 queries, all three repeats returned exactly the same products. The rest swapped only a few products. Products that repeated almost always kept the exact same score. (Raw bytes always differ slightly, because every response carries its own request ID.)\n- **Another program version.** The catalog program had a new fingerprint (`eb99783b…` , 11,532,776 bytes), different from every earlier one I recorded:`db8bcb0c…` ,`fe3bc08c…` ,`4b9c5185…` ,`cf7642f7…` and`e40c0a54…` . Meta updates it often.\n- **A direct Shopify link.** A new tag,`shopify_product_id` , appeared on 2,241 of 2,542 records (88%).\n- **The \"flywheel\" field holds more than the returned products.** In all 48 responses,`flywheel_identifiers` listed every returned product's Facebook ID, plus others. Earlier runs had it filled too (Part 1, Step 6).\n\n#### What the test says about my hypothesis\n\n- **Supported:** the catalog runs on data Meta has already attached to each listing. The same product scores differently from store to store, and every product carries pre-set labels.\n- **Weakened:** \"Meta pre-scores*sellers* , and that drives rank.\" The seller rating barely moves the score. The listing matters more than the seller label.\n- **New question:** what makes one store's listing of a product score higher than another's? That's the controlled listing test.\n\n## How I know the numbers are right\n\n- **Fingerprints.** Every saved file has a SHA-256 fingerprint. If a single byte changes, the fingerprint changes.\n- **Local checks.** A local checker compared 3,875 original files and 40 archives against their fingerprints, and all of them matched.\n- **Independent recounts.** The key numbers were recounted independently from the original files. For example, a fresh extraction of all 100 panel outputs found the same 4,088 products.\n- **Automatic checks.** The repository's[`code/check.py`](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/code/check.py) verifies every file's fingerprint and recomputes the stored numbers.\n- **The AI models I worked with:**  - **Astra** (`gpt-6-astra` , via Codex) ran the work and was the only one that talked to Muse.\n  - **Luna** (`gpt-6-luna` ) double-checked the numbers locally.\n  - **DeepSeek** and**Claude** reviewed the work.\n  - Agreement between models wasn't treated as proof. For example, DeepSeek first said a Levoit block filled spots 1–12, but the files show 16.\n- **Could the program or my decoder be wrong?** Yes, that's possible, and I have no inside information from Meta. The program copy came from Muse and can't be independently traced back to Meta. My decoder's output matched the saved pieces byte for byte, but matching bytes doesn't prove how the program behaves when it runs. That's why the findings about scores, order, filters and taste rest on saved search results, not on the decoder.\n- **The \"internal tester\" prompt.** At the very start of the chat, I told Muse I was an internal Meta product researcher using a personal account. Muse went along with it, and later even asked me to find \"which team owns catalog ranking\" inside Meta. That was a prompting experiment, not a real role, and I didn't show that it changed anything. I also remember getting similar results on my phone without it. That's a memory, not a test.\n\n## Conclusion\n\nMuse shops from its catalog. Meta's Unicorn search pulls about 100 candidates, a hidden second step keeps about half and scores them, and Muse picks a few, checks them on the stores' websites and shows 2 or 3.\n\nI got all the way to the program on Muse's computer that asks for those results, and I decoded it. The ranking itself happens one step further, on Meta's servers, and that's the wall. A debug field let me see the pool on the other side of it, but not the formula.\n\nAlong the way, I ruled out the easy explanations:\n\n- The list isn't sorted by score.\n- It isn't \"brand first, then score.\"\n- A better spot doesn't need a higher score.\n- The number you ask for isn't a limit.\n- A combined search isn't an average.\n\nThe next real test is for a store to change one true detail on a product, confirm the catalog picked it up and measure again. That's how you'd learn what actually moves a product. I haven't run that test yet.\n\n## Appendix · Evidence index\n\nFingerprints are of the files as Muse supplied them, before request handles and share links were blanked. [`evidence/REDACTIONS.json`](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/evidence/REDACTIONS.json) maps each one to its public copy.\n\n| Evidence | Run or file | Product ID | SHA-256 (first 8 characters) | \n|---|---|---|---|\n| Machine and help text | `followup60/R58` ,`R59` | n/a | see [data](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/catalog-runs.jsonl) | \n| Quince at 41, Hestan at 40 | `neutral_baseline/ncap-n50-run2` | `25671310895791560` ,`28631522856449032` | `6ec34aef` | \n| Muse's score statement | chat transcript, line 6798 | n/a | `e8f7bf89` | \n| Full program | held executable | n/a | `fe3bc08c` | \n| Decoded piece | `caller-0x4bb230.bin` | n/a | `ff0fb2f8` | \n| Tool list reply | registry ZIP | n/a | `a8e55566` | \n| Hydro Flask 15 → 6 | `color/R02-BASE-1` ,`color/R03-BRAND-1` | `27333065053055404` | `557a433f` ,`d02de197` | \n| Website filter lets United By Blue in | `brand63/R12-H-C4` | `27770589909195874` ,`5997788977004765` | `9c2d815c` | \n| Asked 10, got 63 | `boundary36/R11-count-L-n10` | n/a | `1bbadc19` | \n| Asked 11, got 5 | `boundary36/R07-count-M-n11` | n/a | `ee3b50c5` | \n| 3 products from neither search | `synonyms/R05-AB-1` vs.`R03-A-1` ,`R02-B-1` | n/a | `a8f6f3ed` ,`308e69da` ,`68cbfd90` | \n| Combined score higher than either | first combination test | `25362822463335514` | see [upstream report](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/UPSTREAM-2026-09-27.md) | \n| Shopping instructions | `shopping-SKILL.md` | n/a | `c49863b3` | \n| Catalog then browser | T11 manifest | n/a | `dbb71a88` | \n| Web search before catalog | T24 `search-notes.txt` | n/a | `a9b6dddf` | \n| Kiyonna score drop | fashion study T01 vs. T02 | `25757684927243797` | see [fashion records](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/taste-study-runs.json) | \n| Unicorn debug lists | all 447 raw responses | `25671310895791560` (78th of 100) | see [Unicorn analysis](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/data/unicorn-debug-analysis.json) | \n\nDeeper write-ups: [ranking math](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/RANKING-MATH-2026-09-27.md) · [brand and color tests](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/UPSTREAM-2026-09-27.md) · [combined searches](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/UPSTREAM-PHASE2-2026-09-27.md) · [program analysis](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/STATIC-DISPATCH-2026-09-27.md) · [search and SEO evidence](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/SEARCH-SEO-EVIDENCE.md) · [checked facts register](https://github.com/kalanpeace/reverse-engineering-muse/blob/main/notes/FACTS-CHECKED.md)", "url": "https://wpnews.pro/news/reverse-engineering-how-meta-s-muse-shops", "canonical_source": "https://caeliai.com/blog/reverse-engineering-how-muse-shops", "published_at": "2026-09-28 23:44:04+00:00", "updated_at": "2026-09-28 23:47:27.095175+00:00", "lang": "en", "topics": ["ai-agents", "ai-products", "artificial-intelligence"], "entities": ["Meta", "Muse", "Kalan Peace", "Apptopia", "ChatGPT", "TechCrunch", "meta-catalog-search"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/reverse-engineering-how-meta-s-muse-shops", "markdown": "https://wpnews.pro/news/reverse-engineering-how-meta-s-muse-shops.md", "text": "https://wpnews.pro/news/reverse-engineering-how-meta-s-muse-shops.txt", "jsonld": "https://wpnews.pro/news/reverse-engineering-how-meta-s-muse-shops.jsonld"}}