Reverse Engineering How Meta's Muse Shops Independent researcher Kalan Peace reverse-engineered Meta's Muse AI shopping agent and found the product ranking is driven by a `ranking_score` field attached to each product, with the catalog search handled by an 11.7 MB compiled Rust program at `/opt/hatch/bin/meta-catalog-search`. Muse, Meta's AI agent that gives each user a cloud Linux computer, passed 2.8 million installs worldwide in its first 12 days after launching in early September 2026, with about 642,000 daily U.S. users, according to Apptopia. Peace also reported that raw catalog results returned internal debug links to Meta's internal search system, which sit behind a Meta employee sign-in and showed only an "Internal Login" page when one link was opened on 28 September. Back to Blog https://caeliai.com/blog Reverse Engineering How Muse Shops Research orchestration: Astra gpt-6-astra . Independent checks: Luna gpt-6-luna, max . Reviews: Claude and DeepSeek. Editing: Claude claude-opus-5-5 . Evidence, code and data: github.com/kalanpeace/reverse-engineering-muse https://github.com/kalanpeace/reverse-engineering-muse . A note for anyone at Meta. The raw catalog results Muse gave me include internal debug links to Meta's own search system Part 1, Step 6 step-6--looking-over-the-wall-unicorn . They all go to Meta's internal network, which sits behind a Meta employee sign-in: the one link that was opened, on 28 September, showed only Meta's "Internal Login" page. I don't have a Meta login, so all I have are the addresses, not what's behind them. In the public data, the request handle in each link is blanked, and the originals are kept privately. If anyone working at Meta would like to see more of what Muse gave me, contact me at \ email protected\ https://caeliai.com/cdn-cgi/l/email-protection 9ef5fff2fff0eefbfffdfbdef9f3fff7f2b0fdf1f3 . I don't think this qualifies for a bug bounty, but I wanted to flag it. What Muse is Muse is Meta's new AI agent. It doesn't just answer questions like a chatbot. Meta gives each user's Muse its own Linux computer in the cloud, with a web browser, a place to save files and a set of tools. Muse can run commands, open websites, write files and send smaller helper agents off to do parts of a job. You talk to it through Meta's mobile and web apps. Muse is also one of the fastest-growing AI apps around. It launched in early September 2026, only about three weeks before this report. By Apptopia's count, it passed 2.8 million installs worldwide in its first 12 days, with about 642,000 daily users in the U.S. That's faster than ChatGPT's early mobile launch. TechCrunch https://techcrunch.com/2026/09/21/metas-muse-is-outpacing-chatgpts-early-mobile-launch/ A few facts about how it's built, from Meta's own description Meta https://research.meta.ai/blog/security-and-safety-for-ai-agents-our-approach-with-muse , published September 8, 2026 : - Each user's Muse runs in its own isolated section of a dedicated virtual machine. - Muse has a browser, a working folder of files, command-line tools, skills and helper agents. - Being in charge of your own Muse computer is not the same as controlling Meta's servers. Built-in connectors run in separate, protected "workers." One of the things Muse can do is shop . Ask it for a frying pan or a dress, and it searches for products, checks their pages and shows you a few as product cards. What I did I wanted to know how Muse decides which products to show. So I took it apart from the outside. - I ran hundreds of catalog searches through Muse and saved exactly what came back. Every file got a SHA-256 fingerprint, so any change to it would show. - I found the file names behind Muse's shopping. They include the catalog program /opt/hatch/bin/meta-catalog-search , the protected worker it talks to at /run/hatch/privsep/meta-catalog-search.sock , the shopping instructions Muse follows shopping-SKILL.md and the ranking score attached to each product. - I got a copy of the catalog program itself. It's an 11.7 MB compiled Rust program. I wrote my own decoder to read its machine code without ever running it. - I ran controlled experiments. I changed one thing at a time, like the wording, a brand filter or the number of products requested, and repeated each test three times. - I followed full shopping tasks from the question, to the catalog, to the page checks, to the cards on screen. What I found A simple way to understand how Muse shops: - The shopper asks a question. - Muse turns it into a search of what I'll call the Muse catalog , the product catalog behind meta-catalog-search . - Muse turns it into a search of what I'll call - The Muse catalog returns a list. - In my shopping tasks, that was 40 to 71 products, each with a score. - Muse picks a few. - Usually it picked 3 to check. - A browser checks those products on the stores' own websites. - It confirms price, stock and the exact variant. - Muse shows the shopper a few cards: 2 or 3 in the tasks I traced. What stood out: - The catalog is the backbone. In every shopping task I traced, Muse searched the catalog first and then checked products in a browser. - Muse says it doesn't use Google's rankings. In its words: "discovery ranking isn't Google ranking … I'm not pulling a Google SERP and reading off positions." Its browser goes straight to store websites. It does use web search for context, like fashion trends, and once reported "a Google Shopping cross-check" to confirm a retailer. - Meta rates every seller, and the AI never sees it. Each product arrives with a hidden seller quality label, usually elite , good , acceptable or poor . It comes from Meta's servers, isn't named in the catalog program on Muse's computer, and is left out of what the AI reads when it shops. - The same product scores differently depending on the store. Two coffee shops selling the same Moccamaster grinder, with the same seller rating, scored 0.551 and 0.711. - Muse already has brands in mind. When the catalog went down, it named brands straight from memory "The honest verdict from everything I know: a Vitamix…" and sent its browser directly to those brands' sites. At first I thought the outage was a hallucination. It wasn't: three health checks confirmed it. - The score doesn't decide the order. Products with higher scores often sit below products with lower scores. - Being returned isn't being shown. Products at spots 30, 35 and 65 became cards 1, 1 and 2. A 3 product with a broken page was never shown. - Taste words swap out the results. Adding "luxury" or "timeless" to a dress search replaced every product in it. Repeating the same search kept about 91%. - Muse's sense of "good brands" comes from its training. It said so itself: "Honestly? Recognition … it's a rich-get-richer loop." - The program I got is not the ranking system. It's the program on Muse's computer that asks Meta's servers for results. It sits one step before the ranking. - That's where I hit the wall. The formula behind the score lives on Meta's servers. I couldn't reach it, and neither can anyone outside Meta. - But I could see over it. A hidden debug field shows the first step behind the wall: Meta's Unicorn search pulls about 100 candidates, and a second step keeps about half and reorders them. One real saved search, "stainless steel frying pan" on 27 September, product by product. Part 1, Step 6 explains the three lists. The rest of this report is highly technical: raw output, math and code, explained in full. Own a store? Book a free call → https://cal.com/caeliai/free-analysis We'll see if you qualify for a two-week sprint. I apply what I learned here to your products: I find where they drop out of Muse's results, fix what's fixable in your listings and measure again, so you can see what actually changed. Not everyone will qualify, and no one can honestly promise a ranking. What I can promise is that you'll know exactly where you stand. Technical report The short version The scores come pre-made. Every ranking score I saw arrived already calculated. It comes from upstream, on Meta's servers, which aren't open to the public. That makes sense, and it's normal for any company's ranking system. The only part I could fully see was the catalog program on Muse's computer, meta-catalog-search . I got the whole program and decoded it, but it doesn't rank anything. It sends the search off, receives the results and interprets what it gets back: it reads the products, numbers them and prints them. Everything above it, where the actual ranking happens, is server-side. The formula never showed up in anything I could reach, but a debug field did show the pool it picks from Part 1, Step 6 . So this report maps everything around the ranking: what goes in, what comes out, what Muse does with it and what changes the results. It doesn't reveal the formula itself. Where these files came from These aren't answers I got by asking a chatbot questions. Here's how the files for the controlled experiments were made: - Script first. I wrote each test as a script ahead of time: the exact searches, the order and how many repeats. - Muse ran it on its own computer. It saved the raw output of every search to a file, along with a receipt showing the exact command, start and end times and exit code. - Muse packed and handed over the files. It bundled the files into ZIP archives with a manifest listing every file and its SHA-256 fingerprint. - I checked every file. I downloaded the archives and checked every file against the manifest and every archive for corruption. I also confirmed that the script Muse ran was byte-for-byte the one I wrote. For example, all 60 searches in the fashion study's main schedule passed every check: receipt, command, schedule, script, output fingerprint and a 183-file manifest. You can watch this happen in the full chat transcript, where Muse creates, packages and hands over each batch of files. Where a record was written down by hand instead of saved automatically, like some browser reports, I say so. What could still be wrong. I have no inside information from Meta, and nobody at Meta confirmed any of this. The program copy came from Muse, so I can't independently prove it's the exact program Meta runs, and later tests reported different versions. My decoder could also misread something. That's why the main findings in Parts 2 to 4 rest on saved search results, which don't depend on the decoder at all. Follow along Everything behind this report is public on GitHub: github.com/kalanpeace/reverse-engineering-muse https://github.com/kalanpeace/reverse-engineering-muse . It includes: - The full chat transcript with Muse https://github.com/kalanpeace/reverse-engineering-muse/tree/main/transcript , word for word, so anyone can read how I got to each point. Every quote in this report comes from it or from a saved file. - Every file Muse gave me https://github.com/kalanpeace/reverse-engineering-muse/tree/main/evidence : the raw search outputs, receipts, manifests and archives. - My decoder https://github.com/kalanpeace/reverse-engineering-muse/tree/main/code and the analysis and chart code. - The data and fingerprints https://github.com/kalanpeace/reverse-engineering-muse/tree/main/data behind every number. Every run ID like neutral baseline/ncap-n50-run2 and product ID like 25671310895791560 can be looked up there. A note about the chat. Some of the messages on my side were typed by Astra, my lead AI agent, using computer use to operate the Muse chat on my behalf. 23 messages are signed "Astra here," and 3 more open with "Astra." Others may not be labeled. Astra followed the research plan I set, and it was the only agent allowed to talk to Muse. How sure I am - Very sure: the numbers from saved search results. That covers scores, spots, counts, overlaps, and which products were returned and shown. These are counted directly from files with fingerprints. - Fairly sure: the reading of the catalog program. The bytes check out, but reading machine code is interpretation, and the copy could differ from what Meta runs. - Muse's word only: Muse's descriptions of itself, like "discovery ranking isn't Google ranking," "recognition" and "chosen server-side." They're quoted exactly, but they're claims, not measurements. When I pushed Muse for evidence, it sometimes walked a claim back. For example, it admitted its "verbatim server body" claim rested on a search it never saved. Those corrections are in the transcript too. - Hypotheses: anything labeled that way, like meaning-based matching or a recognition loop. They fit the evidence, but they aren't proven. How this report works - Part 1: how I got into the Muse catalog, what came out, what the program looks like inside, and what can be seen over the wall. - Part 2: what the catalog does. - Part 3: what Muse does with the catalog's results, followed by a summary of what we know and don't know. - Part 4: taste. This is the more experimental part, about how words like "fashionable" or "luxury" change what Muse finds. - Part 5: my hypothesis about how it all fits together, and how to test it. Every finding is laid out the same way: - What Muse output: the exact command and the exact result. - What Muse said: Muse's own words about it, or its own instructions. - What it means: my interpretation, in plain words. - What it doesn't prove: where the evidence stops. The experiments were separate, and I never add their results together: | Experiment | Searches | Products returned | |---|---|---| | Early catalog runs | 182 | 6,314 | | Combined-search, color and brand tests | 27 | 1,200 | | Second combined-search tests | 21 | 1,104 | | Broad store panel 20 product types × 5 question types | 100 | 4,088 3,827 different products | | Fashion wording study | 72 | 3,827 1,125 different products | | Home-goods shopping tasks | 5 | 270 → 14 cards shown | | Fashion shopping tasks | 7 in 6 tasks | 375 → 15 cards shown | | Raw-tags test 12 categories + 4 products, 3 repeats each | 48 | 2,542 | The two 3,827s are a coincidence: one counts different products in the panel, the other counts all product records in the fashion study. Part 1 · Getting into the Muse catalog Step 1 · Finding the tool What Muse output. Muse's own computer answered two basic questions: followup60/R58 uname -srm → Linux, x86-64 followup60/R59 /opt/hatch/bin/meta-catalog-search --help → the tool's own manual Muse pasted the tool's manual its --help text into our chat word for word. It starts: Meta 1P product catalog search. Usage: meta-catalog-search OPTIONS -q, --query