# Ask Threads for 50 Replies and Half of What Comes Back Is Other People's Posts

> Source: <https://dev.to/nikita_iakovlev_415524c19/ask-threads-for-50-replies-and-half-of-what-comes-back-is-other-peoples-posts-pao>
> Published: 2026-10-10 21:13:39+00:00

I run a visa agency in Bali. The part of my week that has nothing to do with visas goes into scrapers on Apify, and one of them reads the part of Meta Threads nobody screenshots: the conversation layer. What an account replies to other people, the replies under a given post, and what it reposts.

Posts are the easy half of a social network. Replies are where the data model fights you, because a reply on its own is not information. "Here you go! [https://dev.meta.ai](https://dev.meta.ai)" means nothing until you also have the question it answered.

Every number below is from runs I can point at: a 50-row replies run on 6 October 2026, the same input on 23 September, and a 7-row thread run I made today, 10 October 2026, on build 0.1.26.

Threads has a tab at `/@username/replies`. Visit it logged out and you do not see a list of that person's replies. You see a list of little conversations: somebody else's post, and underneath it the reply the account wrote. Threads serves it that way because it is the only way the tab reads as anything.

So the scraper has a choice, and both obvious answers are bad: drop the parent and the export is a column of orphan sentences, keep only the parent and you have lost the reply you came for.

I keep both and label them. Here is the real shape of a run with `maxPosts: 50` against `zuck` on 6 October:

| `source` | rows | 
|---|---|
| `replies` — the account's own replies | 25 | 
| `reply-context` — the post each one answers | 23 | 

Forty-eight rows, roughly half of them written by somebody else. (Twenty-three rather than twenty-five because two of those parents had already appeared earlier in the run and are de-duplicated.) Shrink the input and the pairing is plain to see — `maxPosts: 2` returns exactly this:

```
reply-context | merab.dvalishvili | isReply=false | "Always fun sparring with Mark 🦾⚔️"
replies       | zuck              | isReply=true  | "Always fun when you visit 🙏"
```

The practical consequence is about money and expectations, so I will say it bluntly: **`maxPosts` counts rows, not replies.** Ask for 50 and you get about 25 replies plus their context. If you only want the account's own words, filter `source == "replies"` after the fact — but you were billed for the parents, because fetching and delivering them is the work. Knowing that in advance is the difference between a budget and a surprise.

Every row also carries where it sits:

```
{
  "source": "replies",
  "threadId": "3960495719916521600",
  "positionInThread": 1,
  "threadLength": 2,
  "isReply": true,
  "replyToUsername": "merab.dvalishvili",
  "rootPostUsername": "merab.dvalishvili",
  "sourceUsername": "zuck"
}
```

`sourceUsername` is the account you asked about, not the author of the row — with ten usernames in one run it is the only field that tells you which input produced which pair.

Now the part that cost me a week of not noticing.

The logged-out replies tab embeds **four** conversations in the HTML it serves. Four. If you want the fifth you have to ask for more, and Threads has two ways of answering. There is the refetch query the site's own front end calls when you scroll — `BarcelonaProfileRepliesTabRefetchableDirectQuery` — which takes the variables already sitting in the page's preloaders and returns up to 25 conversations. And there is an older feed query that also returns a list of threads and looks, in the console, like the same thing.

It is not the same thing. Here is one row — same post id, same text, same like count — as the old feed returned it on 23 September and as the refetch query returns it now:

| field | legacy feed | refetch query | 
|---|---|---|
| `fullName` | `null` | `"Randi Zuckerberg"` | 
| `replyCount` | `null` | `12` | 
| `repostCount` /`quoteCount` /`shareCount` | `null` | `1` /`0` /`6` | 
| `replyControl` | `null` | `"everyone"` | 
| `rootPostUsername` | `null` | `"randizuckerberg"` | 
| `isReply` | **`false`** | **` true`** | 
| `likeCount` | `64` | `64` | 

`likeCount` matching is the control: this is not a number that drifted over two weeks, it is the same object served with less in it. The old endpoint returns a trimmed post — no `user.full_name`, no engagement counters beyond likes, and no `text_post_app_info.is_reply`.

That last one is the dangerous row in the table, and it is worth separating from the others. Six missing fields are honest: `null` tells you to go look. But `is_reply` absent from the payload becomes `Boolean(undefined)` → `false`, and `false` is a claim. A row that is a reply, in a dataset of replies, asserting that it is not a reply. Any pipeline that splits originals from replies on that flag silently got it wrong, and nothing in the output looked broken.

In the 23 September run, 42 rows came back and **34 of them were stripped** — the first 8 (the four conversations embedded in the HTML) were complete, everything after that was the degraded feed. The 6 October run on the same input: 48 rows, zero nulls in those fields. The fix was not clever: call the query the site itself calls, and keep the legacy feed strictly as a last resort.

If you scrape anything with a GraphQL front end: two endpoints returning the same object type do not return the same object. Diff them field by field on a row you can identify in both — and check the booleans, not just the fields you happen to read.

The other direction is a post URL in, conversation out. Today's run, verbatim from the log:

```
INFO  1 job(s) in thread mode: https://www.threads.net/@dikaiosvne/post/DYSUpHvm6lD
INFO  https://www.threads.net/@dikaiosvne/post/DYSUpHvm6lD: 7 rows (post + replies)
INFO  Done: {"posts":7,"pages":1,"requests":8,"retries":1,"bytes":2027542,
             "viewLookups":7,"viewBytes":1236844,"viewsCut":7,"errors":0}
```

One page gave the post and its replies — 2.0 MB of embedded JSON, no browser, no login. Root post and top reply:

```
thread | dikaiosvne | 177 likes | 125,753 views | "when do we get a muse spark api?"
reply  | zuck       | 640 likes | 130,081 views | "Here you go! https://dev.meta.ai"
```

The reply out-viewed the post it answered, which is Threads working exactly as designed and a decent argument for why reply data is worth pulling at all.

But look at `requests: 8` against `pages: 1`. Seven of those eight requests exist only for `viewCount`. Threads does not put view counts in the feed JSON it serves a logged-out visitor; the number is printed on each post's own page, inside a block keyed `BarcelonaLoggedOutExpansionGating`. So one request per row — and if you fetch each of those pages whole you pay for megabytes to read six digits. Instead the response body is read as a stream and the connection is cut the moment that block arrives: `viewsCut: 7` means all seven lookups aborted early, `viewBytes: 1236844` means they cost 1.2 MB between them rather than ~14 MB. Set `includeViews: false` and the run is one request and a few seconds; `viewCount` comes back `null` and nothing else changes. Those 45 seconds for 7 rows are almost entirely view lookups.

One wart I will own rather than paper over: thread mode fills `parentPostId` and `parentPostUrl` on each reply and leaves `threadId` / `positionInThread` null, while replies mode does the opposite. Both answer "what does this answer", through different fields, because they come from differently-shaped payloads. If you consume both modes, coalesce them.

[Threads Replies Scraper](https://apify.com/lergassy/threads-replies-scraper?fpr=wmeplu), one request, rows back:

```
curl -X POST "https://api.apify.com/v2/acts/lergassy~threads-replies-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"mode":"thread","postUrls":["https://www.threads.net/@dikaiosvne/post/DYSUpHvm6lD"],"maxRepliesPerPost":6}'
```

That exact call is the run above: 7 rows, $0.014, 45 seconds, at $0.002 a row with no start fee. `filterKeywords` is applied before billing, so pulling a brand's replies filtered to `refund`, `broken`, `delay` costs what the complaints cost and not what the account's whole reply history costs.

When it is the wrong tool, plainly. There is no reply search on Threads — you cannot ask for every reply mentioning your brand; you start from an account or a post. Private accounts contribute nothing. Deleted and hidden replies are simply absent, because Threads does not serve them to a visitor, so a reply count of 72 on a post does not promise you 72 rows. And very large conversations are paginated with a tail that is not always public: `maxRepliesPerPost` caps what you read, Threads caps what exists for a logged-out reader, and the smaller of the two wins.

If you take one thing from this: when you scrape a conversation, decide what your row *is* before you decide which fields it has. I got the fields right months before I got the unit right, and the unit is what the invoice counts.
