Mick Donahueas part of
Write for Apify- a program for developers sharing original articles about what they've built with Apify.
I built a Reddit scraper, put it on Apify Store, and used it for months without trouble.
Then I exposed it to an AI agent as a tool and watched it return 25 rows about Minecraft, Xbox promos, and a wisdom teeth story.
Nothing errored. The Actor ran, returned data, reported success. The results were useless. Every reason for that traced back to a sentence my input schema never wrote down.
Prerequisites #
- An Apify account and API token. The free plan is $5 of monthly usage credit.
- An MCP-capable AI client. I used Claude Code. Claude Desktop, Cursor, and Windsurf behave the same.
- Optional, for the second half: an n8n instance.
- No Reddit API credentials. The Actor does not use them. It reads publicly visible pages, so check Reddit's User Agreement and robots.txt to confirm your use case is covered.
The Actor #
Reddit Scraper (labrat011/reddit-scraper) has four modes: subreddit posts, search, user profiles, and comment trees. Search mode is where this all happened.
The inputs that matter:
| Field | Type | Default |
|---|---|---|
| mode | enum | subreddit_posts , required |
| searchQuery | string | one phrase |
| searchQueriesList | array | batch several phrases |
| searchSort | enum | relevance |
| timeFilter | enum | week |
| maxResults | integer | 100 |
Remember searchQueriesList, searchSort, timeFilter, and maxResults. All four behave in ways I did not expect. I wrote all four.
Connecting the Actor to an agent over MCP #
Apify's MCP server exposes any Actor as a callable tool over the Model Context Protocol. No wrapper, no server to run. You point a client at an endpoint scoped to the Actors you want:
https://mcp.apify.com?tools=labrat011/reddit-scraper
One command in Claude Code:
claude mcp add --scope user --transport http apify \
"https://mcp.apify.com?tools=labrat011/reddit-scraper" \
--header "Authorization: Bearer YOUR_APIFY_TOKEN"
That is the whole integration. Which is the problem. The agent builds its own input from your input schema, and the schema is the only thing it has.
What I was hunting for #
Most Reddit lead-gen advice says monitor your product name and your competitors' names. That works and it is late. By the time somebody posts "Notion vs Obsidian" they have a shortlist and you are arguing about position on it.
The earlier signal is the complaint, before anyone names a product:
"Looking for an alternative to Nextcloud"
"Is there a tool that reconciles Stripe and Shopify payouts?"
So I searched for complaint language instead of product names. Six phrases.
Failure 1: Nobody told the agent to quote anything #
The first config looked fine:
{
"mode": "search",
"searchQueriesList": [
"is there a tool that",
"does anyone else hate",
"so frustrated with",
"looking for an alternative to",
"so sick of",
"anyone know a better way to"
],
"searchSort": "new",
"maxResults": 25
}
It succeeded. Twenty five rows. Two matched a phrase.
r/DeadlockTheGame. r/Dentists. r/almosthomeless. A letter to MAGA. The same Xbox promo crossposted four times inside sixty seconds.
Reddit does not phrase-match unquoted input. "Is there a tool that" is five common words, matched loosely. Pair that with searchSort: new and I was sampling whatever hit Reddit in the last fifteen minutes. The results looked random because they were.
Two characters fix it:
{ "searchQuery": "\"looking for an alternative to\"" }
Same Actor. Same workflow. Twenty five rows, twenty five matches.
Real people leaving products they pay for. r/selfhosted off Nextcloud after an upgrade deleted his files. r/webdev off WordPress after the ACF fight. r/CreditCards listing exactly what Citi broke.
Now the part that matters for agents.
A human hits this once, sees nonsense, tries quotes. An agent does not. It writes natural language, gets a success response with 25 rows attached, and reports the task done. No error. No warning. Nothing to notice.
My schema said:
"Keyword or phrase to search across Reddit."
It said "phrase". It never said Reddit will not treat it as one.
Failure 2: maxResults is a budget, not a limit #
Quoting fixed, I ran all six phrases with maxResults: 25. 24 rows came back. Each one matched the first phrase in the list.
Phrases 2 through 6 never ran.
maxResults is a total for the run, spent one query at a time. My own code:
async for item in scraper.scrape():
if count >= max_results:
break
One counter. One generator. Queries in order. The first query drained the budget, and the rest starved.
The schema never said so. An agent batching six phrases gets a plausible sheet, no error, and five-sixths of its request silently dropped.
Failure 3: timeFilter only fires on one sort value #
Quoting fixed relevance and broke freshness. relevance favors engagement. Engagement accumulates for years, so I was pulling on-topic posts from 2014.
I set timeFilter: "month". Nothing changed. Three of 25 rows landed inside the last month. Oldest was April 2014.
I assumed search mode ignored it. Then I read my own scraper:
if self.config.search_sort.value == "top":
params["t"] = self.config.time_filter.value
timeFilter is only sent when the sort is top. I had asked for relevance.
My description did say "Only applies when sort is 'Top'". It says sort. In search mode, the field is searchSort. I wrote both fields and still misread it. An agent had no chance.
Sort by top, and the filter fires:
{
"mode": "search",
"searchQueriesList": ["\"looking for an alternative to\""],
"searchSort": "top",
"timeFilter": "month",
"maxResults": 25
}
| Config | Rows inside last month | Oldest row |
|---|---|---|
| relevance, no timeFilter | 3 of 25 | 2014-04-22 |
| top + timeFilter: month | 24 of 24 | 2026-07-14 |
Precision drops slightly: 22 of 24 exact matches under top, against 25 of 25 under relevance. For lead gen, a bounded window is worth two points.
The fresh batch: r/Lightroom paying €15/month for Adobe and benchmarking ON1, DxO, and RawTherapee. r/degoogle off Google Docs. r/Piracy leaving WeMod over new pricing. r/Horticulture pricing a Dosatron replacement.
What I changed #
Four description rewrites and one editor change. No Python. Builds 1.3.6 through 1.3.14 in an afternoon.
- "description": "Keyword or phrase to search across Reddit."
+ "description": "Keyword or phrase to search across Reddit. IMPORTANT:
+ Reddit matches loose words by default, not phrases. To match an exact
+ phrase, wrap it in double quotes: \"looking for an alternative to\".
+ An unquoted multi-word query returns largely unrelated posts."
- "description": "Maximum number of results to return."
+ "description": "Maximum number of results for the whole run. This is a
+ shared total, not a per-query limit: when using searchQueriesList the
+ budget is consumed one query at a time."
- "description": "Time range when sorting by Top. Only applies when sort is 'Top'."
+ "description": "Time range filter. Applies ONLY when the sort is Top:
+ that means sort='top' in Subreddit Posts mode, or searchSort='top' in
+ Search mode. With any other sort value this field is accepted but has
+ no effect."
- "description": "How to sort search results."
+ "description": "How to sort search results. Relevance, the default,
+ favors highly upvoted posts, which are often several years old. For
+ recent results use Top together with timeFilter."
Describe the failure, not the field. "Time range filter" tells an agent nothing. "Accepted but has no effect unless sort is top" tells it everything. An agent cannot look at your results and judge them. The schema is the only feedback channel you have, and it fires before the mistake instead of after.
The fix that looked like it did not work #
searchQueriesList is an array of strings. Over MCP, my Actor advertised it as an array of objects.
I declared the type explicitly:
{
"searchQueriesList": { "type": "array", "items": { "type": "string" } }
}
Rebuilt. MCP still said object.
The platform derives item type from the Console editor, not from the declared schema. My three working array fields use editor: "stringList". This one used editor: "json". I switched it, which fixed the Console UX and removed the JSON escaping for humans.
Then I re-checked the endpoint minutes later and it still said object. I concluded an Actor author cannot fix this at all, and I nearly published that as a platform bug.
I was wrong. Checking again hours later, every array field reports string, including that one. The build had been correct the whole time. The schema the MCP server hands out lagged behind it.
If you build Actors for agents: read what your input schema becomes after conversion, because it is not always what you declared. Call tools/list on the endpoint and look at the JSON Schema your agent actually receives. Then give it time before concluding anything. In my case, the schema still showed the old type minutes after the build and was correct when I checked again later.
What happened when I asked an agent for leads #
Schema fixed. I gave an agent one sentence and no instructions on how to search:
Find me Reddit posts from the last month where people are looking to replace a paid tool they currently use.
Its first move:
"Search Reddit now. Batch exact-phrase queries, sort=top, timeFilter=month."
Exact-phrase quoting. top. A time filter. Nobody told it that. Those are the three things the rewritten descriptions teach, and the three things I got wrong by hand that same morning. I ran it twice from different directories to rule out context bleed from my own work. Same opening both times.
Then it did something my scheduled pipeline cannot. It read its own results and reformulated:
"Noise heavy. Filter rest of dataset for software/SaaS signal."
Ran again, different phrasing, reported back:
"Two runs done (search mode, top/month, residential proxy). 225 posts scraped, most noise. Real signal below."
225 posts, then a filtered table of people naming a product, a price, and a reason to leave. r/webdev paying $150/month for MapBox autocomplete. r/SEO on Semrush, "pricing hard to justify". r/CRMSoftware, Salesforce not justified for a small team.
It wrote its own caveats too. Excluded seven subreddits it judged to be vendor content marketing rather than real buyers. Flagged one post as a false positive after reading the author's own edit.
That is the real difference between the two paths, and it is not intelligence. My n8n workflow runs one query I picked in advance, forever. The agent ran twice, judged the first result, and threw out what looked like marketing. A pipeline cannot notice its results are wrong. An agent can, but only if the schema gets it to a good first attempt.
The finding I did not expect #
Three searches, one phrase each, quoted, same sort. Precision was 100% every time. Usefulness was not close.
| Phrase | What came back |
|---|---|
| "looking for an alternative to" | Nextcloud, WordPress, Adobe, Citi, Optus, Claude Code |
| "anyone know a better way to" | Warframe, Minecraft, Hay Day, Pokémon TCG, crochet |
| "is there a better way" | Elden Ring, Animal Crossing, TF2, grilling asparagus |
21 of 25 rows for the second phrase were gaming or hobby posts. The third was worse.
The difference is what the phrase presupposes. "Alternative to" presupposes an incumbent. Whoever types it already pays for something and is shopping. "Better way to" presupposes a chore. Usually unpaid. Frequently, a video game.
I built that phrase list believing frustration equals buying intent. The data says frustration equals engagement. Only the subset implying an existing paid product equals intent.
The agent found the same thing independently and put it better than I did:
Broad phrases ("cheaper alternative to") return mostly non-software: watercolor paper, fish, boots. Software-scoped phrases work better.
The same Actor on a schedule #
An agent calling a tool mid-conversation is one shape. A pipeline that runs whether anyone is watching is the other. I built that in n8n with Apify's community node, installed through n8n's community nodes system. The workflow is in the repo under examples/n8n:
Settings > Community nodes > Install > @apify/n8n-nodes-apify
The Apify node is a verified community node, so it isn't bundled with n8n's core package. n8n-nodes-base has 307 node packages and none is Apify, so a workflow referencing n8n-nodes-base.apify dies with "Unrecognized node type".
One trap: the Actor ID takes a tilde, not the slash the Store URL shows you.
labrat011~reddit-scraper -> 200 username~actor-name
dejMd0QoBemGH3zTn -> 200 the internal Actor ID
labrat011/reddit-scraper -> 404 the form the Store URL shows you
Either of the first two works. Paste the internal ID if you want it safe, there is no separator to get wrong, and it is what n8n stores once it resolves the Actor anyway.
Every finding above applies here identically. Same Actor, same schema, same traps. The schema fixes helped both paths at once.
What it costs #
The Origin column is the whole argument: n8n, CLI, and MCP hitting one Actor inside the same hour.
Runs cost $0.003 to $0.006 each at 3 to 25 results. A daily scheduled workflow runs under $2 a year. Pricing is $1.50 per 1,000 results plus $0.02 per Actor start.
Lessons #
The code was never the problem. Every failure came from a sentence the schema did not contain. Every fix was a description rewrite.
Success is not correctness. My Actor reported success on all 25 Minecraft rows. Exit code zero, data in the dataset, task complete. The only way to tell the difference was to read the rows.
Agents have one shot and no feedback. A human runs your Actor, looks at the output, adjusts. An agent constructs input once, gets a success response, and cannot judge whether the rows are any good. Anything ambiguous in your schema becomes a confident wrong answer.
Read the schema your agent receives, not the one you wrote. Call tools/list on the MCP endpoint and check what your input schema became after conversion. Then check it again later. Mine reported the wrong item type for hours after the fix had already shipped, and I nearly published that as a platform bug.
Document parameter interactions or they do not exist. A field silently ignored because of another field's value is the most expensive thing you can leave unwritten. timeFilter cost me a full run and a wrong assumption about my own code.
The Actor source, including every schema commit above, is on GitHub. The Actor is labrat011/reddit-scraper on Apify Store. The MCP endpoint is live if you want to point an agent at it yourself.