cd /news/ai-agents/aliexpress-scraper-returns-zero-reco… · home › topics › ai-agents › article
[ARTICLE · art-147598] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=· neutral

AliExpress Scraper Returns Zero Records on CSR Pages from `ja` or `ko`

A developer documented that the AliExpress Scraper Actor returns zero records when an AI agent requests regional storefronts such as `ja` (Japan) or `ko` (Korea), because those pages serve client-side rendered HTML that omits the `_init_data_` JSON state the Actor parses. The writeup, covering integration of the Actor as an agent tool via Apify's MCP endpoint at `https://mcp.apify.com` with the `?tools=owner/actor-name` parameter, recommends agents restrict requests to regions known to serve server-side rendered JSON, including `com`, `us`, `de`, `fr`, `es`, `it`, `nl`, `pt`, `pl`, `tr`, `ar`, `vi`, `th`, and `id`.

by read17 min views4 publishedOct 8, 2026

Large language models (LLMs) are increasingly capable of acting as autonomous agents, performing complex tasks by breaking them down into sub-problems and calling external tools. Exposing Apify Actors as callable tools for these agents presents specific integration challenges and opportunities, particularly around how input schemas translate into tool signatures and the predictable limits an agent will encounter.

This article details how an AI agent can invoke the AliExpress Scraper via Apify's Managed Cloud Platform (MCP), focusing on the critical ?tools= configuration, the derived input schema for an agent, and the operational limitations an agent must be designed to handle for robust operation.

If an AI agent requests a regional storefront that occasionally serves a client-side rendered (CSR) page, such as ja or ko, the AliExpress Scraper Actor will return zero records. This occurs because these pages omit the _init_data_ JSON state that the Actor relies on for parsing. For reliable data extraction, agents should prioritize regions known to consistently serve server-side rendered (SSR) JSON, such as com, us, de, fr, es, it, nl, pt, pl, tr, ar, vi, th, or id.

The aliexpress-scraper Actor is designed to extract structured data from AliExpress, offering various modes for querying. When integrating this Actor as an AI agent tool, the agent needs to understand not just the functionality, but also the nuances of its failure modes. A particularly sharp edge arises with certain regional storefronts. The Actor's region input field dictates which AliExpress subdomain to query. While many regions reliably provide server-side rendered (SSR) HTML containing embedded JSON for data extraction, some, like ja (Japan) or ko (Korea), can occasionally serve a client-side rendered (CSR) page. This means the critical _init_data_ JSON state, which the Actor uses for parsing, is entirely absent from the initial HTML.

When this happens, the Actor run will complete successfully, but it will yield zero records in its dataset. An AI agent, unaware of this specific limitation, might interpret this as a successful run with no matching data, rather than a parsing failure due to a rendering discrepancy. To mitigate this, an agent should be programmed to:

com, us, or de which are explicitly documented to ship SSR JSON reliably.autoEscalateOnBlock carefully Here's how an agent might select a region, preferring reliable ones:

def select_aliexpress_region(desired_region: str) -> str:
    """
    Selects an AliExpress region, prioritizing known reliable SSR storefronts.
    """
    reliable_regions = ["com", "us", "de", "fr", "es", "it", "nl", "pt", "pl", "tr", "ar", "vi", "th", "id"]
    if desired_region in reliable_regions:
        return desired_region
    elif desired_region in ["ja", "ko", "he", "ru"]: # Known problematic or special cases
        print(f"Warning: Region '{desired_region}' may serve CSR-only pages, potentially returning zero records.")
        if desired_region == "ru":
            print("Note: For Russian language results, use region='com' with language='ru_RU'.")
            return "com" # Redirect internally to .com with language override
        return desired_region # Still allow, but with warning
    else:
        print(f"Unknown region '{desired_region}', defaulting to 'com'.")
        return "com"

agent_desired_region = "ja"
actual_region_to_use = select_aliexpress_region(agent_desired_region)
print(f"Agent will attempt to scrape region: {actual_region_to_use}")

You expose an Apify Actor as an AI agent tool by calling the Apify MCP server at https://mcp.apify.com and specifying the Actor using the ?tools=owner/actor-name query parameter. This endpoint translates the Actor's input schema into a tool signature that an agent can understand and execute, provided the agent has the necessary Apify API token for authenticated calls.

The core of enabling AI agents to use Apify Actors lies in Apify's Managed Cloud Platform (MCP). The MCP server acts as an intermediary, presenting Actors as structured tools. For the aliexpress-scraper Actor, the endpoint https://mcp.apify.com?tools=crawlerbros/aliexpress-scraper is the entry point. When an agent queries this endpoint, MCP responds with a structured description of the aliexpress-scraper tool, derived directly from its input schema. This description includes the tool's name, a natural language description, and most critically, its parameter schema, which is a JSON Schema representation of the Actor's input.

An AI agent's orchestration logic would then parse this schema to understand what arguments (mode, searchQuery, region, etc.) the aliexpress-scraper tool expects, their types, and any constraints or default values. Executing the tool involves making a POST request to the MCP server with the appropriate X-Apify-Api-Token header and a JSON body corresponding to the Actor's input.

Here's a simplified representation of how the aliexpress-scraper input schema translates into a tool signature for an agent:

{
  "name": "aliexpress-scraper",
  "description": "Scrape AliExpress search results, product details, store profiles, and customer reviews. Multi-region (com / us / ru / es / fr / de / it / nl / pt / pl / ar / tr / ko / ja / vi / th / id / he), multi-currency, with sort, price, rating, and ship-from/ship-to filters.",
  "input_schema": {
    "type": "object",
    "properties": {
      "mode": {
        "type": "string",
        "description": "What to scrape. search: text-query results. byProduct: product detail by ID. byStore: store profile by ID. byReviews: customer reviews for product IDs. byUrl: parse any AliExpress",
        "enum": ["search", "byProduct", "byStore", "byReviews", "byUrl"],
        "default": "search"
      },
      "searchQuery": {
        "type": "string",
        "description": "Free-text search query, e.g. \"phone case\". Required when mode=search."
      },
      "productIds": {
        "type": "array",
        "items": { "type": "string" },
        "description": "AliExpress numeric product IDs (e.g. 1005010155028387)."
      },
      "region": {
        "type": "string",
        "description": "AliExpress regional sub-domain to query.",
        "default": "com"
      },
      "currency": {
        "type": "string",
        "description": "ISO-4217 currency code. AliExpress may override it with the proxy IP's local currency.",
        "default": "USD"
      },
      "language": {
        "type": "string",
        "description": "Locale that AliExpress should render text in (b_locale cookie).",
        "default": "en_US"
      },
      "priceMin": {
        "type": "integer",
        "description": "Minimum price (in the selected currency). 0 disables.",
        "default": 0
      },
      "maxItems": {
        "type": "integer",
        "description": "Hard cap on emitted records.",
        "default": 25
      },
      "maxPages": {
        "type": "integer",
        "description": "Maximum pagination pages per search query (60 items per page).",
        "default": 3
      }
    },
    "required": ["mode"]
  }
}

This structured tool description allows the agent to dynamically construct valid requests. However, it's crucial for the agent to also understand the implications of default values versus explicitly passed values. The prefill values specified in the Actor's console UI are not applied to API calls; only the default values in the schema are. An agent should always construct a complete input dictionary, explicitly setting all necessary parameters, to avoid relying on implicit server-side prefill behavior that won't be triggered by an API call.

An AI agent should switch from synchronous to asynchronous Actor execution when a single run is expected to exceed 300 seconds. The synchronous run endpoint for Apify Actors has a hard-coded timeout of 300 seconds (5 minutes); exceeding this limit will result in an HTTP 408 response, forcing the agent to restart or abandon the task. For longer-running tasks, the agent must initiate the Actor run via a POST request to /v2/acts/<actor>/runs and then poll the run's status or use a webhook.

AI agents, particularly when interacting with external services, need to be mindful of execution limits. For Apify Actors, a critical threshold is the 300-second (5-minute) timeout for synchronous runs. If an agent attempts to call an Actor using the synchronous endpoint and the Actor runs for longer than 300 seconds, the connection will be terminated with an HTTP 408 error. This isn't a graceful shutdown; the Actor will continue running on Apify's platform, but the agent's connection will be lost, making it impossible to retrieve the results synchronously.

For tasks that inherently take longer – such as scraping thousands of AliExpress products across many search pages, or fetching reviews for a large list of product IDs – an AI agent must opt for asynchronous execution. This involves:

https://api.apify.com/v2/acts/crawlerbros/aliexpress-scraper/runs (with an API token). This returns a runId. https://api.apify.com/v2/actor-runs/<runId> to check the status field.SUCCEEDED or FAILED, retrieve the run's dataset items via the dataset URL provided in the run object. This asynchronous pattern adds complexity but is essential for robust operation. An agent's decision-making logic should incorporate an estimate of run duration based on input parameters (e.g., maxItems, maxPages, number of productIds). If these parameters suggest a run might approach or exceed the 300-second threshold, the agent should proactively choose the asynchronous approach. For instance, maxPages=3 for mode=search implies up to 180 items (60 items/page), which is unlikely to hit the limit. However, if maxItems is set to a much higher number, or productIds contains hundreds of entries, an asynchronous execution strategy becomes imperative.

import requests
import time

APIFY_API_TOKEN = "YOUR_APIFY_API_TOKEN"
ACTOR_ID = "crawlerbros/aliexpress-scraper"

def run_aliexpress_scraper_async(run_input: dict):
    """
    Initiates an asynchronous run of the AliExpress Scraper Actor and polls for its completion.
    """
    print("Initiating asynchronous Actor run...")
    start_response = requests.post(
        f"https://api.apify.com/v2/acts/{ACTOR_ID}/runs",
        headers={"Authorization": f"Bearer {APIFY_API_TOKEN}"},
        json=run_input
    ).json()

    run_id = start_response["data"]["id"]
    print(f"Actor run started with ID: {run_id}")

    dataset_id = start_response["data"]["defaultDatasetId"]
    print(f"Default Dataset ID: {dataset_id}")

    while True:
        run_status_response = requests.get(
            f"https://api.apify.com/v2/actor-runs/{run_id}",
            headers={"Authorization": f"Bearer {APIFY_API_TOKEN}"}
        ).json()
        status = run_status_response["data"]["status"]
        print(f"Current run status: {status}")

        if status in ["SUCCEEDED", "FAILED", "ABORTED"]:
            print(f"Run finished with status: {status}")
            break
        time.sleep(10) # Wait 10 seconds before polling again

    if status == "SUCCEEDED":
        results_url = f"https://api.apify.com/v2/datasets/{dataset_id}/items"
        print(f"Retrieving results from: {results_url}")
        results = requests.get(results_url, headers={"Authorization": f"Bearer {APIFY_API_TOKEN}"}).json()
        print(f"Retrieved {len(results)} items.")
        return results
    else:
        print("Actor run did not succeed.")
        return []

long_run_input = {
  "mode": "search",
  "searchQuery": "laptop bag",
  "region": "com",
  "maxPages": 10, # Max items could be 600, potentially a longer run
  "maxItems": 500
}

maxTotalChargeUsd parameter protect against unexpected costs? The maxTotalChargeUsd parameter protects against unexpected costs by setting an upper limit on the monetary charge for an Actor run. When the total cost incurred by the run reaches this cap, the Actor run is terminated. This provides a crucial safeguard for AI agents operating with budget constraints, although termination is not instantaneous, and some resources may still be consumed briefly after the cap is tripped.

Cost control is a critical aspect of autonomous agent operation, especially when calling external APIs or Actors that incur charges. The aliexpress-scraper Actor, like many on Apify, charges based on specific events. To prevent an agent from inadvertently spending beyond a defined budget, the maxTotalChargeUsd query parameter (or ACTOR_MAX_TOTAL_CHARGE_USD environment variable within the Actor) is invaluable.

An AI agent, before initiating a potentially large scrape, can estimate the likely cost based on the number of expected output records and the per-result price. It can then set maxTotalChargeUsd to a value that reflects its allocated budget for that specific task. If, for any reason (e.g., more results found than expected, higher proxy usage), the actual cost begins to approach this limit, the Apify platform will automatically terminate the run.

It's crucial for the agent to understand that this termination is not immediate. The Actor will continue to consume resources for a brief period while it's winding down. Therefore, setting maxTotalChargeUsd slightly above the absolute hard limit can provide a small buffer. This mechanism allows for fine-grained financial control and prevents runaway costs in scenarios where an agent's logic might accidentally request an overly expansive scrape.

Consider a scenario where an agent is instructed to find "cheap phone cases" but due to a misinterpretation, it searches for a very generic term and sets maxItems extremely high. Without maxTotalChargeUsd, this could lead to a significant bill. With it, the agent ensures that even if its logic falters, the financial impact is contained.

import requests

APIFY_API_TOKEN = "YOUR_APIFY_API_TOKEN"
ACTOR_ID = "crawlerbros/aliexpress-scraper"

def run_aliexpress_scraper_with_budget(run_input: dict, max_charge_usd: float):
    """
    Runs the AliExpress Scraper Actor with a maximum total charge limit.
    """
    print(f"Initiating Actor run with max charge: ${max_charge_usd}")
    start_response = requests.post(
        f"https://api.apify.com/v2/acts/{ACTOR_ID}/runs?maxTotalChargeUsd={max_charge_usd}",
        headers={"Authorization": f"Bearer {APIFY_API_TOKEN}"},
        json=run_input
    ).json()

    print(f"Run started: {start_response['data']['id']}")
    print("Agent should poll run status and check for 'ABORTED' due to budget.")

budget_constrained_input = {
  "mode": "search",
  "searchQuery": "smartwatch",
  "region": "us",
  "maxItems": 1000 # Potentially many results, but capped by budget
}

The AliExpress Scraper charges for two events: a "result" event at $0.005 per single item in the default dataset, and an "Actor Start" event at $0.005 per GB of memory allocated to the run. These event charges are in addition to platform usage fees for run time and memory, which are billed separately at the user's Apify plan rates.

The number of "result" events scales directly with maxItems, maxPages, and maxReviewsPerProduct, with discounts available for FREE, BRONZE, SILVER, GOLD, PLATINUM, and DIAMOND tier users.

Understanding the cost structure is crucial for any AI agent that aims to operate efficiently and predictably. The aliexpress-scraper Actor uses a PAY_PER_EVENT pricing model.

Here's a breakdown:

$0.005. This is the most significant variable cost. mode=search: The maxItems and maxPages inputs directly multiply the number of potential result events. Since each page can yield up to 60 items, maxPages=3 could result in up to 180 product records, each triggering a result event.mode=byProduct: Each ID in productIds typically fetches one product record.mode=byReviews: Each ID in maxReviewsPerProduct review records. If enrichReviews is true for An intelligent agent can use these details to make informed decisions. For example, if the goal is to get a brief overview of products, setting maxItems to a low number significantly reduces the "result" event cost. If only product details are needed, mode=byProduct with a specific list of productIds might be more cost-effective than a broad search query followed by post-processing.

def estimate_aliexpress_scraper_event_cost(run_input: dict, user_tier: str = "FREE") -> float:
    """
    Estimates the event-based cost for an AliExpress Scraper run.
    Does not account for platform usage.
    """
    result_price_per_item = {
        "FREE": 0.005,
        "BRONZE": 0.00433,
        "SILVER": 0.00367,
        "GOLD": 0.003,
        "PLATINUM": 0.003,
        "DIAMOND": 0.003
    }.get(user_tier.upper(), 0.005) # Default to FREE tier if unknown

    estimated_results = 0
    if run_input.get("mode") == "search":
        max_items = run_input.get("maxItems", 25)
        max_pages = run_input.get("maxPages", 3)
        estimated_products = min(max_items, max_pages * 60)
        estimated_results += estimated_products

        if run_input.get("enrichReviews", False):
            max_reviews_per_product = run_input.get("maxReviewsPerProduct", 0)
            if max_reviews_per_product > 0:
                estimated_results += estimated_products * max_reviews_per_product

    elif run_input.get("mode") == "byProduct":
        estimated_results += len(run_input.get("productIds", []))

    elif run_input.get("mode") == "byReviews":
        max_reviews_per_product = run_input.get("maxReviewsPerProduct", 0)
        if max_reviews_per_product > 0:
            estimated_results += len(run_input.get("productIds", [])) * max_reviews_per_product

    elif run_input.get("mode") == "byStore":
        estimated_results += len(run_input.get("storeIds", []))

    elif run_input.get("mode") == "byUrl":
        estimated_results += len(run_input.get("urls", [])) # Simplified assumption, depends on URL type

    actor_start_cost = 0.005

    total_event_cost = (estimated_results * result_price_per_item) + actor_start_cost
    return total_event_cost

example_input_search = {
  "mode": "search",
  "searchQuery": "bluetooth speaker",
  "maxItems": 100,
  "maxPages": 2,
  "enrichReviews": True,
  "maxReviewsPerProduct": 5
}

An AI agent using the AliExpress Scraper must account for several limitations: client-side rendering issues for product details and some regional storefronts (potentially leading to empty records), best-effort currency localization (AliExpress may override requested currency based on proxy IP), and the absence of variation/SKU details or shipping costs due to API call signing requirements. Aggressive crawls can also be rate-limited, requiring explicit proxy usage.

Beyond the cost structure and execution model, an AI agent needs a deep understanding of the aliexpress-scraper's functional limitations to avoid generating incorrect assumptions or making repeated, futile requests. These limitations are intrinsic to how AliExpress renders its content and how the Actor can access it:

aliexpress-scraper primarily relies on server-side rendered (SSR) JSON embedded in the HTML. This means: byProduct yields less detail.ja, ko, he) sometimes serve CSR-only pages, causing runs to return zero records. An agent should be programmed to detect this failure mode and potentially re-attempt with a more reliable region or language setting (e.g., language=ru_RU with region=com for Russian content).currency input is sent, AliExpress might override it based on the proxy IP's geolocation. For deterministic currency results, an agent must explicitly enable useProxy=true and configure a residential proxy in the desired country, incurring additional proxy costs and potentially shorter proxy session lifespans (~30 minutes for residential).true) helps by switching to residential proxies, but for sustained, high-volume scraping, an agent should proactively enable Unnamed storage expirations significantly impact data retention for AI agents, as data from default datasets, request queues, and key-value stores from older runs will be automatically deleted. On the free Apify plan, only the 10 most recent runs are retained for four months. To ensure long-term data persistence, AI agents must be configured to use named storages, which are exempt from automatic deletion.

For AI agents managing ongoing data collection or requiring access to historical data, the Apify platform's storage policies are a critical design consideration. By default, Actor runs create "unnamed" storages. This includes the default dataset where the aliexpress-scraper outputs its records, any request queues it uses, and temporary key-value stores. These unnamed storages are subject to expiration. Specifically, on the free plan, only the data from the 10 most recent runs is retained, and even then, only for four months. Older data, or data from runs beyond the 10-run limit, is automatically purged.

This has direct implications for an AI agent's operational strategy:

output/datasetId (or requestQueueId, keyValueStoreId) with a An agent that periodically scrapes AliExpress data for market analysis, for example, would ideally create a named dataset (e.g., my-product-trends-2026) and instruct the aliexpress-scraper to output its data there for every run. This ensures that even if hundreds of runs occur over months, the collected data remains accessible.

import requests

APIFY_API_TOKEN = "YOUR_APIFY_API_TOKEN"
ACTOR_ID = "crawlerbros/aliexpress-scraper"

def run_aliexpress_scraper_with_named_dataset(run_input: dict, dataset_name: str):
    """
    Initiates an Actor run, pushing results to a named dataset for persistence.
    The dataset will be created if it doesn't exist.
    """
    print(f"Ensuring named dataset '{dataset_name}' exists or is created...")
    dataset_info = requests.post(
        f"https://api.apify.com/v2/datasets?token={APIFY_API_TOKEN}",
        json={"name": dataset_name}
    ).json()
    named_dataset_id = dataset_info["data"]["id"]
    print(f"Using dataset ID: {named_dataset_id}")

    run_input_with_dataset = run_input.copy()
    run_input_with_dataset["output/datasetId"] = named_dataset_id

    print(f"Initiating Actor run with output to named dataset '{dataset_name}'...")
    start_response = requests.post(
        f"https://api.apify.com/v2/acts/{ACTOR_ID}/runs",
        headers={"Authorization": f"Bearer {APIFY_API_TOKEN}"},
        json=run_input_with_dataset
    ).json()

    run_id = start_response["data"]["id"]
    print(f"Actor run started with ID: {run_id}")

tracking_input = {
  "mode": "search",
  "searchQuery": "phone case",
  "region": "com",
  "maxItems": 10
}

Proxy session limitations can affect long-running AliExpress Scraper tasks because datacenter proxies persist for 26 hours, while residential proxies typically last around 30 minutes. If an AI agent uses residential proxies for sustained scraping due to rate limiting or currency localization needs, it must be prepared to handle frequent proxy rotation and potential session drops during extended runs.

When an AI agent uses the aliexpress-scraper for tasks that involve a significant number of requests or require specific IP geo-locations, the choice and management of proxies become crucial. AliExpress can rate-limit aggressive crawls, and autoEscalateOnBlock provides a fallback to residential proxies. However, there are important distinctions in proxy session longevity that an agent must consider:

If an AI agent enables useProxy=true with a residential proxy group from the outset, or if autoEscalateOnBlock triggers the use of residential proxies, the agent's logic needs to be robust enough to handle these proxy rotations. While the Actor itself is designed to manage proxy changes internally, the agent should be aware that the egress IP address used for requests will change often, which can impact consistent currency localization if AliExpress relies heavily on the IP's geolocation. For truly deterministic currency, the agent must ensure a residential proxy from the exact target country is selected, and acknowledge that even then, the underlying proxy IP will rotate.

This limitation means that an agent cannot assume a stable IP identity for the duration of a multi-hour scrape if residential proxies are involved. It might need to implement logic to detect if proxy-dependent features (like specific currency localization) become inconsistent and potentially adjust its strategy, such as restarting segments of the scrape or issuing warnings.

Checked against the Actor's input schema and Apify docs on 2026-10-08.

The Actor's README is the source of truth for its inputs, outputs and limits. Need a hand wiring this into your stack? Email info@crawlerbros.com

── more in #ai-agents 4 stories · sorted by recency
── more on @aliexpress 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/aliexpress-scraper-r…] indexed:0 read:17min 2026-10-08 · —