{"slug": "how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-no", "title": "How to export any public Telegram channel to JSON and read it in English (no login, no bot token)", "summary": "A developer built a Python tool that exports any public Telegram channel to JSON without a login or bot token by scraping the t.me/s/<channel> web preview with httpx and BeautifulSoup, then paginating backwards via the ?before=<post id> parameter. The writeup documents five ways real channels broke the initial parser — reply blocks containing nested message text, video and sticker elements, abbreviated view counts like \"15.4K\", and HTTP 302 redirects for channels without a public preview — and describes adding LLM-based translation and per-post summaries batched into a single JSON response to cut costs.", "body_md": "A lot of useful information lives in Telegram channels: tech news, market commentary, local announcements. Much of it isn't in English. If you want to follow ten Russian or Ukrainian tech channels, copying posts into a translator one by one gets old fast.\n\nThe obvious options both have catches:\n\n`api_id` from my.telegram.org and a login with your phone number. That's more setup than a quick export deserves, and it ties the scraper to your personal account.\nThere's a third option that needs neither: every public channel has a web preview at `https://t.me/s/<channel>`. It's plain HTML, about 20 posts per page, and it pages backwards with a `?before=<post id>` parameter.\n\nI built a small tool on top of it. Here is how it works, the problems I ran into, and how I added AI translation and summaries.\n\n``` python\nimport re\nimport time\n\nimport httpx\nfrom bs4 import BeautifulSoup\n\ndef parse_page(html: str):\n    soup = BeautifulSoup(html, \"html.parser\")\n    posts = []\n    for msg in soup.select(\".tgme_widget_message[data-post]\"):\n        if \"service_message\" in msg.get(\"class\", []):\n            continue\n        text_el = msg.select_one(\".tgme_widget_message_text.js-message_text\")\n        if text_el:\n            for br in text_el.find_all(\"br\"):\n                br.replace_with(\"\\n\")\n        time_el = msg.select_one(\".tgme_widget_message_date time\")\n        views_el = msg.select_one(\".tgme_widget_message_views\")\n        posts.append({\n            \"id\": int(msg[\"data-post\"].split(\"/\")[1]),\n            \"url\": \"https://t.me/\" + msg[\"data-post\"],\n            \"date\": time_el[\"datetime\"] if time_el else None,\n            \"text\": text_el.get_text().strip() if text_el else \"\",\n            \"views\": views_el.get_text(strip=True) if views_el else None,  # e.g. \"15.4K\"\n        })\n    more = soup.select_one(\"link[rel=prev]\")  # href=\"https://dev.to/s/<channel>?before=<id>\"\n    before = re.search(r\"before=(\\d+)\", more[\"href\"]).group(1) if more else None\n    return posts, before\n\ndef fetch_channel(channel: str, limit: int = 100):\n    url, posts = f\"https://t.me/s/{channel}\", []\n    with httpx.Client(headers={\"User-Agent\": \"Mozilla/5.0\"}, follow_redirects=False) as client:\n        while url and len(posts) < limit:\n            resp = client.get(url)\n            if resp.status_code != 200:  # 302 = no public preview for this channel\n                break\n            page, before = parse_page(resp.text)\n            posts += sorted(page, key=lambda p: p[\"id\"], reverse=True)\n            url = f\"https://t.me/s/{channel}?before={before}\" if before else None\n            time.sleep(1)  # be polite\n    return posts[:limit]\n\nif __name__ == \"__main__\":\n    for post in fetch_channel(\"durov\", limit=5):\n        print(post[\"date\"], post[\"views\"], post[\"text\"][:80])\n```\n\n`pip install httpx beautifulsoup4` and it runs. Each page lists posts oldest-first, so I sort each page newest-first and then follow `before` to older pages.\n\nThe parser above worked on my first test page. Real channels then broke it in five ways.\n\n`.tgme_widget_message_reply` block, and that block also contains a `.tgme_widget_message_text`. Grab the first match and you save the `.js-message_text`, or skip anything inside the reply block.`<video>` is a video.`.tgme_widget_message_video_player`.` 15.4K` or `3.1M`, sometimes with a space as a thousands separator. Convert them before you sort or chart anything.`t.me/s/<channel>` redirects (HTTP 302) to the plain channel card. Treat that as \"skip\", not as an error.\nAlso: send requests at most about once a second, and keep `<br>` tags as line breaks, as the parser does. Otherwise multi-line posts come out as one long line.\n\nOnce posts are structured, an LLM can do the reading. For each post I ask for:\n\nA few things made this cheap and reliable:\n\n`{\"results\": [...]}` keyed by post id. The system prompt is paid once per batch instead of once per post.`response_format: {\"type\": \"json_object\"}` with HTTP 400. On a 400 I retry once without it, then strip code fences and take the outermost `{...}`.` retryDelay` field of the error body (`\"retryDelay\": \"34s\"`), and OpenAI-style APIs send a `Retry-After` header. Honor it, cap the wait, and after several 429s in a row turn AI off for the run instead of hanging.\nWith a small, fast model the AI part costs very roughly $0.20–1 per 1,000 posts, depending on the model and post length. Gemini's free tier costs nothing within its rate limits.\n\nI packaged all of this as an Actor on Apify: **[Telegram Channel Scraper + AI Translate & Summarize](https://apify.com/kassieiii/telegram-channel-scraper-ai)**. You paste channel names, optionally add your AI key, and download JSON, CSV or Excel, or send the results to Google Sheets, Make, Zapier or a webhook.\n\nBeyond the parser above, it handles:\n\nPricing is pay-per-result: $1 per 1,000 posts, plus $0.50 per 1,000 AI-enriched posts (your own key pays the tokens), plus $0.001 per run. Apify's free plan includes $5 a month to spend in the Store, which covers roughly 5,000 posts.\n\nIt also works from AI agents through Apify's MCP server, so you can ask Claude or Cursor to \"get the last 50 posts from @habr_com in English and summarize the main themes\".\n\n`BTC` or `IPO`, daily\nOnly public channels and only what they publish. No member lists, no private groups. Use the data in line with your local laws and Telegram's terms.\n\nIf you try it, I'd like to hear which channels and languages you use it on, and what breaks. I built it, so issues reach me directly.", "url": "https://wpnews.pro/news/how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-no", "canonical_source": "https://dev.to/kassieiii/how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-login-no-bot-token-582l", "published_at": "2026-10-06 12:42:18+00:00", "updated_at": "2026-10-06 12:48:31.429271+00:00", "lang": "en", "topics": ["ai-tools", "natural-language-processing", "developer-tools"], "entities": ["Telegram", "httpx", "BeautifulSoup"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-no", "markdown": "https://wpnews.pro/news/how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-no.md", "text": "https://wpnews.pro/news/how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-no.txt", "jsonld": "https://wpnews.pro/news/how-to-export-any-public-telegram-channel-to-json-and-read-it-in-english-no-no.jsonld"}}