cd /news/ai-tools/how-to-export-any-public-telegram-ch… · home › topics › ai-tools › article
[ARTICLE · art-146019] src=dev.to ↗ pub= topic=ai-tools verified=true sentiment=· neutral

How to export any public Telegram channel to JSON and read it in English (no login, no bot token)

A developer built a Python tool that exports any public Telegram channel to JSON without a login or bot token by scraping the t.me/s/<channel> web preview with httpx and BeautifulSoup, then paginating backwards via the ?before=<post id> parameter. The writeup documents five ways real channels broke the initial parser — reply blocks containing nested message text, video and sticker elements, abbreviated view counts like "15.4K", and HTTP 302 redirects for channels without a public preview — and describes adding LLM-based translation and per-post summaries batched into a single JSON response to cut costs.

by read4 min views1 publishedOct 6, 2026

A lot of useful information lives in Telegram channels: tech news, market commentary, local announcements. Much of it isn't in English. If you want to follow ten Russian or Ukrainian tech channels, copying posts into a translator one by one gets old fast.

The obvious options both have catches:

api_id from my.telegram.org and a login with your phone number. That's more setup than a quick export deserves, and it ties the scraper to your personal account. There's a third option that needs neither: every public channel has a web preview at https://t.me/s/<channel>. It's plain HTML, about 20 posts per page, and it pages backwards with a ?before=<post id> parameter.

I built a small tool on top of it. Here is how it works, the problems I ran into, and how I added AI translation and summaries.

import re
import time

import httpx
from bs4 import BeautifulSoup

def parse_page(html: str):
    soup = BeautifulSoup(html, "html.parser")
    posts = []
    for msg in soup.select(".tgme_widget_message[data-post]"):
        if "service_message" in msg.get("class", []):
            continue
        text_el = msg.select_one(".tgme_widget_message_text.js-message_text")
        if text_el:
            for br in text_el.find_all("br"):
                br.replace_with("\n")
        time_el = msg.select_one(".tgme_widget_message_date time")
        views_el = msg.select_one(".tgme_widget_message_views")
        posts.append({
            "id": int(msg["data-post"].split("/")[1]),
            "url": "https://t.me/" + msg["data-post"],
            "date": time_el["datetime"] if time_el else None,
            "text": text_el.get_text().strip() if text_el else "",
            "views": views_el.get_text(strip=True) if views_el else None,  # e.g. "15.4K"
        })
    more = soup.select_one("link[rel=prev]")  # href="https://dev.to/s/<channel>?before=<id>"
    before = re.search(r"before=(\d+)", more["href"]).group(1) if more else None
    return posts, before

def fetch_channel(channel: str, limit: int = 100):
    url, posts = f"https://t.me/s/{channel}", []
    with httpx.Client(headers={"User-Agent": "Mozilla/5.0"}, follow_redirects=False) as client:
        while url and len(posts) < limit:
            resp = client.get(url)
            if resp.status_code != 200:  # 302 = no public preview for this channel
                break
            page, before = parse_page(resp.text)
            posts += sorted(page, key=lambda p: p["id"], reverse=True)
            url = f"https://t.me/s/{channel}?before={before}" if before else None
            time.sleep(1)  # be polite
    return posts[:limit]

if __name__ == "__main__":
    for post in fetch_channel("durov", limit=5):
        print(post["date"], post["views"], post["text"][:80])

pip install httpx beautifulsoup4 and it runs. Each page lists posts oldest-first, so I sort each page newest-first and then follow before to older pages.

The parser above worked on my first test page. Real channels then broke it in five ways.

.tgme_widget_message_reply block, and that block also contains a .tgme_widget_message_text. Grab the first match and you save the .js-message_text, or skip anything inside the reply block.<video> is a video..tgme_widget_message_video_player. 15.4K or 3.1M, sometimes with a space as a thousands separator. Convert them before you sort or chart anything.t.me/s/<channel> redirects (HTTP 302) to the plain channel card. Treat that as "skip", not as an error. Also: send requests at most about once a second, and keep <br> tags as line breaks, as the parser does. Otherwise multi-line posts come out as one long line.

Once posts are structured, an LLM can do the reading. For each post I ask for:

A few things made this cheap and reliable:

{"results": [...]} keyed by post id. The system prompt is paid once per batch instead of once per post.response_format: {"type": "json_object"} with HTTP 400. On a 400 I retry once without it, then strip code fences and take the outermost {...}. retryDelay field of the error body ("retryDelay": "34s"), and OpenAI-style APIs send a Retry-After header. Honor it, cap the wait, and after several 429s in a row turn AI off for the run instead of hanging. With a small, fast model the AI part costs very roughly $0.20–1 per 1,000 posts, depending on the model and post length. Gemini's free tier costs nothing within its rate limits.

I packaged all of this as an Actor on Apify: Telegram Channel Scraper + AI Translate & Summarize. You paste channel names, optionally add your AI key, and download JSON, CSV or Excel, or send the results to Google Sheets, Make, Zapier or a webhook.

Beyond the parser above, it handles:

Pricing is pay-per-result: $1 per 1,000 posts, plus $0.50 per 1,000 AI-enriched posts (your own key pays the tokens), plus $0.001 per run. Apify's free plan includes $5 a month to spend in the Store, which covers roughly 5,000 posts.

It also works from AI agents through Apify's MCP server, so you can ask Claude or Cursor to "get the last 50 posts from @habr_com in English and summarize the main themes".

BTC or IPO, daily Only public channels and only what they publish. No member lists, no private groups. Use the data in line with your local laws and Telegram's terms.

If you try it, I'd like to hear which channels and languages you use it on, and what breaks. I built it, so issues reach me directly.

── more in #ai-tools 4 stories · sorted by recency
── more on @telegram 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/how-to-export-any-pu…] indexed:0 read:4min 2026-10-06 · —