A lot of useful information lives in Telegram channels: tech news, market commentary, local announcements. Much of it isn't in English. If you want to follow ten Russian or Ukrainian tech channels, copying posts into a translator one by one gets old fast.
The obvious options both have catches:
api_id from my.telegram.org and a login with your phone number. That's more setup than a quick export deserves, and it ties the scraper to your personal account.
There's a third option that needs neither: every public channel has a web preview at https://t.me/s/<channel>. It's plain HTML, about 20 posts per page, and it pages backwards with a ?before=<post id> parameter.
I built a small tool on top of it. Here is how it works, the problems I ran into, and how I added AI translation and summaries.
import re
import time
import httpx
from bs4 import BeautifulSoup
def parse_page(html: str):
soup = BeautifulSoup(html, "html.parser")
posts = []
for msg in soup.select(".tgme_widget_message[data-post]"):
if "service_message" in msg.get("class", []):
continue
text_el = msg.select_one(".tgme_widget_message_text.js-message_text")
if text_el:
for br in text_el.find_all("br"):
br.replace_with("\n")
time_el = msg.select_one(".tgme_widget_message_date time")
views_el = msg.select_one(".tgme_widget_message_views")
posts.append({
"id": int(msg["data-post"].split("/")[1]),
"url": "https://t.me/" + msg["data-post"],
"date": time_el["datetime"] if time_el else None,
"text": text_el.get_text().strip() if text_el else "",
"views": views_el.get_text(strip=True) if views_el else None, # e.g. "15.4K"
})
more = soup.select_one("link[rel=prev]") # href="https://dev.to/s/<channel>?before=<id>"
before = re.search(r"before=(\d+)", more["href"]).group(1) if more else None
return posts, before
def fetch_channel(channel: str, limit: int = 100):
url, posts = f"https://t.me/s/{channel}", []
with httpx.Client(headers={"User-Agent": "Mozilla/5.0"}, follow_redirects=False) as client:
while url and len(posts) < limit:
resp = client.get(url)
if resp.status_code != 200: # 302 = no public preview for this channel
break
page, before = parse_page(resp.text)
posts += sorted(page, key=lambda p: p["id"], reverse=True)
url = f"https://t.me/s/{channel}?before={before}" if before else None
time.sleep(1) # be polite
return posts[:limit]
if __name__ == "__main__":
for post in fetch_channel("durov", limit=5):
print(post["date"], post["views"], post["text"][:80])
pip install httpx beautifulsoup4 and it runs. Each page lists posts oldest-first, so I sort each page newest-first and then follow before to older pages.
The parser above worked on my first test page. Real channels then broke it in five ways.
.tgme_widget_message_reply block, and that block also contains a .tgme_widget_message_text. Grab the first match and you save the .js-message_text, or skip anything inside the reply block.<video> is a video..tgme_widget_message_video_player. 15.4K or 3.1M, sometimes with a space as a thousands separator. Convert them before you sort or chart anything.t.me/s/<channel> redirects (HTTP 302) to the plain channel card. Treat that as "skip", not as an error.
Also: send requests at most about once a second, and keep <br> tags as line breaks, as the parser does. Otherwise multi-line posts come out as one long line.
Once posts are structured, an LLM can do the reading. For each post I ask for:
A few things made this cheap and reliable:
{"results": [...]} keyed by post id. The system prompt is paid once per batch instead of once per post.response_format: {"type": "json_object"} with HTTP 400. On a 400 I retry once without it, then strip code fences and take the outermost {...}. retryDelay field of the error body ("retryDelay": "34s"), and OpenAI-style APIs send a Retry-After header. Honor it, cap the wait, and after several 429s in a row turn AI off for the run instead of hanging.
With a small, fast model the AI part costs very roughly $0.20–1 per 1,000 posts, depending on the model and post length. Gemini's free tier costs nothing within its rate limits.
I packaged all of this as an Actor on Apify: Telegram Channel Scraper + AI Translate & Summarize. You paste channel names, optionally add your AI key, and download JSON, CSV or Excel, or send the results to Google Sheets, Make, Zapier or a webhook.
Beyond the parser above, it handles:
Pricing is pay-per-result: $1 per 1,000 posts, plus $0.50 per 1,000 AI-enriched posts (your own key pays the tokens), plus $0.001 per run. Apify's free plan includes $5 a month to spend in the Store, which covers roughly 5,000 posts.
It also works from AI agents through Apify's MCP server, so you can ask Claude or Cursor to "get the last 50 posts from @habr_com in English and summarize the main themes".
BTC or IPO, daily
Only public channels and only what they publish. No member lists, no private groups. Use the data in line with your local laws and Telegram's terms.
If you try it, I'd like to hear which channels and languages you use it on, and what breaks. I built it, so issues reach me directly.