# How I used an AI agent to recover 18 years of a .NET user group's history

> Source: <https://dev.to/kapral1818/how-i-used-an-ai-agent-to-recover-18-years-of-a-net-user-groups-history-59pb>
> Published: 2026-10-02 21:01:36+00:00

I run [wrocnet.org](https://wrocnet.org/), the site of the Wrocław .NET user group. The group has been meeting since 2007, but its 18-year history was never kept in one place. I wanted a real archive: every meeting, talk, and speaker in one spot, backed by sources, and built to last. Getting there meant digging through old websites, forum posts, and archived pages, and I used an AI agent to do most of that digging. This post explains how I used the agent to research sources, how I verified its findings, and what the final numbers mean.

My research across archived websites, event pages, social media, email, and project records uncovered:

**Put AI agents to work in real software engineering.**

I help teams apply AI to research, coding, reviews, and automation, backed by deep experience in .NET architecture, performance, and production systems.

The problem was that the site itself had moved three times over the years, and each move left part of the history behind:

`wroc.net.isvclub.com`, the very first site, from November 2007 to early 2008` wrocnet.org` on Community Server, from 2008 to 2009
I found no surviving snapshot from late 2009 through most of 2010, so it is unclear what, if anything, the group used during that stretch.

*Every platform move carried some of the group's history forward and left some of it behind.*

None of these systems kept a full, clean copy of its own history, and nothing bridges them automatically. I assigned the repetitive work to the agent. It opened Wayback Machine snapshots, compared timestamps for the same URL, downloaded Meetup event pages, and checked one date against three unrelated sources before writing anything into a Jekyll post.

The repository already had a basic list of old meetings: numbers, dates, and one-line topics pulled from an [archived snapshot](https://web.archive.org/web/20241125161720/https://wrocnet.github.io/) of an earlier version of the site, `wrocnet.github.io`. That snapshot lists meetings like "17 Dec 2019 » 122. spotkanie - Azure Cognitive Services, Azure Sphere, Multi-tenant Azure," and it gave the repository its numbers, dates, and topics for roughly meetings 48 through 122, long before anyone went looking for speakers, agendas, or sources.

Once I compared file names, front matter numbers, and post content side by side, gaps appeared everywhere. Some meeting numbers had no post at all. Some posts had a date but no number. Some entries were a single sentence with no agenda or speaker. Until then, I had used the AI agent for coding. I now pointed it at the research itself: open a browser, use a terminal, and go through old pages one by one.

Across many sessions, the agent:

`curl` and parsed the embedded JSON with a small Node.js script to extract the event title, date, venue, and description reliably, instead of scraping visible HTML
None of these sources tells the whole story on its own. A gallery proves a meeting happened, but it says nothing about what was discussed there. A forum thread shows people were talking about a date; it does not confirm the meeting took place. Most of the work was judging how much to trust each source, instead of combining everything into one confident-sounding answer that might be wrong.

The research stayed useful only because we were strict about evidence. A few rules repeated across every session:

`date_estimated: true`), and only remove the flag once a stronger source confirms it.
These rules made the work slower, but they are the reason I can trust the result today.

A contributor, Tymoteusz Wojnarowski, used Claude Code independently to recover a large chunk of the earliest history in [PR #40](https://github.com/lukasz-pyrzyk/wrocnet.org/pull/40) (meetings and talks from 2007 to 2011, built from a speaker's personal site, old Wayback captures, and GoldenLine), [PR #41](https://github.com/lukasz-pyrzyk/wrocnet.org/pull/41) (16 meetings from the BlogEngine.NET era), and [PR #52](https://github.com/lukasz-pyrzyk/wrocnet.org/pull/52) (meetings 9, 11, and 34, plus two posts that got dropped in an earlier merge).

I continued the work myself with agent-assisted pull requests, recovering meetings 54 to 170, the earliest meetings (1 through 11), and the 20-to-32 stretch from the Twitter archive, among others:

GitHub's automated Copilot code review checked many of these pull requests alongside my own, repeatedly catching small but real issues: a misspelled street name, a missing archive link, and a technology name that should have been in bold on first mention.

The research leaned on a handful of source types, reused across almost every session:

`wrocnet.github.io` index
Writing code was the smaller part of this project. Most of the effort went into research: finding sources that disagreed with each other, deciding which one to trust, and writing "we do not know" when that was simply true. The agent could browse the web, fetch pages, and read structured data on its own, but it still needed me to set the rules for what counts as proof.

If you want to see the result, the archive is live at [wrocnet.org](https://wrocnet.org/). The full investigation is public too, pull request by pull request, on [GitHub](https://github.com/lukasz-pyrzyk/wrocnet.org/pulls).

*Illustrations in this post were generated with GPT.*
