Hackathons in the vibe coding era: our token economics experiment Quesma hosted its first token economics hackathon, Session #5 of the Vibe Coding Summer Jam series, where five student teams analyzed 18 GB of coding-agent transcripts over one Friday evening to find where tokens are being wasted. Participants received up to $100 each in OpenRouter credits from Quesma and worked with SWE-chat data gathered at Stanford from open-source contributors using tools including Claude Code, Codex, Cursor, and Pi. The team "Palantir" found that expressive or rude language in corrections helps agent performance, while "Hackengersi" found inconclusive results on whether chaotic, drunk-style prompts affect output. Students had one Friday evening, 18 GB of coding-agent transcripts, and a single question: where are we burning tokens? They investigated wasted tokens, agents stuck in loops, and whether swearing at AI changes its behavior. This was Quesma’s first token economics hackathon. My peak hackathon years: 2011-2012 Back then, building an application overnight required pragmatic choices and skills. You needed a manageable idea, a team working well together, and tools that let you move quickly. Meteor, PubNub, or Twilio were our secret weapons. At pre-IPO Facebook in the summer of 2011, hackathons for interns were every Thursday. Zuckerberg encouraged the hacker ethos: move fast and break things. The ability to test many buggy versions was a competitive advantage against Google’s social product. Students with a few weeks of industry experience shipped the edit-post feature. I built an internal screenshot tool for reporting bugs. In 2012, Greylock Hackfest at Dropbox’s office was a ticket to meet prospective investors. The Ivy League teams were strong competition, and the demos had a wow factor. Today, vibe coding makes demos cheap and easy, so a demo alone is not enough. The format needs to evolve. Hackathons helped you learn, show your talent, and meet employers, investors, or future co-founders. I have fond memories of that era. Then startups took over my calendar. Who comes to hackathons in 2026? Most participants were students around 20, still figuring out what they wanted to do. AI was complicating those choices. One participant told me they had turned down a place at a prestigious, expensive UK university. Was it worth the investment in the AI era? Compared with my generation in Poland, the young people I meet seem more entrepreneurial, resourceful, and internationally minded. They trust each other more. My generation had it easier in one respect: we had conviction. I believed the internet would keep growing. After the iPhone launched, smartphones were an obvious bet. So was cloud computing. AI keeps surprising everybody. I get the doubts, but you need to believe and bet on something. We have a responsibility to help students find that direction. The event: Masters of Token Economics Our investor, Inovo https://inovo.vc/ , and Startup Founders Stars https://startupstars.pl/ organize the Vibe Coding Summer Jam series, where companies bring real business problems for participants to solve. Quesma hosted Session 5. We ran from 5 pm until past midnight. Friday evening worked well for everyone’s schedules. I opened with a presentation about Quesma and the challenge. Then participants formed five teams, each with its own room. They worked on their own laptops, with up to $100 in OpenRouter credits from Quesma. I helped teams explore different directions. Inovo partner Tomasz Swieboda was also available throughout the evening and brought two teenagers as observers. We provided pizza and non-alcoholic drinks. Our business case: find insights in SWE-chat At Quesma, we analyze coding-agent trajectories to improve token efficiency. Many companies are hitting their token budgets in 2026. For the hackathon, we presented SWE-chat https://swe-chat.com , originally gathered at Stanford from open-source contributors with research papers behind it. These are records of coding sessions involving tools such as Claude Code, Codex, Cursor, and Pi. We asked participants to download the data and fish for opportunities. An insight, a visualization, a token-saving technique: we wanted original ideas. What the five teams found “Hackengersi” asked whether chaotic prompts, written as though someone were drunk, affected results. Their findings were inconclusive. My takeaway: AI is good at interpreting messy language. Ideas and requirements may matter more than proper sentences. “Palantir” investigated swearing, threats, and other forceful corrections. They found that expressive language helps and that being rude when prompting agents is effective. It reminded me of Sergey Brin’s claim that models “tend to do better if you threaten them.” https://www.youtube.com/watch?v=8g7a0IWKDRE&t=500s “Zespół 1” had my favorite visual presentation. They examined agents rereading files after compaction and repeatedly checking whether a background job had finished. They suggested there might be an alternative to compaction that avoids some rereads. Another improvement is to write a script once to poll a background job without keeping the LLM in the loop. “ZENTRA” tackled agents stuck in loops: rerunning tests without changing the code. Their Progress Guard demo showed how a monitor could detect this, nudge the agent, or stop the loop. It reminded me of Jarred Sumner at Anthropic, who mostly prompted Claude with encouragement such as “keep going” and “believe in yourself” during his Riemann-hypothesis run. https://www.anthropic.com/research/riemann-zeta “Ekipa” won the vote by a small margin. Their dashboard explored context size, compaction, and how long agents worked between user messages. They questioned whether waiting until roughly 170,000 tokens to compact was too late and suggested shorter stretches of independent work. They found historical patterns that backed up developers’ intuitions, though the limits have changed in the latest models. What we learned and what comes next The tooling was uneven. Some had Claude Code or Codex ready; one struggled with copy-pasting from Gemini Flash. A $100 to $200 monthly subscription is a barrier for some students. Next time, I’d offer a short, optional workshop during the event: setting up tools and writing a program to analyze the data with OpenRouter and a cheaper model such as DeepSeek. Many teams analyzed only a subset. One reported spending about $40 on its analysis. I enjoyed the evening, and participants told me they did too. Companies that cut junior hiring risk missing out on their energy and ideas. I believe these students can help us adopt AI and cut token costs. Quesma is also sponsoring Warsaw Model Trainers. Our founding engineer Piotr Migdał is co-leading a model-training workshop https://luma.com/Warsaw-Model-Trainers-w3 with Ania Olchowik, ahead of the September 25-27 hackathon https://luma.com/Warsaw-Model-Trainers-hackathon . We’re exploring a token economics hackathon in San Francisco in the second half of October 2026. Thank you to Kewin Czupryński, Bartosz Podgórski, Tomasz Świeboda, and Malgorzata Piotrowska for organizing the event.