{"slug": "what-we-learned-running-an-agentic-hackathon-with-30k-hackers", "title": "What we learned running an agentic hackathon with 30k+ hackers", "summary": "HackerRank's fourth Orchestrate hackathon, which concluded September 12, drew 35,700 registrations and 3,062 participants, up from 12,871 registrations and 1,349 participants in the first May edition. HackerRank reported that early-career developers — students and those with up to two years of experience — grew from 52% of the Top 50 in May to 78% by September, and that the strongest submissions depended less on tool choice than on how developers directed, tested, and retained ownership of AI-generated work. The four evaluation metrics were positively but weakly to moderately correlated, with a maximum Spearman correlation of 0.512, indicating no single metric captured the full evaluation.", "body_md": "We wrapped up the fourth edition of Orchestrate on September 12th. What began as an experiment in helping our developer community become AI-first has grown into a recurring view of how developers actually build with AI and keep coming back to.\n\nAcross the four editions, participants used different models, coding harnesses, architectures, and development styles. Registrations grew from 12,871 to 35,700, while participation increased from 1,349 to 3,062.\n\nParticipants shared development transcripts, working implementations, task outputs, AI interviews, model configurations, and architectural decisions. These artifacts helped us to examine what strong AI-assisted engineering looks like and how it changes with experience, repetition, and tool choice.\n\nStrong submissions showed a complete engineering loop: planning, implementation, testing, debugging, correction, and review. They also constrained model output, built fallbacks, and defended their decisions with evidence.\n\nThe central finding was simple: strong AI-assisted engineering depended less on which tool developers chose and more on how well they directed it, tested its work, and retained ownership of the important decisions.\n\n## **What is Orchestrate?**\n\n**What is Orchestrate?**\n\n[Orchestrate](https://www.hackerrank.com/hackerrank-orchestrate-october26?utm_source=x&utm_medium=blog&utm_campaign=orchestrate-october26&utm_content=blog-post) is a 24-hour, AI-first hackathon on HackerRank where developers build real-world AI agents and are evaluated not only on their code, but also on how they think and work with AI.\n\nParticipants can use any IDE, coding tool, or model, including OpenAI, Claude, and Gemini etc. At the end of the challenge, they submit their code, AI chat transcript, and outputs for a provided test dataset. Our AI Interviewer (Chakra) reviews the code and development transcript, followed by a 30-minute live interview focused on the participant’s decisions, understanding of the problem, and use of AI.\n\n### **Growing from one edition to the next**\n\n**Growing from one edition to the next**\n\nThe interest in Orchestrate grew rapidly across the four editions. Registrations increased from 12,871 in May to 35,700 in September.\n\nParticipation grew from 1,349 participants in May to 3,062 in September. By the fourth edition, Orchestrate had attracted 2.8 times as many registrations and more than twice as many participants as the first.\n\nThe community remained strongly early-career throughout this growth. Students represented around six in ten participants in every edition, while developers with up to two years of professional experience formed the second-largest group.\n\n### **Early-career developers increasingly reached the Top 50**\n\nEarly-career developers weren't just the largest group entering; they made up a growing share of the strongest performers too. Students and developers with up to two years of experience were 52% of the Top 50 in May, climbing to 70% in June and August, and 78% by September.\n\n## **The four scores measured different strengths**\n\n**The four scores measured different strengths**\n\nTo see how the four evaluation metrics related to each other, we ran Spearman correlations across all four editions.\n\nSpearman measures relative ranking rather than linear relationship, which suited this better than Pearson; we wanted to know whether doing well on one metric tended to track doing well on another, not whether the raw scores moved together.\n\nThe four metrics were positively related, but the correlations were weak to moderate. The highest correlation across all four editions was 0.512, showing that no single metric represented the complete evaluation.\n\n*Chat Transcript* and *Code ZIP*  tracked each other most consistently, suggesting clearer development records went with stronger implementations. *Code ZIP* and *Output CSV* were closer in June and August, but that relationship weakened by September.\n\nThe *AI Interview score* stood apart from the other three. The interview seems to capture something the code or output alone can't show: whether a participant actually understood their system, could defend their decisions, and could name their own evidence and limitations.\n\n## **The harness remained fluid across the leaderboard**\n\n**The harness remained fluid across the leaderboard**\n\nTop submissions did not emerge from a single harness. Across the four editions, participants used a wide range of coding harnesses, and their preferences changed from one edition to the next.\n\nA coding harness is the environment through which a developer works with an AI model, for example, Claude Code, Codex, Cursor, or GitHub Copilot.\n\nAmong transcripts where a harness could be identified, Antigravity was the most commonly used tool in three of the four editions, while Claude Code briefly moved ahead in August. Codex maintained a comparatively stable presence throughout.\n\nNo harness won fully; the usage stayed fragmented beyond the top few, with participants mixing interface, model access, autonomy, and workflow in different combinations. The tooling landscape moved faster than it converged.\n\n### **The most popular harness was not the most common among the Top 50**\n\nThe harness distribution, though, looked different among the highest-ranked submissions. Although Antigravity was the most commonly used harness overall in three editions, Claude Code was the most common among the Top 50 in every edition.\n\nHarness choice may also reflect participant experience, access to free and paid plans, model availability, and challenge requirements.\n\n### **Experience was associated with harness choice**\n\n**Experience was associated with harness choice**\n\nWe tried looking at how the harness choice changed with the experience year. Students consistently preferred Antigravity, while participants with more than four years of experience consistently preferred Claude Code.\n\nAcross the four editions, Antigravity’s share fell from 39% among students to 11% among participants with 7+ years of experience. Claude Code moved in the opposite direction, rising from 14% to 41%.\n\nCodex did not follow either trend as strongly. Its usage remained comparatively stable across experience groups, suggesting that it appealed to a broader cross-section of participants.\n\nThe correlation may reflect differences in free and paid access, familiarity with terminal-based development, or the level of direct control participants wanted over the build.\n\n## **Stronger submissions followed a complete engineering loop**\n\nTool choice varied widely, but how developers on the top of the leaderboard worked was far more consistent. Regardless of harness, model, or experience level, stronger submissions were more likely to show a complete loop: plan, implement, run, debug, correct, review.\n\n**Plan → implement → run → debug → correct → review**\n\nRather than labeling conversations subjectively, we checked development transcripts for evidence of these six stages.\n\nAverage scores rose with the number of visible stages and not just on the Chat Transcript evaluation. AI Interview, Output CSV, and Code ZIP scores all rose too.\n\nEven excluding the *Chat Transcript score* entirely, submissions with five or six visible stages averaged a score of 42.12 across the remaining evaluations, against 27.38 for submissions with two or fewer.\n\nThis pattern was held in every edition. Spearman correlation between visible stages and total score ranged from 0.446 to 0.559, and stayed moderately positive (0.397–0.492) even with Chat Transcript excluded.\n\n### **This pattern was common near the top**\n\n**This pattern was common near the top**\n\nA complete engineering loop appeared in 45 of the Top 50 transcripts in May, 44 of the 49 available transcripts in June, 48 of 50 in August, and 49 of 50 in September.\n\nBy September, the loop had also become common in the 51–200 range, showing that its presence alone no longer separated the very top of the leaderboard, but its absence stayed rare among the top submissions.\n\n## **What are the models used in the agent across the editions?**\n\nParticipants used a broad and shifting mix of model families, and no single provider became the default among high-performing submissions.\n\nModel preferences changed across the four editions. Gemini became the most commonly identified family from June onward, while Llama’s presence declined. GPT and Claude remained comparatively stable, and Qwen became more visible over time.\n\nSome submissions used more than one model family, so these percentages aren't mutually exclusive.\n\n### **The Top 50 used a broad mix of models**\n\n**The Top 50 used a broad mix of models**\n\nThe highest-ranked submissions didn't converge on one model family either; Gemini, GPT, Claude, Llama, and Qwen all showed up in the Top 50 across all four editions.\n\nClaude was more heavily represented in the Top 50 than in the wider field in every single edition. GPT showed up slightly more often near the top too, while Gemini stayed widely used overall but less concentrated at the top.\n\n### **Developers matched model families to different roles**\n\nRather than using one model end-to-end, many participants assigned different models to different stages. No model family had an exclusive job, but clear tendencies emerged:\n\n1. **Gemini** : multimodal interpretation, image/document understanding, structured extraction, classification, routing\n2. **GPT** : reasoning, classification, evidence verification, routing, explanation generation; often the verification layer after other components gathered context\n3. **Claude** : grounded response generation, architectural reasoning, multimodal extraction, review; also used for routing decisions and output verification\n4. **Llama** : support-ticket classification, routing, and response generation in RAG pipelines; a local/cost-conscious option\n5. **Qwen** :  image understanding, structured extraction, classification, routing; another lower-cost local alternative\n\nSpecialist models handled narrower jobs too; in August, 641 submissions paired a Whisper-family model for speech transcription with a separate language model for reasoning: Whisper converted audio to text, the other model interpreted it.\n\nEmbedding and reranking models played a similar role, retrieving relevant documents or past messages before handing context to a language model.\n\nMost systems were built around distinct responsibilities, with each model handling the part of the problem it suited best.\n\n## **The strongest interviews made engineering judgment visible**\n\nThe strongest participants didn't rely on polished delivery or a broad description of what their agent was supposed to do. They gave answers that could be checked by pointing to their architecture and code, explaining specific decisions, citing evidence from testing, and naming where their systems could fail.\n\n### **What stronger candidates did differently**\n\n**What stronger candidates did differently**\n\nThe biggest gaps were in evaluation evidence, quantified results, and the ability to present what we defined as a complete technical defence: architecture, the reasoning behind key decisions, evidence from testing, and known failure modes.\n\nStronger candidates consistently connected high-level claims to implementation details, test results, and known constraints.\n\n### **They gave more complete answers with less prompting**\n\nStronger candidates were more likely to answer the question directly, add implementation detail, explain their reasoning, and back it with evidence without needing repeated follow-ups.\n\nThe common structure: direct answer, then implementation detail, then the reasoning or trade-off behind it, then supporting evidence, then a limitation or fallback.\n\nFor example, a strong candidate wouldn't just say the system used a two-stage pipeline; they explained why the stages were split, where the handoff happened, what failure the split prevented, and what test or example backed the decision.\n\nWeaker interviews tended to stay at the level of a project description; what the agent was meant to do, which model or framework it used, the rough sequence of the system, and even with follow-ups, rarely got down to specific implementation details, measured results, or fallback behavior.\n\n### **Ownership was visible in the language**\n\n**Ownership was visible in the language**\n\nStronger candidates were more likely to say clearly what they chose, tested, changed, rejected, or fixed, separating their own decisions from work they had delegated to an AI coding tool.\n\nNaming a limitation didn't weaken an answer. Top candidates named what their system couldn't handle, what it fell back to instead, and how they'd fix it; these interviews had a clear \"gap, current fallback, next step\" structure.\n\nThe technical questions changed with each challenge, but this answer pattern held steady: a complete technical defence appeared in 96% of May's Top 50 interviews, 90% in June, 98% in August, 80% in September — against 32%, 18%, 10%, and 10% in the Bottom 50.\n\n## **How does the AI Interview score & pattern vary across the experience years?**\n\nParticipants with two to six years of experience generally scored best on the *AI interview*.\n\nBut the relationship between experience and interview performance was weak (correlations of 0.045 to 0.125). Within every experience group, candidates who explained their architecture, justified decisions, presented evaluation evidence, and named failure modes scored 2.6 to 3.6 points higher than those who didn't.\n\nThe AI Interview rewarded visible engineering judgment more than it rewarded years of experience.\n\n## **Experience helped with implementation**\n\n**Experience helped with implementation**\n\nProfessional experience tracked more closely with quality of implementation than with interview performance. Participants with four to six years led on *Code ZIP* score in May, August, and September; those with seven-plus years narrowly led in June.\n\nThe gap showed up in distribution too: only around one in five students reached the top quarter of *Code ZIP scores*, against more than half of participants with four to six years of experience.\n\nExperience seemed to help with architecture, modularity, validation, and failure handling.\n\n## **The architecture changed with the problem**\n\n**The architecture changed with the problem**\n\nThe architecture participants chose changed with the problem they were trying to solve.\n\nWe classified each submitted implementation into one of six mutually exclusive architecture patterns:\n\n- **Single Agent:** One model handled the main reasoning and produced the final decision without retrieval, external tools, or specialist agents.\n- **Multi-Agent:** Multiple named agents or model-led stages handled different parts of the problem before their results were combined.\n- **Deterministic Pipeline:** Application code, calculations, thresholds, and fixed rules controlled the decision without an explicit LLM call.\n- **Model-Assisted Pipeline:** Application code controlled the workflow while models performed bounded tasks such as interpreting an image, classifying a message, or generating an explanation.\n- **ML Model:** A trained statistical or classical machine-learning model served as the primary decision engine.\n- **Single Agent with Tools or RAG:** One primary agent used retrieval or external tools to gather context, inspect evidence, perform safety checks, or validate its output.\n\nThe architecture closely followed the challenge. The documentation-grounded problem favoured agents with retrieval or tools. The visual-evidence problem produced a sharp increase in model-assisted pipelines, while the challenge combining media, message history, personalisation, and policy rules resulted in a more varied architecture mix.\n\nThe financial-planning problem shifted the balance toward deterministic and model-assisted pipelines. Application code typically handled calculations and constraints, while models performed narrower tasks such as reading receipts, interpreting messages, or generating explanations.\n\n## **The strongest codebases were built around model uncertainty**\n\nThe highest-ranked Code ZIP submissions did not have more agents or a more complicated framework. They stood out because the system was engineered to work around the model.\n\nStronger implementations were more likely to constrain model output, validate the result, recover from failures, and provide a repeatable evaluation path.\n\n### **Higher-ranked submissions controlled the output**\n\n**Higher-ranked submissions controlled the output**\n\nThe clearest split was output control. Lower-ranked submissions tended to stop once the model produced an answer. Higher-ranked ones checked whether that answer was structurally valid, safe, complete, and usable downstream.\n\nEvaluation showed the same divide (comparing June, August, and September, since May's submission format didn't expose this field consistently): an explicit evaluation pipeline appeared in 98.9% of the top quarter, against 68.6% of the bottom quarter.\n\nTop submissions included a repeatable way to run examples, compare outputs, check field-level accuracy, and investigate failures before submitting.\n\n### **Architecture also changed across the leaderboard**\n\n**Architecture also changed across the leaderboard**\n\nPurely deterministic pipelines clustered in the bottom quarter. Model-assisted pipelines and single agents with tools/RAG showed up more in the top quarter.\n\nAcross the 200 Top 50 submissions from all four editions: 46.5% used a Single Agent with Tools or RAG, 35.0% a Model-Assisted Pipeline, 14.0% a Multi-Agent architecture, and 4.5% a direct Single Agent. All 200 used structured model output, guardrails, or validation and included deterministic rules where necessary.\n\n## **What high-performing submissions had in common**\n\n**What high-performing submissions had in common**\n\nThe strongest submissions performed well across the development transcript, *AI interview*, *final output*, and *submitted code*.\n\nOf the 200 Top 50 submissions across the four editions, 149 ranked in the top quarter on all four evaluation metrics. Another 47 ranked in the top quarter on three metrics, while only four did so on two.\n\nAcross the four editions, five patterns repeatedly appeared among high-performing submissions:\n\n- They completed the engineering loop instead of stopping at the first generated solution.\n- They used models for interpretation while relying on code for constraints, validation, and control.\n- They treated the required output as a contract and built fallbacks around model failure.\n- They defended their decisions with implementation details, evidence, and known limitations.\n- They preserved human ownership by making clear what they chose, tested, changed, and delegated to AI.\n\n## **Repeat participation helped in improving**\n\n**Repeat participation helped in improving**\n\nAcross the four editions, 915 developers returned to participate again. Comparing each participant’s first edition with their latest revealed a consistent pattern: repeat participation improved how developers built, documented, and defended their systems more reliably.\n\nThe main patterns were:\n\n- Development records became more complete**.** Structured session logs increased from 62% on a participant’s first attempt to 77% on their latest.\n- Transcripts showing planning, implementation, testing, debugging, correction, and verification increased from 56% to 67%.\n- *AI Interview* performance improved for 58% of returning developers.\n- *Code ZIP* performance increased for 57% of returning developers.\n- *Output CSV* performance was evenly split, with roughly half improving and half declining.\n\nOverall, 56% of returning participants moved higher relative to the field. Returning therefore helped, but did not guarantee a better leaderboard position.\n\n## **What four editions taught us**\n\n**What four editions taught us**\n\nNo single model, harness, or architecture determined success. The strongest participants stood out through their process: they tested and corrected their work, controlled and validated model output, and defended their decisions with evidence.\n\nReady to put these patterns into practice? [Register for the next edition of Orchestrate](https://www.hackerrank.com/hackerrank-orchestrate-october26?utm_source=x&utm_medium=blog&utm_campaign=orchestrate-october26&utm_content=blog-post) and show us how you build with AI.", "url": "https://wpnews.pro/news/what-we-learned-running-an-agentic-hackathon-with-30k-hackers", "canonical_source": "https://twitter.com/rvivek/status/2106264394414108925", "published_at": "2026-10-03 15:52:59+00:00", "updated_at": "2026-10-03 16:06:45.646349+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "artificial-intelligence", "developer-tools"], "entities": ["HackerRank", "Orchestrate", "Chakra", "OpenAI", "Claude", "Gemini"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/what-we-learned-running-an-agentic-hackathon-with-30k-hackers", "markdown": "https://wpnews.pro/news/what-we-learned-running-an-agentic-hackathon-with-30k-hackers.md", "text": "https://wpnews.pro/news/what-we-learned-running-an-agentic-hackathon-with-30k-hackers.txt", "jsonld": "https://wpnews.pro/news/what-we-learned-running-an-agentic-hackathon-with-30k-hackers.jsonld"}}