{"slug": "evaluating-and-optimising-ai-in-the-real-world", "title": "Evaluating and Optimising AI in the Real World", "summary": "Tessl will host the AI Native Dev community meetup on 30 April from 17:00 to 20:00 GMT+1 at Tessl AI Limited in London, featuring three technical talks on real-world AI performance. Tessl member of technical staff Rob Willoughby will present lessons from running large-scale evaluations of AI coding agents at 6:30pm, NVIDIA senior AI DevTech engineer Dominic Brown will present Skip Softmax Attention for accelerating long-context LLM inference at 7pm, and Hugging Face's Shaun Smith will discuss using Skills and MCP for open source ML at 7:30pm.", "body_md": "Hosted by\n\n[Contact organizer](mailto:sam@tessl.io)\n\nWorkshop\n\n# Evaluating and Optimising AI in the Real World\n\nJoin the AI Native Dev community in London for a night of practical, technical insight into how modern AI systems actually perform in the real world. From NVIDIA’s work on scaling long-context LLM inference, to Hugging Face’s perspective on Skills and MCP, to Tessl’s lessons from running large-scale evals on AI coding agents, this meetup is all about what breaks, what works, and what developers need to know. Come for the talks, stay for the conversations, and meet others building the next generation of AI-native software.\n\n30 April\n\n17:00 - 20:00 GMT+1\n\nTessl AI Limited, London, UK\n\nAbout the event\n\nJoin the **AI Native Dev** community in London for a night of practical, technical insight into how modern AI systems actually perform in the real world. From [NVIDIA’s](https://www.nvidia.com/en-gb/?utm_source=tessl) work on scaling long-context LLM inference, to [Hugging Face’s](https://huggingface.co/?utm_source=tessl) perspective on Skills and MCP, to [Tessl’s](https://tessl.io/?utm_source=tessl)  lessons from running large-scale evals on AI coding agents, this meetup is all about what breaks, what works, and what developers need to know. Come for the talks, stay for the conversations, and meet others building the next generation of AI-native software.\n\n**Agenda**\n\n- **6pm:** Doors Open\n- **Talks kick off**\n  - **6:30:** Talk 1:**Evaluating AI Skills in the Wild: What We Learned Running Evals at Scale** by Rob Willoughby\n  - **7pm:** Talk 2:**Accelerating Long-Context Inference with Skip Softmax Attention** by Dom Brown\n  - **7:30pm:** Talk 3:**Using Skills and MCP for Open Source ML** by Shaun Smith\n- \n- **8pm:** Networking\n- **9pm:** THE END\n\n**Evaluating AI Skills in the Wild: What We Learned Running Evals at Scale**\n\nHow do you know if an AI coding agent is actually following your instructions? Or if that skill you wrote is having an impact? What can you learn from synthetic eval cases compared to running them in your project on your code? We'll share practical lessons from building and running large-scale evaluations of coding agents, covering eval design, the ways things break at scale, and what the results reveal about where today's models actually struggle.\n\n**Rob Willoughby, Member of Technical Staff, AI Research at Tessl**\n\nRob works on evaluation research at Tessl, designing and running large-scale assessments of how coding agents behave in real-world codebases. His work focuses on figuring out what \"good\" looks like when an AI agent works with your code from eval design and rubric systems to understanding where models (and infrastructure) break down at scale.\n\n**Accelerating Long-Context Inference with Skip Softmax Attention**\n\nThe growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks inherent to the standard attention mechanism. To address this challenge, we introduce Skip Softmax Attention, a drop-in sparse attention method that dynamically prunes the attention matrix without any pre-computation or proxy scores. Our method uses a fixed threshold and existing information from online softmax to identify negligible attention scores, skipping softmax computation, value block loading, and the subsequent matrix multiplication. This fits seamlessly into existing FlashAttention kernel designs with negligible latency overhead.\n\n**Dominic Brown, Senior AI DevTech Engineer at NVIDIA**\n\nDom Brown is a senior AI developer technology engineer at NVIDIA. His work focuses on optimizing large language model inference. Dom studied computer science at the University of Warwick, UK. In his PhD thesis, he investigated the acceleration of particle-in-cell algorithms on CPU and GPU systems.\n\n**Upskilling Models, Agents and the ML Pipeline.**\n\nRecent advances in Language Models and Coding Agents have impacted the ML Stack itself, opening up new opportunities for automation and training.\n\nAt Hugging Face we've incorporated Skills and MCP in a number of areas - from automating model training, and having large models tutor smaller models for certain tasks. We'll share examples of MCP for model-to-model communication, practical tools for upskilling agents and improving our ML pipeline and share some of the tools lessons we've learned along the way.\n\n**Shaun Smith, Open source MCP Lead at HF**\n\nShaun Smith leads Open Source MCP at Hugging Face, and is an MCP Steering Committee member serving as a Community Moderator and within the Transports Working Group.\n\nThis event is brought to you as part of the [AI Native Dev Community](https://ainativedev.io/?utm_source=tessl) . Consider subscribing to the Mailing List, [Podcast](https://www.tessl.io/podcast?utm_source=tessl) , and joining our [Discord Community](https://tessl.co/4ghikjh?utm_source=tessl)\n\nSpeakers", "url": "https://wpnews.pro/news/evaluating-and-optimising-ai-in-the-real-world", "canonical_source": "https://tessl.io/events/evaluating-and-optimising-ai-in-the-real-world", "published_at": "2026-09-23 20:29:03+00:00", "updated_at": "2026-09-23 20:30:30.078921+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "agent-protocols", "ai-research"], "entities": ["Tessl", "NVIDIA", "Hugging Face", "Rob Willoughby", "Dominic Brown", "Shaun Smith", "AI Native Dev", "Skip Softmax Attention"], "alternates": {"html": "https://wpnews.pro/news/evaluating-and-optimising-ai-in-the-real-world", "markdown": "https://wpnews.pro/news/evaluating-and-optimising-ai-in-the-real-world.md", "text": "https://wpnews.pro/news/evaluating-and-optimising-ai-in-the-real-world.txt", "jsonld": "https://wpnews.pro/news/evaluating-and-optimising-ai-in-the-real-world.jsonld"}}