cd /news/artificial-intelligence/evaluating-and-optimising-ai-in-the-… · home topics artificial-intelligence article
[ARTICLE · art-138521] src=tessl.io ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Evaluating and Optimising AI in the Real World

Tessl will host the AI Native Dev community meetup on 30 April from 17:00 to 20:00 GMT+1 at Tessl AI Limited in London, featuring three technical talks on real-world AI performance. Tessl member of technical staff Rob Willoughby will present lessons from running large-scale evaluations of AI coding agents at 6:30pm, NVIDIA senior AI DevTech engineer Dominic Brown will present Skip Softmax Attention for accelerating long-context LLM inference at 7pm, and Hugging Face's Shaun Smith will discuss using Skills and MCP for open source ML at 7:30pm.

by read4 min views3 publishedSep 23, 2026
Evaluating and Optimising AI in the Real World
Image: source

Hosted by

Contact organizer Workshop

Join the AI Native Dev community in London for a night of practical, technical insight into how modern AI systems actually perform in the real world. From NVIDIA’s work on scaling long-context LLM inference, to Hugging Face’s perspective on Skills and MCP, to Tessl’s lessons from running large-scale evals on AI coding agents, this meetup is all about what breaks, what works, and what developers need to know. Come for the talks, stay for the conversations, and meet others building the next generation of AI-native software.

30 April

17:00 - 20:00 GMT+1 Tessl AI Limited, London, UK

About the event

Join the AI Native Dev community in London for a night of practical, technical insight into how modern AI systems actually perform in the real world. From NVIDIA’s work on scaling long-context LLM inference, to Hugging Face’s perspective on Skills and MCP, to Tessl’s lessons from running large-scale evals on AI coding agents, this meetup is all about what breaks, what works, and what developers need to know. Come for the talks, stay for the conversations, and meet others building the next generation of AI-native software.

Agenda

  • 6pm: Doors Open
  • Talks kick off
    • 6:30: Talk 1:Evaluating AI Skills in the Wild: What We Learned Running Evals at Scale by Rob Willoughby
    • 7pm: Talk 2:Accelerating Long-Context Inference with Skip Softmax Attention by Dom Brown
    • 7:30pm: Talk 3:Using Skills and MCP for Open Source ML by Shaun Smith
- **8pm:** Networking
- **9pm:** THE END

Evaluating AI Skills in the Wild: What We Learned Running Evals at Scale

How do you know if an AI coding agent is actually following your instructions? Or if that skill you wrote is having an impact? What can you learn from synthetic eval cases compared to running them in your project on your code? We'll share practical lessons from building and running large-scale evaluations of coding agents, covering eval design, the ways things break at scale, and what the results reveal about where today's models actually struggle.

Rob Willoughby, Member of Technical Staff, AI Research at Tessl

Rob works on evaluation research at Tessl, designing and running large-scale assessments of how coding agents behave in real-world codebases. His work focuses on figuring out what "good" looks like when an AI agent works with your code from eval design and rubric systems to understanding where models (and infrastructure) break down at scale.

Accelerating Long-Context Inference with Skip Softmax Attention

The growing demand for long-context inference capabilities in Large Language Models (LLMs) has intensified the computational and memory bottlenecks inherent to the standard attention mechanism. To address this challenge, we introduce Skip Softmax Attention, a drop-in sparse attention method that dynamically prunes the attention matrix without any pre-computation or proxy scores. Our method uses a fixed threshold and existing information from online softmax to identify negligible attention scores, skipping softmax computation, value block , and the subsequent matrix multiplication. This fits seamlessly into existing FlashAttention kernel designs with negligible latency overhead.

Dominic Brown, Senior AI DevTech Engineer at NVIDIA

Dom Brown is a senior AI developer technology engineer at NVIDIA. His work focuses on optimizing large language model inference. Dom studied computer science at the University of Warwick, UK. In his PhD thesis, he investigated the acceleration of particle-in-cell algorithms on CPU and GPU systems.

Upskilling Models, Agents and the ML Pipeline.

Recent advances in Language Models and Coding Agents have impacted the ML Stack itself, opening up new opportunities for automation and training.

At Hugging Face we've incorporated Skills and MCP in a number of areas - from automating model training, and having large models tutor smaller models for certain tasks. We'll share examples of MCP for model-to-model communication, practical tools for upskilling agents and improving our ML pipeline and share some of the tools lessons we've learned along the way.

Shaun Smith, Open source MCP Lead at HF

Shaun Smith leads Open Source MCP at Hugging Face, and is an MCP Steering Committee member serving as a Community Moderator and within the Transports Working Group.

This event is brought to you as part of the AI Native Dev Community . Consider subscribing to the Mailing List, Podcast , and joining our Discord Community

Speakers

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @tessl 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evaluating-and-optim…] indexed:0 read:4min 2026-09-23 ·