Poolside Releases Laguna S 2.1, an Open-Weight Agentic Coding Model Punching Above Its Weight Class on SWE-Bench Multilingual Poolside released Laguna S 2.1, a 118B-parameter open-weight Mixture-of-Experts model with 8B activated parameters per token, achieving 78.5% on SWE-Bench Multilingual and 70.2% on Terminal-Bench 2.1, topping the published leaderboard for open models. The model, trained on 4,096 NVIDIA H200 GPUs starting May 22, 2026, supports up to 1M token context and runs on a single NVIDIA DGX Spark. Poolside https://poolside.ai/ has released Laguna S 2.1 https://poolside.ai/blog/introducing-laguna-s-2-1 , a 118B-parameter open-weight model built for agentic coding. It is a Mixture-of-Experts MoE model with 8B activated parameters per token. It supports a context window of up to 1M tokens in both thinking and no-thinking modes. The weights are on Hugging Face https://huggingface.co/poolside/Laguna-S-2.1 under an OpenMDW-1.1 license, and the model is small enough to run on a single NVIDIA DGX Spark https://www.nvidia.com/en-us/products/workstations/dgx-spark/ . On long-horizon coding benchmarks, Laguna S 2.1 holds its own against models several times its size, including DeepSeek-V4-Pro-Max, NVIDIA’s Nemotron 3 Ultra, and Thinking Machines’ Inkling. Laguna S 2.1 is a scale-up of the Laguna XS family, trained on the same pre-training data as XS 2.1. What is Laguna S 2.1 The model activates roughly 6.8% of its parameters on any given token. All 118B parameters remain resident in memory, but only ~8B route through the network per step. That sparsity is why a mid-size model can behave like a larger one while staying cheap to serve. Poolside team publishes weights in BF16, FP8, INT4, and NVFP4, along with official GGUF and MLX conversions and DFlash draft models. It went from the start of training to launch in under nine weeks. Pre-training began on 22 May 2026 on 4,096 NVIDIA H200 GPUs. It is the first Poolside model where reinforcement learning ran in FP8 precision. Performance Laguna S 2.1 scores 70.2% on Terminal-Bench 2.1 https://poolside.ai/blog/introducing-laguna-s-2-1 with thinking enabled. That places it first among open, disclosed-size models on Poolside’s compiled leaderboard, behind only larger or closed systems. On SWE-Bench Multilingual it scores 78.5%, topping the published table outright. The full comparison Poolside released is below. | Benchmark | Laguna S 2.1 118B-A8B | Tencent Hy3 295B-A21B | Inkling 975B-A41B | Nemotron 3 Ultra 550B-A55B | DeepSeek-V4-Pro-Max 1.6T-A49B | Kimi K3 2.8T-A50B | Qwen 3.7 Max | Muse Spark 1.1 | Claude Fable 5 | |---|---|---|---|---|---|---|---|---|---| | Terminal-Bench 2.1 | 70.2 | 71.7 | 63.8 | 56.4 | 64.0 | 88.3 | 74.5 | 80.0 | 88.0 | | SWE-Bench Multilingual | 78.5 | 75.8 | – | 67.7 | 76.2 | – | 78.3 | – | – | | SWE-Bench Pro Public | 59.4 | 57.9 | 54.3 | – | 55.4 | – | 60.6 | 61.5 | 80.3 | | DeepSWE v1.1 | 40.4 | – | – | – | 9.0 | 69.0 | – | 53.3 | 70.0 | | SWE Atlas Codebase QnA | 46.2 | – | – | – | 27.2 | – | – | 42.2 | – | | Toolathlon Verified | 49.7 | – | 45.5 | 34.3 | 55.9 | – | – | 75.6 | – | The clearest signal is DeepSWE v1.1, which still has real headroom. There, Laguna S 2.1 scores 40.4% against DeepSeek-V4-Pro-Max’s 9.0%, with roughly one-sixth the active parameters. Closed frontier models such as Claude Fable 5 and Kimi K3 still lead on several benchmarks. Poolside’s claim is about the weight class, not the outright top. Trajectories from the final evaluation set is published at trajectories.poolside.ai https://trajectories.poolside.ai/ . Two thinking modes, and where the score comes from Laguna S 2.1 has two modes: off and max , with max enabled by default. In max mode the model sets its own test-time compute budget. Poolside is shipping without user-configurable low/medium/high effort control for now. Max thinking lifts Terminal-Bench 2.1 from 60.4% to 70.2%. It lifts DeepSWE from 16.5% to 40.4%. Those gains cost tokens: DeepSWE trajectories run about 249k completion tokens in thinking mode against 99k without. Poolside team reports coherent, productive reasoning running for hours and hundreds of thousands of tokens. Three published runs Poolside team shared three unedited trajectories to show behavior rather than scores. In one, the model built a working HTML/CSS browser engine from an empty folder. That run took 181 steps across a 50-minute session, with no human intervention, and it validated its own output against headless Chromium. In a second, the model optimized Poolside’s own agent harness. It made the harness 5.2% faster with roughly 71% lower memory allocation, replacing O n² string concatenation with buffers. In a third, it re-derived Erdős Problem 397 https://www.erdosproblems.com/397 offline in Perl over 68 minutes, since the sandbox had no Python. That result is an independent rediscovery; GPT-5.2 Pro solved the same problem in January 2026, and Laguna’s knowledge cutoff is November 2025. Can you actually deploy it? Sizing uses the full 118B parameters, not the 8B active count, because every expert stays in memory. At 4-bit NVFP4 or INT4 the weights need about 59 GB, which fits comfortably on a single DGX Spark’s 128 GB of unified memory. At FP8 they need about 118 GB, still within a single Spark or a single H200 https://www.nvidia.com/en-us/data-center/h200/ . At BF16 they need about 236 GB, which calls for two linked Sparks or a multi-GPU datacenter node. Poolside worked with NVIDIA to optimize inference from TRT-LLM serving to NVFP4 on Blackwell, down to a single DGX Spark. It shipped day-one support for vLLM https://github.com/vllm-project/vllm , SGLang https://github.com/sgl-project/sglang , and Ollama https://ollama.com/ . Hosted access runs through OpenRouter https://openrouter.ai/poolside , free at 256K context and paid at the full 1M context for $0.10 / $0.20 / $0.01 per 1M input / output / cache-read tokens. The model is also on Baseten https://www.baseten.co/ , Kilo https://kilocode.ai/ , Prime Intellect https://www.primeintellect.ai/ Prime Lab, and ZML https://zml.ai/ . Key Takeaways - Laguna S 2.1 is a 118B-total / 8B-active MoE coding model with a 1M-token context, open under OpenMDW-1.1. - It scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual, leading open disclosed-size models. - At 4-bit it runs on a single NVIDIA DGX Spark; FP8 fits one Spark or H200, BF16 needs two. - Default ‘max thinking’ drives most of the score, at a real token cost DeepSWE: 16.5% → 40.4% . - Trained in under nine weeks on 4,096 H200 GPUs, it is Poolside’s third model in under three months. Sources: Poolside Technical— Introducing Laguna S 2.1 https://poolside.ai/blog/introducing-laguna-s-2-1 , Robert McHardy on X https://x.com/robert mchardy/status/2079614440295506171 , Poolside on X https://x.com/poolsideai/status/2079613777343848465 and Hugging Face model card https://huggingface.co/poolside/Laguna-S-2.1 . Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.