{"slug": "rubyllm-2-1-mcp-judgments-evaluations-and-less-work-on-every-call", "title": "RubyLLM 2.1: MCP, Judgments, Evaluations, and Less Work on Every Call", "summary": "RubyLLM 2.1 was released, adding an MCP client, typed judgments, evaluations, tool progress, OpenTelemetry tracing, and a nineteenth provider, while reducing per-call work with no code changes. The MCP client lets developers define an MCP server as a Ruby class that controls which tools the model sees, and the release notes warn users of the ruby_llm-mcp gem to remove it before updating because it defines the same RubyLLM::MCP constant. Judgments return model-measured probabilities, choices, and scores over a single request, defaulting to Jev models from TypeSafe, and evaluations run YAML-defined cases with model grading via bin/rails \"ruby_llm:eval[DocsEvaluation]\", exiting non-zero on failure for CI use.", "body_md": "[RubyLLM](https://rubyllm.com) 2.1 is out. I released it on stage at Deccan Queen on [Rails](https://rubyonrails.org) in Pune.\n\nIt builds on [2.0](https://paolino.me/rubyllm-2-0/) and adds an MCP client, typed judgments, evaluations, tool progress, OpenTelemetry tracing, and a nineteenth provider. It’s also faster, with no code changes. In this post:\n\n## An MCP Client Where You Decide What the Model Sees\n\nIn RubyLLM 2.1 an MCP server is a [Ruby](https://www.ruby-lang.org/en/) class you own: it lives in your repo, and it says exactly which tools the model gets.\n\n```\nclass Linear < RubyLLM::MCP\n  url \"https://mcp.linear.app/mcp\"\n  inputs :user\n  oauth owner: :user\n\n  only :list_issues, :get_issue, :create_issue # the model sees these three\n  requires_approval :create_issue\nend\n\nchat = RubyLLM.chat.with_mcp(Linear.new(user: current_user)) # one connection per user\nchat.ask \"What's blocking the release?\"\n```\n\nTo try a server before shaping it, connect from the console. Every tool becomes a Ruby method:\n\n```\n>> docs = RubyLLM.mcp(url: \"https://learn.microsoft.com/api/mcp\")\n>> docs.microsoft_docs_search(query: \"Azure Blob Storage\").text\n=> \"...\"\n```\n\nThe same class can rename tools, rewrite their descriptions, fix arguments, and wrap results, which is how 2.1 follows the advice from April that your agent’s [context window is not a junk drawer](https://paolino.me/your-agents-context-window-is-not-a-junk-drawer/): prototype with MCP, then craft the tools you control. Before this, [ruby_llm-mcp](https://github.com/patvice/ruby_llm-mcp) by @patvice carried MCP for RubyLLM users for a long time. Thank you. If you use it, remove it before updating, since it defines the same `RubyLLM::MCP` constant. The [MCP Client guide](https://rubyllm.com/mcp/) covers OAuth, input requests, resources, prompts, MCP Apps, and Tasks.\n\n## Judgments\n\nJudgments answer questions about your own data (is this urgent, which team owns it) with probabilities a model measured, instead of a confidence number it wrote as text:\n\n```\nclass TicketTriage < RubyLLM::Judge\n  probability :urgent, \"Does this need attention today?\"\n\n  choice :department, \"Which team should handle this?\" do\n    billing   \"Payments and refunds\"\n    technical \"Bugs and integrations\"\n    other     \"Everything else\"\n  end\n\n  score :frustration, \"How frustrated is the customer?\",\n    [\"Calm\", \"Frustrated\", \"Angry\"]\nend\n\njudgment = TicketTriage.judge(\"Please refund the duplicate charge today.\")\njudgment.urgent.probability # => 0.96\njudgment.department.choice  # => :billing\n```\n\n`probability`, `choice`, and `score` questions are all asked over the same input in one request, and your code acts on them with thresholds you pick. Judges default to Jev models from TypeSafe, a new built-in provider. Kieran Klaassen added OpenAI’s `gpt-6-luna` through the Decisions API ([#1008](https://github.com/crmne/ruby_llm/pull/1008)). See the [Judgments guide](https://rubyllm.com/judgments/).\n\n## Evaluations\n\nAn evaluation runs your agent on cases with known answers and has a model grade each response, so you can tell whether a prompt or model change made things better. Cases live in a YAML file next to the class, and `bin/rails \"ruby_llm:eval[DocsEvaluation]\"` runs them and exits non-zero on failure, so it works in CI:\n\n```\nclass DocsEvaluation < RubyLLM::Evaluation\n  evaluation :correctness # keeps the default check\n  evaluation :grounded, \"Every claim is supported by the documents in metadata\"\n  evaluation :cites_sources, \"The answer links to at least one document\"\n\n  def perform(question)\n    DocsAgent.new.ask(question)\n  end\nend\n```\n\nEvaluations can also grade tool calls with plain Ruby assertions, use an Agent or a Judge as the grader, and run as RSpec or Minitest tests. See the [Evaluations guide](https://rubyllm.com/evaluations/).\n\n## Faster by Default\n\n2.1 does less work on every call, with no code changes. RubyLLM’s own work, 2.0.0 against 2.1 on the same machine:\n\n| Workload | 2.0 | 2.1 | \n|---|---|---|\n| Stream a 2 MB event that arrives in 16 KB pieces | 116 ms | 2.0 ms | \n| Stream 500 Perplexity chunks that each cite 20 sources | 137 ms | 17 ms | \n| Ask a Bedrock chat with 200 messages of history | 2.9 ms | 0.40 ms | \n| Memory kept by a streamed 40-turn chat with a 256 KB image | 28 MB | 0.48 MB | \n| Eight threads loading the model registry at once | 273 ms | 32 ms | \n\nConnections are now shared across calls, threads, and fibers. To keep them open, pick a persistent adapter (from the `faraday-net_http_persistent` gem):\n\n```\nconfig.faraday_adapter = :net_http_persistent # keeps connections open\n```\n\nThe benchmarks need no API keys: clone RubyLLM and run `bundle exec rake \"benchmark:compare[v2.0.0]\"` to compare any version with your checkout. The [connection guide](https://rubyllm.com/configuration-connection/#connection-reuse) covers adapters.\n\n## Tools That Report Progress\n\nA slow tool can now tell your user what it’s doing instead of leaving them with a spinner:\n\n``` python\nclass ReadReport < RubyLLM::Tool\n  def execute(url:)\n    pages = Scanner.pages(url)\n    pages.each_with_index.map do |page, index|\n      progress \"Reading page #{index + 1} of #{pages.size}\", value: index + 1, total: pages.size # report progress\n      page.text\n    end.join(\"\\n\")\n  end\nend\n\nchat.with_tools(ReadReport).after_tool_progress do |tool_call, progress| # receives every update\n  puts \"#{tool_call.name}: #{progress.message}\"\nend\n```\n\nMCP tools report through the same callback. See [Reporting Progress](https://rubyllm.com/tool-execution/#reporting-progress).\n\n## OpenTelemetry Tracing\n\nOne line sends every model call, tool run, and workflow as a span to the tracing backend you already use:\n\n```\nRubyLLM::OpenTelemetry.enable\n```\n\nSpans join the current trace, so the HTTP call your tool makes nests under that tool. They carry metadata such as models, tokens, and tool names, and never prompts, responses, or tool arguments. The [OpenTelemetry guide](https://rubyllm.com/opentelemetry/) has the full span reference.\n\n## Smaller Changes\n\n- **Hetzner** is provider number nineteen:`RubyLLM.chat(model: \"Qwen3.8-27B\", provider: :hetzner)` .\n- **Per-tenant agents** : a`context` block can use the agent’s inputs, so each workspace can bring its own API key. Thanks to mikemikimike (#903).\n- **A secondary database** can hold RubyLLM’s supporting records next to your chats.\n- **Usage beyond chats** : in Rails, embeddings, transcriptions, and other one-shot calls go into the usage ledger too.\n- **`error.request_shape`** lists every turn’s parts and sizes when a provider rejects a request.\n- **Provider uploads are reused** across processes in Rails, so a large PDF isn’t uploaded again on every job.\n- **Perplexity chat runs on the Agent API** , since Sonar retires.\n- **Ruby 3.2 or later** is required, since 3.1 reached end of life.\n- **JSON 3 is allowed** , thanks to Filipe Kalicki (#968), who also made model lookups use an index (#981).\n\n## Upgrading from 2.0\n\n```\ngem \"ruby_llm\", \"~> 2.1.0\"\nbundle update ruby_llm\nbin/rails generate ruby_llm:upgrade\nbin/rails db:encryption:init   # only for MCP OAuth, if your app has no encryption keys yet\nbin/rails db:migrate\n```\n\nMost 2.0 apps need no code changes. The few that do are listed in the [upgrade guide](https://rubyllm.com/upgrading/), and everything new is in [What’s New in 2.1](https://rubyllm.com/whats-new-in-2-1/).\n\nThanks to everyone who sent code for this release: Andrii Furmanets, Andrey Samsonov, Andy Wang, Anton Kopylov, Filipe Kalicki, Guilherme Lages Santos, Islam Gagiev, Kieran Klaassen, Marc Köhlbrugge, mikemikimike, Mikhail Topolskiy, Muhammad Zain Ul Abidin, Paul Arterburn, and Viktor Schmidt.", "url": "https://wpnews.pro/news/rubyllm-2-1-mcp-judgments-evaluations-and-less-work-on-every-call", "canonical_source": "https://paolino.me/rubyllm-2-1/", "published_at": "2026-10-08 09:10:00+00:00", "updated_at": "2026-10-08 15:47:30.256583+00:00", "lang": "en", "topics": ["ai-agents", "agent-protocols", "developer-tools", "large-language-models", "ai-tools"], "entities": ["RubyLLM", "RubyLLM 2.1", "RubyLLM::MCP", "ruby_llm-mcp", "TypeSafe", "Jev", "Kieran Klaassen", "OpenTelemetry"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/rubyllm-2-1-mcp-judgments-evaluations-and-less-work-on-every-call", "markdown": "https://wpnews.pro/news/rubyllm-2-1-mcp-judgments-evaluations-and-less-work-on-every-call.md", "text": "https://wpnews.pro/news/rubyllm-2-1-mcp-judgments-evaluations-and-less-work-on-every-call.txt", "jsonld": "https://wpnews.pro/news/rubyllm-2-1-mcp-judgments-evaluations-and-less-work-on-every-call.jsonld"}}