Finding the right AI models for specific tasks using LLMs as a ‘judge’ Northeastern University computer science student Ryan Amiri used large language models as a "judge" during a co-op at HubSpot to evaluate the quality, relevance and correctness of outputs from popular AI agents, producing pass-or-fail labels with reasoning on criteria including tone, accuracy and safety. Northeastern professor Aleks Gollu said Amiri "was able to leverage the LLM, and what used to be probably five or 10 people doing this, do it by himself," helping HubSpot identify which AI tools best suit tasks such as fetching YouTube videos or reading a user's Gmail inbox. HubSpot representatives did not respond to a request for comment on Amiri's contributions prior to publication. Finding the right AI models for specific tasks using an LLM as the ‘judge’ This student used large language models as a “judge” to evaluate customers’ problems with some of the most popular AI models. OAKLAND – The same abilities that make AI agents useful – their autonomy, intelligence and flexibility – can also make them difficult to evaluate for accuracy. This according to Anthropic https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents , whose flagship product is the AI assistant Claude. AI agents perform many functions in rapid sequence as they attempt to interpret and respond to a user’s prompt, but their responses are not deterministic and are often prone to error. That’s why evaluating these tools to ultimately improve them is a complex process, and involves separating the good responses from the bad using a mixture of qualitative and quantitative review types, experts say. “A bad tool is one where the description is bad, if the inputs are not accurate, if the input descriptions are too short,” said Ryan Amiri, a Northeastern University computer science student who began his college journey in Oakland. “The agents reading them are like, I don’t know what this tool is supposed to do. That’s where their performance becomes a detriment.” Amiri, who grew up in Orange County, California, said he wanted to combine his longstanding passion for technology with his interest in this type of evaluatory problem-solving as he searched for co-op opportunities in the AI industry. Amiri found a position as a software and machine learning engineer at HubSpot, a data broker and customer relationship management firm https://www.hubspot.com/ . He used large language models LLMs – AI systems trained on text to understand and generate human-like responses – to evaluate a variety of problems experienced by customers when interacting with some of the most popular AI models, each of which has its own strengths and weaknesses, Amiri said. Using the LLMs as a “judge,” Amiri helped HubSpot to assess the quality, relevance and correctness of outputs produced by the AI agents when performing particular tasks. The result was a better understanding of which AI agents were best suited for a given task – which are called tools – such as fetching YouTube videos or reading a user’s GMail inbox, and identifying areas for improvement that the AI companies could draw from, Amiri said. “ Amiri was able to leverage the LLM, and what used to be probably five or 10 people doing this, do it by himself,” said Aleks Gollu https://damore-mckim.northeastern.edu/people/aleks-gollu/ , professor of entrepreneurship and innovation on Northeastern’s Oakland campus. Part of Amiri’s success came from his understanding of how to fulfill users’ requests based on their individual needs, Gollu added. Representatives at HubSpot didn’t respond to a request for comment on Amiri’s contributions to the company prior to publication. Just like how companies may want different levels of precision when using GPS to locate their assets – think a librarian locating a single book on a shelf, compared to a dockworker looking for a massive stack of shipping containers – users want AI agents to tailor results to their specific needs, Gollu said. One size does not necessarily fit all. That’s why testing AI agents for quality means tapping a wide range of review types, he said. These can include user feedback, structured human studies and manual transcript review, in which humans read through agent conversations and look for errors. LLM-as-judge is another review method, one that applies rigid grading logic to agents’ responses and scores them on a rubric, Amiri said. At HubSpot, Amiri fed prompts and responses into his LLM judge, which evaluated the text using explicit instructions and criteria, such as the agent’s tone, accuracy and safety. The judge then generated a pass or fail label along with reasoning for its decision, which Amiri took to HubSpot’s creation studio. The engineers there could then decide to reject these tools, disallowing them from responding to such a request until their capabilities could be improved, or approve them and add them to the agents’ toolkit. Amiri said the six-month co-op forced him to grow as an engineer and translate the concepts he gained in Gollu’s class from theory into practice. It also helped him secure a summer experiential learning assignment as a research engineer at Ramp Labs, the hub of AI research at corporate finance and expense management automation company Ramp. As part of his internship, Amiri helped Ramp launch a new product called Ramp Router https://router.com/ , which directs customers to the most capable AI model for a given task based on the lab’s own evaluation and judgement processes – similar to those employed by Amiri at his co-op. “We are trying to provide an accurate assessment of how the models are currently being used by every single frontier lab right now,” Amiri said. His performance at Ramp led his manager to give him a return offer as a mid-level engineer, he added. Representatives at Ramp didn’t respond to a request for comment on Amiri’s performance at the company prior to publication. “He is that model example student who understands what’s at stake during the four years in his education, and is always thinking one step ahead,” Gollu said. “Clearly he is good enough to go all the way.”