When designing EvalPort, the grader system was the hardest part to get right. Every eval framework has its own way of scoring LLM outputs β DeepEval uses metric classes, Promptfoo uses assertion objects, Inspect AI uses solver functions. We needed a system expressive enough to cover 90%+ of real-world eval needs, but simple enough that any framework could implement it.
The result: 11 grader types that carry their own semantics. A grader isn't just a name β it specifies its parameters, its model, its threshold. An eval suite is self-describing.
exact_match β Compare output to expected output, optionally ignoring case.
contains β Check if the output contains a substring.
regex β Match against a regular expression.
semantic_similarity β Embed output and expected output, compare cosine similarity against a threshold.
llm_judge β Use an LLM to evaluate the output against a prompt template. The most powerful grader.
json_schema β Validate that the output is valid JSON matching a JSON Schema.
json_path β Extract a value from JSON output using a JSONPath expression, then compare it.
code β Run a function to evaluate the output.
human β Defer to human review.
model_graded β Compare the output to a reference answer using a model.
custom β Escape hatch for graders not covered by built-in types.
A test case references graders by ID. Multiple graders can evaluate the same test case. The ResultSet records each grader's score separately.
Self-describing: An eval suite carries everything a framework needs to execute it.
Framework-agnostic: Any framework can implement any subset of grader types.
Extensible: The custom type lets frameworks bring their own graders.
Comparable: Results from different frameworks use the same grader IDs.
pip install evalport-sdk
npm install evalport-sdk
Spec: https://github.com/adhabnr-ux/evalport/blob/main/spec/SPEC.md