Financial research agents need evaluation frameworks that reflect institution-specific standards and fix values as of an information cutoff. Fixed benchmarks with hand-written rubrics are expensive to extend and cannot encode proprietary evaluation criteria. FinAutoRubric solves this by letting experts write reusable guidance once, then generating per-query rubrics automatically through a multi-agent loop with code-enforced validation.
When you deploy a financial research agent, you need to know whether its answers meet your institution's standards. Generic benchmarks do not help because:
Existing expert-reviewed finance benchmarks rely on fixed, per-item rubrics. Each rubric is written once for a specific query and cannot be reused. Extending the benchmark requires writing new rubrics from scratch.
FinAutoRubric separates reusable expert guidance from query-specific rubric generation. The architecture has three layers:
def generate_rubric(query, expert_guidance, task_bank):
"""
Generate a rubric for a financial research query.
Args:
query: The research question (e.g., "What is AAPL's P/E ratio?")
expert_guidance: Prompts and rules from analysts
task_bank: Reusable criteria library
Returns:
Rubric with expected values and scoring criteria
"""
draft_rubric = writer_agent.generate(
query=query,
guidance=expert_guidance,
task_bank=task_bank,
information_cutoff="2026-09-28"
)
review_result = reviewer_agent.validate(
rubric=draft_rubric,
guidance=expert_guidance,
sources=draft_rubric.sources
)
if review_result.confidence < THRESHOLD:
return escalate_to_human(draft_rubric, review_result)
return draft_rubric
The writer agent queries financial data sources, computes expected values, and drafts rubric criteria. The reviewer agent checks whether the expected values match the sources and whether the rubric follows expert guidance. If the reviewer's confidence falls below a threshold, the system escalates to a human analyst.
The expert guidance layer solves the problem of encoding proprietary standards without leaking them into model training data. Analysts write:
These artifacts live in your infrastructure, not in the model. The agent reads them at runtime, so you can update standards without retraining.
The Task Bank stores reusable criteria templates:
| Criterion Type | Example | Reusability Scope |
|---|---|---|
| Data source preference | "Use Bloomberg for real-time prices" | All price queries |
| Temporal constraint | "Fix values as of market close" | All time-sensitive queries |
| Calculation method | "Use trailing twelve months for P/E" | All ratio queries |
| Compliance check | "Verify data is not material non-public" | All queries |
When generating a rubric, the writer agent pulls relevant criteria from the Task Bank and instantiates them for the specific query.
Financial data changes constantly. A rubric must fix the expected value as of a specific date to prevent temporal drift. FinAutoRubric handles this in three ways:
information_cutoff timestamp.
If the writer agent cannot find a source dated before the cutoff, it flags the criterion as unverifiable and escalates.
The core challenge is validating that an auto-generated rubric matches expert judgment without requiring the expert to review every rubric. FinAutoRubric uses three mechanisms:
The paper reports that on three expert-authored finance benchmarks, FinAutoRubric's rubrics track expert scoring as closely as the strongest evaluated generator. In-house analysts preferred the auto-generated rubrics in a blind review.
| Failure Mode | Cause | Mitigation |
|---|---|---|
| Hallucinated expected value | Writer agent invents a number not in sources | Reviewer agent checks source citations; escalate if confidence is low |
| Stale data | Writer agent uses data after cutoff | Source versioning + reviewer validation |
| Misaligned criteria | Rubric does not match expert guidance | Code-enforced rules + Task Bank templates |
| Reviewer false positive | Reviewer approves a bad rubric | Human spot-checks on escalated cases; log all rubrics for audit |
The escalation mechanism is critical. If the reviewer agent's confidence is below threshold, a human analyst reviews the rubric before it is used. This creates a feedback loop: analysts see which rubrics fail and update the expert guidance or Task Bank accordingly.
A production deployment needs:
The paper's released FinAutoRubric Benchmark includes 100 queries across 78 tasks and eight asset classes. Rubrics generated by an earlier model generation still leave headroom for a later one, which suggests that the framework is extensible as models improve.
You need to monitor:
Log every rubric generation with full provenance: query, expert guidance version, Task Bank version, writer agent output, reviewer agent output, and final rubric. This audit trail is essential for compliance and debugging.
Use FinAutoRubric when:
Avoid it when:
The framework trades generation latency for extensibility. If you need to evaluate thousands of queries per second, pre-generate rubrics offline and cache them. If you need real-time evaluation, this approach will bottleneck on the writer and reviewer agents.