Alignment is a human-AI workflow: how Custom Views make agent review possible Arize AI aligned the automated evaluators for its Alyx AI assistant in Arize AX by sampling production sessions into a labeling queue and building a Custom View that places the full session transcript, tool sequence, errors, generated artifacts, and annotation rubric on one page for human reviewers. The company uses those human labels as ground truth to test and improve its evaluators, then monitors the resulting scores in production through a second, session-level Custom View. Automated evals can score every agent session, but they do not automatically reflect what your team considers a good outcome. That takes human judgment. To align our evals for Alyx https://arize.com/docs/ax/alyx , the AI assistant in Arize AX, we sample production sessions https://arize.com/blog/new-in-arize-ax-september-2026-updates/ into a labeling queue and ask reviewers to judge what actually happened. Those labels become the ground truth we use to test and improve our evaluators. The problem is that agent sessions are difficult to review. A single session can contain several user turns, tool calls, errors, status changes, and generated artifacts. Reconstructing the outcome from raw trace data https://arize.com/blog/agent-traces-without-llm-judge/ makes labeling slow and inconsistent. We solved that problem with a Custom View https://arize.com/docs/ax/observe/create-custom-views built for the labeling queue. It puts the full journey and annotation rubric on one page, making human review faster and more consistent. We then use those labels to align our evals and a second, session-level Custom View to monitor the resulting scores in production. In this post, we’ll walk through that loop, and how to use Custom View to queue a sample of production sessions, review them against the rubric, align the evals to those labels, monitor the scores, and repeat. Build better agents with Arize Trace, evaluate, and learn. Build agents that work with Arize AX and start tracing your runs today. Prefer open source? Try Arize Phoenix for self-hosted, open source agent observability https://arize.com/phoenix?utm source=blog&utm medium=referral&utm campaign=ax-inline-cta&utm content=alignment-is-a-human-ai-workflow-custom-views-inline-cta-phoenix . 1. Create a labeling queue We start by running session-level evals https://arize.com/docs/ax/evaluate/trace-and-session-evals on the Alyx production project, then add a mix of sessions to a labeling queue https://arize.com/docs/ax/evaluate/human-review labeling-queues : successes, failures, and ambiguous cases, where a reviewer could defend more than one label. The queue manages assignment and annotation. More importantly, it gives us the human ground truth needed to test whether our automated evals reflect the outcomes we actually care about. But the queue alone does not make a session easy to judge. An Alyx session might begin with a request to fix an evaluator, continue with “it still does not work,” and include dataset inspection, variable-mapping updates, task creation, and a final explanation. In the raw trace, a reviewer has to reconstruct that journey across many spans. 2. Make labeling faster with a Custom View For the labeling queue, we asked Alyx for a structured transcript of the session, with the rubric on the same record. That makes the same purpose-built layout the default review surface for every labeler, rather than something each reviewer has to find or configure. The difference is easiest to see side by side. Without a Custom View, the reviewer starts from the general-purpose session and trace interface. The evidence is available, but it is distributed across turns, spans, attributes, tool calls, and result panels. The reviewer has to navigate the trace and mentally reconstruct the user’s journey before applying the rubric. The before-and-after comparison above shows the same underlying session organized in two ways. In the default view, the record is optimized for exploring traces. In the Custom View, the transcript, completed work, tool sequence, errors, generated artifacts, rubric, and outcome label are placed together around the labeling decision. It preserves the evidence while reducing the navigation required to interpret it. Consider an evaluator-repair session. A confident final response does not prove the work was completed. The reviewer needs to see whether Alyx inspected the right evaluator, found the unmapped variables, mapped them to the correct dataset columns, and left the task runnable. The view makes that sequence easy to check. It also exposes partial success: Alyx may diagnose the right issue but fail to create the task, or create the evaluator while leaving mappings incomplete. Reviewers can label these cases with specific failure modes instead of forcing them into a simple pass or fail. Because everyone sees the same evidence in the same layout, subject-matter experts can review sessions without first becoming observability experts. When reviewers disagree, the question is whether the rubric is unclear or the record is missing evidence. 3. Use human labels to align the evals The labels are how the evaluator gets corrected. We compare each automated score with the human label https://arize.com/docs/ax/observe/take-action/annotate-traces and inspect disagreements. A mismatch may mean the evaluator is using the wrong criteria, the rubric is ambiguous, or the review surface omitted an important signal. We update the evaluator or rubric, run it again, and repeat until its results track human judgment https://arize.com/blog/measuring-human-llm-judge-alignment/ reliably. This is why preserving the underlying evidence matters. An AI summary is useful for orientation, but it can omit the detail that changes a label: a user clarification, a failed tool call, or an artifact that was never created. The Custom View stays scannable without turning the session into a lossy summary. 4. Monitor aligned evals at the session level Once the evals reflect our reviewers’ judgment, a second Custom View helps us monitor their results across complete production sessions https://arize.com/blog/session-level-evaluations-with-arize-ax/ . Our Session Evals Dashboard puts the routed user journey, headline judges, journey-specific evals, structural checks, and behavioral signals into one scan. Here, too, the before-and-after comparison shows why the Custom View matters. The default Sessions view is useful for finding and opening sessions, but it does not express how the eval suite fits together; a reviewer has to inspect individual columns and session details to understand which journey ran and whether the signals agree. The screenshot contrasts the default Sessions view with the Session Evals Dashboard Custom View. The default view presents the raw session list and standard fields. The Custom View groups the same results by their role in the evaluation system the user journey, headline judges, journey-specific evals, structural findings, and behavioral signals , so agreement, conflict, and missing coverage are visible in one scan. The two views now serve consecutive stages of the loop: - The labeling view answers what actually happened? - The eval view answers how are the aligned evals behaving across production sessions, and where should we investigate next? When the dashboard surfaces conflicts, regressions, or ambiguous cases, we can route those sessions back into the labeling queue and continue improving the evals. Start with one agent journey Pick one journey that matters, such as evaluator setup, support resolution, grounded research, or coding-task completion. Run your session-level evals, send a small mix of successes, failures, and ambiguous cases to a labeling queue, and create a Custom View around the evidence reviewers need. A prompt can be as direct as: - “Show each session turn as a chat bubble with latency.” - “Put the final answer and annotation rubric side by side.” - “Show the tool sequence, errors, and generated artifacts for this session.” Use the resulting labels to align your evaluators. Then monitor those evaluators with a session-level Custom View and send new disagreements back through the queue. Close the loop A session eval scores the conversation. Whether that score matches what your team would accept shows up only when someone labels the session from the evidence. The labeling view puts the journey, the artifacts, and the rubric on one record, including the partial failures a pass/fail label would hide. The session view is where you see whether the updated eval still agrees with that label. Try Custom Views for yourself today. https://arize.com/docs/ax/observe/create-custom-views