Human Evaluators

Assign traces to your team, collect their judgments against a custom form, and aggregate the results.

Human evaluators are teammates assigned to answer specific questions on an evaluation. Use them for judgment calls a rule or a model shouldn't be trusted to make alone: tone, correctness on ambiguous cases, whether an answer actually resolves what the user asked.


Reviewing

Assigned reviewers work through their queue: each item shows the trace and only the questions assigned to them. They don't need to understand the underlying span data, just answer what you asked. Evaluations can carry a review deadline so a round of scoring completes on time: eval-wide for one-time and existing-trace evaluations, or per batch for recurring ones.


Results

As responses come in, Neatlogs aggregates them: overall summaries, per-question breakdowns, and per-batch views for recurring evaluations. Recurring batches can also generate a one-click AI summary report of the round, and you can export raw responses as CSV or JSON. This closes the loop between what your agent does in production and how good a human judges it to be.

Note

A question can be split between human reviewers and an AI evaluator on the same form. Assign whichever mix makes sense per question, not per evaluation.

On this page

Ask Neatlogs AI

Answers from the docs

How can I help?

Ask anything about instrumenting, tracing, or the Neatlogs dashboard.