Human evaluation of LLM performance made simple

Upload
CSV with reference + generated
Score
Review each row and rate what matters
Export
Download the completed CSV anytime

Start a new evaluation

Upload a CSV with reference (ground-truth) and LLM-generated text.
CSV format. Include two columns: reference_text and generated_text. Aliases (reference, ground_truth, prediction, response) are accepted. Extra columns are preserved on export.
Drop your CSV here
or click to browse · max 10 MB
Keep setup in one flow
Choose your dataset source first, then review the built-in matrices, add custom ones only if needed, and start the evaluation.
Private by design. We do not permanently store uploaded data. Files are used only for the current evaluation session and export workflow.

Core HumanELY metrics

Base rubric

These five built-in metrics are always included. You can also add custom metrics on the left with their own definition and 1–5 scale.

Relevance
How accurate the response is in content, reasoning, and helpfulness.
Coverage
How completely the response covers the key topics and content from the reference.
Coherence
Fluency, grammar, and organization of the generated content.
Harm
Bias, toxicity, privacy, and hallucinations in the generated response.
Comparison
How the generated response compares to human and alternate-LLM responses.
Before adding custom metrics. Explore our matrices, then add only the extra rubric you still need.
Optional custom metrics
Add any extra rubric you want reviewed. Each custom metric includes its own definition and 1–5 scale labels.
No custom metrics added yet. Use Add custom metric above when you need one.