Crowdsourcing of hard to translate inputs (texts, images, audios) at scale.
-
Updated
Oct 6, 2026 - Python
Crowdsourcing of hard to translate inputs (texts, images, audios) at scale.
Self-hosted LLM chatbot arena, with yourself as the only judge
Find informative examples to efficiently (human)-evaluate NLG models.
开源 AI 应用评测平台,支持 RAG、AI Agent、多轮对话、LLM-as-Judge、接口评测、评测报告和人工盲测。Open-source AI evaluation platform for RAG, AI Agents, multi-turn conversations, LLM-as-Judge, endpoint evaluation
Concept-Guided Chain-of-Thought (CGCoT) pairwise annotation tool for systematic text evaluation using LLMs. Generate breakdowns, compare items, compute scores, and validate against human judgments. Supports Ollama, Hugging Face, Google Gemini, OpenAI, and Anthropic models.
Official repository for Deep Research Comparator: A Platform For Fine-grained Human Annotations of Deep Research Agents
Code and data for paper "Achieving Reliable Human Assessment of Open-Domain Dialogue Systems"
Multidimensional Evaluation for Text Style Transfer Using ChatGPT. Human Judgement as a Compass to Navigate Automatic Metrics for Formality Transfer (HumEval 2022)
Code for "QE4PE: Word-level Quality Estimation for Human Post-Editing" ✍️
Success and Failure Linguistic Simplification Annotation 💃
Personalized Chinese humor generation research: human preference experiments, blind evaluation, negative results, and a Next.js + Supabase web prototype.
人工视频Track辅助:离线人工审核目标跟踪输出视频,支持问题分类、断点续审、兼容转码和汇总导出。
A reproducible benchmark for evaluating AI design agents across 7design scenarios. Double-blind SbS voting · 140 tasks · Bootstrap CI
A collection of MTurk templates designed to make complex tasks easier for human annotators.
Research repository accompanying a Master's thesis on automatic Japanese haiku generation and human evaluation across multiple Large Language Models.
Code for evaluating automatic correction of language-learner text, including human-evaluation and metric-comparison tooling.
Chatbot for IIIT Nagpur using Fine Tuning and RAG
Traceable bilingual Amazon review insight agent for cross-border beauty operations
CS 685 Advanced Natural Language Processing Project: Learning Schematic and Contextual Representations for Text-to-SQL Parsing
To associate your repository with the human-evaluation topic, visit your repo's landing page and select "manage topics."