An in-the-wild benchmark for AI agents in the production harness.
-
Updated
Sep 18, 2026 - Python
An in-the-wild benchmark for AI agents in the production harness.
Gen AI Evaluation Toolkit on AWS is a flexible, cloud-native accelerator built on AWS serverless architecture that enables comprehensive evaluation of generative AI applications.
DataClawEval: A Benchmark for Engineering Data Agents in Real Industrial Harness
A SnitchBench-style benchmark inverted for the dark-forest problem: does a listener AI alert humans about an alien signal when alerting may doom humanity?
🚀 AI Evolution Factory - From evaluation tool to continuous AI self-improvement platform. Agentic evaluation, auto-finetuning, global P2P testing, and hardware telemetry for local LLMs
Open-source verification for AI systems. Declare expected behavior, check real runs, and keep the evidence behind each result
Autonomous Multi-Turn Agentic Metamorphic Fuzzer & Swarm Consensus Stress-Testing Engine. Formulates Martingale context drift estimators, metamorphic execution DAGs, and Sybil groupthink inoculation filters.
A small benchmark for agent skills, verification artifacts, and fresh-session resumability.
Runnable lab measuring invalidation/staleness in agent memory — paper: Are We Ready For An Agent-Native Memory System? (arXiv:2606.24775)
Two-tier honest-evaluation harness for agentic RTL design (research, WIP)
Proof-bound evaluator stress testing with oracle-witnessed reward-hacking exploits and replayable evidence.
Paired rollouts for group-relative RL of LLM agents in stochastic environments
Pipeline to investigate structured reasoning and instruction adherence in multimodal LLMs
The specification behind ESAC, a program of compact benchmarks that measure difficult capabilities without measuring budget. ESAC-GI is released and runnable; ESAC-AG has a harness and one of its seven categories. Items are generators rather than stored questions, and every score carries the version tag of the instrument that produced it.
Deterministic synthetic fixtures and hard-gate scoring for agentic biosafety safeguard routing.
Scorekeeper is a server that benchmarks AI agents and assistant platforms
A Harbor-format RL environment for pass@k estimation, with a verifier built to resist reward hacking. 17 tests, 10 adversarial solutions rejected, amd64 CI.
Measure AI agents’ performance with standardized tests across 314 tasks, 33 domains, and 4 difficulty levels for clear, reproducible comparison.
A GPU-backed Harbor/Terminal-Bench RL environment for data-poisoning defense: containerized task, oracle solution, and a 25-check verifier that rejects 17 reward-hacking solutions.
To associate your repository with the agentic-evaluation topic, visit your repo's landing page and select "manage topics."