Production-grade LLM Evaluation & Benchmarking Framework - GPT-4, Claude, Gemini, Mistral. Accuracy, latency, cost, hallucination, reasoning metrics.
-
Updated
Aug 29, 2026 - Python
Production-grade LLM Evaluation & Benchmarking Framework - GPT-4, Claude, Gemini, Mistral. Accuracy, latency, cost, hallucination, reasoning metrics.
Code for NAACL paper When Quantization Affects Confidence of Large Language Models?
Enterprise-grade LLM evaluation framework | Multi-model benchmarking, honest dashboards, system profiling | Academic metrics: MMLU, TruthfulQA, HellaSwag | Zero fake data | PyPI: llm-benchmark-toolkit | Blog: https://dev.to/nahuelgiudizi/building-an-honest-llm-evaluation-framework-from-fake-metrics-to-real-benchmarks-2b90
LLM hallucination evaluation pipeline: two Claude models scored on TruthfulQA (817 questions, 38 categories) via AWS Bedrock and DeepEval LLM-as-judge, with per-category failure analysis and model comparison
A hallucination detection pipeline for Large Language Models (LLMs).
A deterministic honesty layer for LLM agents + a reproducible TruthfulQA eval that measures calibration honestly — reporting the wins and the costs.
Official code for "From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems" (IWSDS 2026)
Premise-direction NLI consensus for response selection in LLMs
Does inter-agent LLM disagreement signal hallucination? An 817-question TruthfulQA study with judge-based evaluation and a fair single-model baseline.
Evaluation of Llama-3.1-8B Base vs Instruct on TruthfulQA using few-shot prompting and automatic judge models
Prototype that asks several LLMs to review an answer and flags likely hallucinations, with a balanced TruthfulQA evaluation.
Does instruction tuning make language models more sycophantic? A paired causal study across Qwen, Llama, and Gemma on TruthfulQA, showing the effect is family-dependent in both magnitude and direction. 7,200 evaluations, 12 ATE estimates with paired t-tests and bootstrap CIs.
A tool to evaluate and compare local LLMs running on Ollama or LM Studio under identical conditions using deepeval's public benchmarks (MMLU, TruthfulQA, GSM8K).
Evaluating how semantically equivalent prompt variations affect LLM predictions and uncertainty estimates.
Truthfulness and confidence calibration of matched base and instruct Qwen models on binary TruthfulQA
A hybrid 4-stage pipeline for detecting and mitigating hallucinations in Large Language Models using RAG retrieval, self-consistency sampling, uncertainty estimation, and NLI-based verification. Evaluated on TruthfulQA.
Adaptive, probe-controlled activation steering that cuts LLM hallucination rate by ~10pp (64.6%→58.5%) on TruthfulQA — steers only when a real-time risk probe flags danger, unlike fixed-strength ITI/CAA/TSV baselines reproduced here for comparison.
Your LLM judge says a model is wrong. Can it show the false sentence? Evidence-backed grading, checked against human labels.
PT-GAT Transformer Diagnostics: task-relative hallucination diagnosis with adequacy triggers, evidence conditioning, and anti-collapse baselines.
Multilingual hallucination evaluation framework for Large Language Models across Indian languages using TruthfulQA, NLLB-200, and mechanistic interpretability.
To associate your repository with the truthfulqa topic, visit your repo's landing page and select "manage topics."