Sitelet https://github.com/yonghongzhang-io
Skip to content
View yonghongzhang-io's full-sized avatar
🎯
Focusing
🎯
Focusing

Block or report yonghongzhang-io

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
yonghongzhang-io/README.md

Yonghong Zhang

PhD candidate in Economics & Business at Universidad Autónoma de Madrid.

I study causal inference, reliable AI agents, and climate-policy evaluation. My work develops benchmarks for causal research workflows and audits whether agent evaluations measure the capabilities they claim to measure.

Homepage · Google Scholar · LinkedIn

Selected research projects

An execution-grounded benchmark that checks whether LLM-generated causal-inference workflows execute and recover the correct numerical estimates. Under review.

Paper · Code · Dataset

Evidence-grounded auditing of identification assumptions in difference-in-differences studies. Accepted, ClimateNLP Workshop @ EMNLP 2026.

Paper · Code · Dataset · HF paper page

A trade-data API environment for evaluating agent tool use under faults and constrained budgets. The subsequent Double Measurement Confound study audits scaffold ownership, scoring validity, and reliability across seeds. Under review.

Audit paper · Code · Demo

🌍 GLEAM (work in progress)

A multilingual ESG-perception research project built around an auditable agentic LLM pipeline. The research repository is currently private while the project is under development.

Other tools

Supporting benchmark infrastructure

Green Comtrade Bench · Purple Comtrade Baseline · AgentBeats Leaderboard

Pinned Loading

  1. comtrade-openenv comtrade-openenv Public

    ComtradeBench OpenEnv: execution-grounded benchmark for reliable LLM tool-use under adversarial trade-data API conditions.

    Python 2

  2. purple-agent-officeqa purple-agent-officeqa Public

    A2A retrieval agent for Treasury Bulletin OfficeQA: document-grounded QA over public finance tables and reports.

    Python

  3. climate-claim-classifier climate-claim-classifier Public

    An LLM/NLP demo: classifying climate-related claims (type, carbon-market relevance, specificity, evidence needed)

  4. phd_kb_starter phd_kb_starter Public

    One-command installer for an LLM-powered personal knowledge base, following Andrej Karpathy's 8-stage workflow. Built for Claude Code + Obsidian.

    Shell 1 2

  5. agentbeats-leaderboard-v2 agentbeats-leaderboard-v2 Public

    Leaderboard infrastructure for the ComtradeBench / AgentBeats agent-evaluation benchmark: task definitions, submission flow, and scoring.

    Python 2

  6. green-comtrade-bench-v2 green-comtrade-bench-v2 Public

    Deterministic offline ComtradeBench judge for evaluating agent robustness under pagination, retries, duplicates, page drift, and totals traps.

    Python 1