PhD candidate in Economics & Business at Universidad Autónoma de Madrid.
I study causal inference, reliable AI agents, and climate-policy evaluation. My work develops benchmarks for causal research workflows and audits whether agent evaluations measure the capabilities they claim to measure.
Homepage · Google Scholar · LinkedIn
An execution-grounded benchmark that checks whether LLM-generated causal-inference workflows execute and recover the correct numerical estimates. Under review.
Evidence-grounded auditing of identification assumptions in difference-in-differences studies. Accepted, ClimateNLP Workshop @ EMNLP 2026.
Paper · Code · Dataset · HF paper page
A trade-data API environment for evaluating agent tool use under faults and constrained budgets. The subsequent Double Measurement Confound study audits scaffold ownership, scoring validity, and reliability across seeds. Under review.
Audit paper · Code · Demo
A multilingual ESG-perception research project built around an auditable agentic LLM pipeline. The research repository is currently private while the project is under development.
- OfficeQA Agent: document-grounded retrieval and reasoning over U.S. Treasury Bulletin data.
- Climate Claim Classifier: a demo for structured classification of climate-related claims.
- PhD Knowledge Base Starter: a research knowledge-base workflow for Claude Code and Obsidian.
Green Comtrade Bench · Purple Comtrade Baseline · AgentBeats Leaderboard

