Abhishek Chandwani
Co-founder, Metaphi AI
I build reinforcement-learning environments and evaluation infrastructure for long-horizon enterprise agents: the systems that decide whether an agent actually did the job. My work covers LLM-as-judge evaluation, programmatic verifiers, environment fidelity, rubric construction, and evaluation harnesses for tool-using agents.
Before Metaphi: BCG and Shell. MBA, Harvard Business School.
Interests
Persona simulation and adversarial persona vectors. Recursive self-improvement. Agent behavior analysis. Scaling agent verifiability.
Research
-
LH-BencharXiv · 2026
Skill-grounded evaluation of long-horizon agents on subjective enterprise tasks. Expert-authored rubrics against observable task artifacts, validated against human preference judgments.
-
COBOLBenchbenchmark
Verifier-backed benchmark for frontier coding agents on realistic enterprise COBOL maintenance tasks. Executable correctness checks over recovered production estates.
-
SimHubwhitepaper
Continual-learning and evaluation harness for agents in production: traces, evaluation, feedback, and improvement loops in one system.
-
EnterpriseSWEforthcoming
Enterprise software-engineering environments for agents.
-
Optimum portfolio construction using genetic algorithmIJSAEM · 2015
Genetic-algorithm construction of optimum stock portfolios from the S&P 500. With Pankaj Sinha and Tanmay Sinha. Int. J. of System Assurance Engineering and Management (Springer).