About Me
I work on systems that let AI agents run the research loop: locating what is failing, proposing a change, running the experiment, and keeping the change only when a held-out metric improves. I have worked on this in the bounded AI4AI post-training loop at XYZ AI Lab, on RD-Agent at MSRA, and on my own project PaperFarm. I also work on the measurement side, where BenchScope quantifies how many independent signals a benchmark provides and Agent² RL-Bench tests whether LLM agents can carry out RL post-training themselves.
I am an undergraduate at Stony Brook University. Before that I spent two years at Sun Yat-sen University's HCP Lab under Dr. Keze Wang, where I learned to run a research problem end to end. I am always glad to hear about new collaborations.
Education
Experience
Work on XYZ-Aquila, a family of open-weight Deep Search agents, within the lab's AI4AI post-training loop, where agents diagnose failures and implement scoped changes under a fixed resource budget, with every step logged so it can be replayed. Contributor to the AI4AI at Scale technical report. LinkedIn 🤗 Hugging Face
Research on RD-Agent, Microsoft's open-source framework for automating the data-driven R&D loop.
Projects
A full-pipeline system for improving the agentic capabilities of LLMs, built out in XYZ-Aquila, two open-weight Deep Search agents based on Qwen3.6-35B-A3B and Qwen3.5-397B-A17B. Humans fix the optimization contract: target capability, private benchmark, allowed interventions, and resource budget. Agents then diagnose failures and implement scoped changes across data, learning, runtime, tools, and infrastructure, and a separate held-out evaluator decides what is kept. XYZ-Aquila-mini leads every benchmark column in the reported <40B open-weight table, and the weights and SFT dataset are released openly. Technical Report (PDF) 🤗 Weights & Data
Runs an unattended research loop on an existing repository through a Scout → Manager → Critic → Experiment cycle: it proposes changes, runs them, and keeps those that improve the target metric. Each experiment is an isolated git commit that rolls back automatically on failure, parallel workers run in separate git worktrees across GPUs, and a terminal dashboard tracks progress. Works with whichever coding agent is installed: Claude Code, Codex CLI, Aider, OpenCode, Kimi CLI, or Gemini CLI. PyPI
Microsoft's open-source framework for automating the data-driven R&D loop, which proposes hypotheses, implements them as code, and runs the resulting experiments. Contributed as a research intern at MSRA.
101,000+ skills across 21 domains, extracted from academic papers (PMC, arXiv, eLife) and technical sources (GitHub, StackOverflow). Provides full-text search over offline SQLite bundles, and every skill links back to the source it was derived from. Described in the SkillCenter preprint. arXiv:2607.07676
Publications
-
AI4AI at Scale: A Full-Pipeline System for Enhancing LLM Agentic Capabilities
-
SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents
-
Agent² RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
-
BenchScope: How Many Independent Signals Does Your Benchmark Provide?
-
FAST-CAD: Fairness-Aware Self-Supervised Framework for Automated Non-Contact Stroke Diagnosis
-
ATOM: A Framework of Detecting Query-Based Model Extraction Attacks for Graph Neural Networks
-
Mitigating Cache Noise in Test-Time Adaptation for Large Vision-Language Models
-
TsKAN: A Transparent Architecture for Improving the Interpretability of Multivariate Time Series Forecasting
-
Orthogonal Filtering Alignment Network for Ship Detection in SAR Images Under Frequency Shift Interference
* Equal contribution.
Awards and Honors
- AAAI-26 Student Scholarship & Volunteer Program — Travel scholarship and selected student volunteer.
- China National Olympiad in Informatics (NOI) Winter Camp 2022 — Silver Medal (roughly comparable to USACO Platinum).
- NOIP 2021 — Provincial First Prize (Guangdong) (roughly comparable to USACO Gold).
- National College Student Mathematical Modeling Competition 2024 — Provincial First Prize.
Interview
Service
Conference Reviewer
Journal Reviewer