Boqin Yuan

MSCS @ UC San Diego

prof_pic.jpg

b4yuan[at]ucsd.edu

I’m a Master student in Computer Science at UC San Diego, exploring the intersection of Machine Learning and Software Engineering. At UCSD, I work with STABLE Lab (Prof. Jishen Zhao) on AI Agent and ML systems research. My work has been accepted at top venues including ICML, DAC, and ICLR workshops, spanning AI agents, ML systems, and agent evaluation.

I’m an enthusiastic open-source contributor to AI evaluation benchmarks. I’m an active contributor to the Harbor adapter for Terminal-Bench, and a co-author of SkillsBench and Agents’ Last Exam, community benchmarks that measure how well LLM agents handle real, expertise-heavy work.

This summer, I’m joining Moody’s Analytics in San Francisco as an SDE Intern (AI Agent) on the Banking team, building and evaluating AI agents for financial services.

I was a full-time Machine Learning Engineer at CambioML (YC S23) from 2024-2025, working on vision language models (AnyParser) and Computer Use Agent (CUA) systems (Energent.ai).

Previously studied Mathematics & Computer Science and Statistics at University of Illinois at Urbana-Champaign (UIUC). I bridge research ideas and production-ready engineering.

Research Interests: Machine Learning, AI Agent, ML Systems, AI Evaluation, Reinforcement Learning

Email GitHub LinkedIn Google Scholar CV

News

May 19, 2026 This summer, I’m joining Moody’s Analytics in San Francisco as an SDE Intern (AI Agent) on the Banking team, building and evaluating AI agents for financial services.
Apr 30, 2026 AMA-Bench accepted to ICML 2026! :tada:
Mar 03, 2026 Diagnosing Retrieval vs. Utilization Bottlenecks accepted to ICLR 2026 MemAgents.

Selected Publications

  1. CAIS-W
    /assets/papers_image/clawtrace.png
    ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation
    Boqin Yuan, Renchu Song, Yue Su, Sen Yang, and Jing Qin
    In ACM CAIS 2026 Workshop on Agent Skills (Oral), 2026
  2. ICLR-W
    /assets/papers_image/agent_memory.png
    Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory
    Boqin Yuan, Yue Su, and Kun Yao
    In ICLR 2026 Workshop MemAgents (Poster), 2026
  3. ICML
    /assets/papers_image/ama.png
    AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications
    Yujie Zhao, Boqin Yuan, Junbo Huang, Haocheng Yuan, Zhongming Yu, Haozhou Xu, and 6 more authors
    In International Conference on Machine Learning (ICML), 2026
  4. DAC
    /assets/papers_image/RTL.png
    PRO-V-R1: Reasoning Enhanced Programming Agent for RTL Verification
    Yujie Zhao, Zhijing Wu, Boqin Yuan, Zhongming Yu, Hejia Zhang, Wentao Ni, and 3 more authors
    In Design Automation Conference (DAC), 2026