Simon Yu
You can also call me U Chi Lok (余知樂) or Simão (in Portuguese)
I am a 3rd year PhD student at Northeastern University, advised by Weiyan Shi. My research goal is to build self-improving agent systems via self-play and interaction with real-world feedback. I closely work with Chris Manning from Stanford and Natasha Jaques from UW. I am currently interning at MSR Redmond with Baolin Peng and Jianfeng Gao, working on AI for AI and self-improvement. Before that, I interned at Orby AI, mentored by Peng Qi.
I work toward this goal from three angles:
- Self-Play and RSI: enabling agents to improve through self-play and build their own training environments (AutoEnvScaling for automating the data flywheel; SPADE for self-play in generated environments; SCOPE for population co-training for user simulator; SPIRAL for reasoning through zero-sum games).
- Meta-Agents: Shepherd turns an agent’s execution into a reversible, Git-like trace, so meta-agents can inspect, fork, replay, and revert other agents’ runs to supervise, optimize, and train them.
- Environment Scaling & Continual Learning: scaling what agents learn from and what they keep, including TextArena for multi-agent environments and evaluation, GEM for unified, scalable environment generation, and PolySkill for continually aggregating experience across new domains.
One of the most influential lessons to me is from The Bitter Lesson by Richard Sutton and The Era of Experience by David Silver and Richard Sutton. The idea is not just limited to AI but can be applied to any choice in life. Always choose the path that benefits in the long run, instead of the path that might be easier in the short run.
news
| Sep 29, 2026 | New! New preprint: AutoEnvScaling: Automating the Data Flywheel with Terminal Agents. We turn data flywheel into a terminal task where agents can build their own training environments. We show RSI on Qwen3.6-35B-A3B on TB2.1 and TB4.0. |
|---|---|
| Sep 28, 2026 | New! I recently gave talks on Shepherd at the CAIRD Workshop, co-hosted by FAR.AI and CBAI, as well as at Tencent and Apple. |
| Sep 24, 2026 | New! 2 papers accepted to NeurIPS 2026: Shepherd: Enabling Programmable Meta-Agents via Reversible Agentic Execution Traces and Coding with “Enemy”: Can Human Developers Detect AI Agent Sabotage?. See you in Sydney! |
selected publications
- EMNLPIn The Conference on Empirical Methods in Natural Language Processing, 2026
- NeurIPSIn NeurIPS, 20262.4K+ GitHub Stars. Featured on X (650K+ Total Views): @_avichawla · @akshay_pachaar
- ICMLIn ICML, 2026X Announcement (330K+ Views). Selected organic media coverage: Forbes, VentureBeat, TLDR AI, Yahoo News, The Neuron, Analytics Vidhya
- ICLR
- Arxiv