MARLlib: a scalable, efficient multi-agent RL library
A widely-used library for multi-agent reinforcement learning research.
Journal of Machine Learning Research · 2023
Founder & CEO, Parallight. AI scientist & serial founder. A decade at the intersection of frontier AI research, production AI at scale, and AI education — now building the AI mastery platform for those who'll define the next era.
In his words
The bottleneck moved off the model. What separates people now isn't which model they can access — it's whether they can command one to do real, complex work, and whether they have the compute to let it run.
I've spent a decade on both halves of that: researching the agents themselves, and teaching thousands of engineers to wield them. Parallight is where those two lines meet.
Machines that work, so people may build.
— Marvin Gao
The through-line
Marvin Gao is the founder and CEO of Parallight (KeploreAI Inc). His research is on the exact problem Parallight teaches — agentic systems, multi-agent reinforcement learning, and robot learning, published at JMLR, AAAI, ICRA, and TPAMI. Before Parallight he built production AI at IBM, Ant Financial, and Alibaba, and founded an AI-education platform acquired by a unicorn. More than 2,000 engineers have trained with him.
A widely-used library for multi-agent reinforcement learning research.
Journal of Machine Learning Research · 2023
Prompt-learning that carries a policy from simulation to the real world.
AAAI Conference on Artificial Intelligence · 2024
First-author work on improving pre-trained robot policies. (Johns Hopkins.)
IEEE ICRA 2025
A game-theoretic framework for evolving diverse red-team LLMs.
IEEE TPAMI · & the original "Red Teaming Game" (2023)
Band-8, IBM China Intelligent Service
Built China Construction Bank's automatic Q&A robot (live in Shanghai, Guangzhou, Shenzhen); nationwide route optimization for bank cash carriers; model-based solar-PV generation forecasting for TBEA.
Alipay risk
Risk-classification models over Alipay transfers; awarded a national invention patent for potential-risk-word discovery.
Automatic news abstraction powering Tmall Genie and Rokid smart speakers; speech-map extraction of viewpoints from news.
Multi-agent reinforcement learning — a unified, reproducible MARL environment and framework.
A platform for experienced engineers to learn ML/AI algorithms, practice, and deployment. Acquired by Huike Group.
A Chinese ed-tech unicorn (valued over $1B). Built one of the largest and most recognized online AI-education programs in China.
The AI mastery platform for those who'll define the next era, built by a Silicon Valley team in San Mateo, California.
Master of Science, Computer Science
Research in the Interactive Computing Robot Lab — visual-language models applied to robot policies.
Computer Science & Software Engineering
Knowledge Graph Laboratory and the State Key Laboratory of Optoelectronics.