🤗 Founder & CEO of Saturday Robotics — a high-signal weekly research forum in San Francisco on robotics & world models, connecting frontier researchers, founders and investors across embodied intelligence (30+ sessions; 3,600+ Luma subscribers as of Oct 2026).
🧩 Program Committee member & session chair, Physical World Models for Scaling Embodied AI Workshop (PWMS2026) at IEEE/RSJ IROS 2026.
Physical AI researcher on World-Action Models, sim-to-real transfer, cross-embodiment policy learning.
Building evaluation-centric embodied AI systems spanning world models, agentic reasoning, real-world deployment.
Master’s in CS from Georgia Tech and Financial Mathematics from UChicago; executive education at Stanford GSB. Previously a Machine Learning Engineer on Tensor Auto’s L4 autonomy stack, and a Machine Learning Quant Researcher at Société Générale in Chicago.
A long-term thinker, resilient collaborator, and builder of high-impact AI systems.
Github: https://github.com/junfanz1/
Weekly in-person reading club & research forum in San Francisco on robot world models and embodied intelligence · Luma · X · YouTube · Discord
- 📈 Growth: 32 sessions through Oct 3, 2026. Luma subscribers grew from ~2,400 (early Aug) to 3,600+ (Oct 2026); @saturdayrobotic followers from 1,500+ (Jul) to 5,000+ (3,000+ LinkedIn, 2,000+ X).
- 🧩 IROS 2026 Workshop (PWMS2026): Program Committee member and session chair for Oral Presentations I–II and the WorldArena 2.0 Challenge.
- 🍾 IROS 2026 Robotics Research Night, Pittsburgh (held alongside IROS; not an official IROS event): ~400 registrations, ~200 attendees.
- 🏔️ CVPR 2026 Denver Research Night: ~600 applicants, ~300 admitted, 6 of 43 lightning-talk proposals selected; talks incl. NVIDIA Cosmos 3. Livestream
- 🤝 Co-hosted sessions with Samsung Research America (Session 20; keynote by Haoru Xue, UC Berkeley, with Prof. Kris Hauser) and Dyna Robotics (Session 28: chief scientist Jason Ma and team on Dyna-2’s 1M-hour scaling law). Recent sessions also covered representation & architecture in robot manipulation (Session 30) and a Dexterity Day (Session 32).
- 🎬 Organizer, ACM SIGGRAPH 2026 Birds of a Feather — “World Models for Robotics: Bridging Graphics, Simulation, and Physical Intelligence” (Los Angeles, Jul 22).
- 🛠️ Next: Robotics Hardware Hackathon (Oct 24, SF).
- 💬 Technical recaps on X liked and reposted by Yann LeCun.
- 🧭 Agents’ Last Exam (ALE): Benchmarking Long-Horizon AI Agents [NeurIPS 2026 · Evaluations & Datasets Track]
A benchmark of long-running professional workflows across 55 fields, built by a large multi-institution collaboration (400+ authors); I contributed as a quantitative-finance domain co-author. Accepted as a poster at NeurIPS 2026, and the first benchmark cited in OpenAI’s GPT-5.6 release announcement (Jul 9, 2026). Website - 🤖 GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling [ICCV 2025]
Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu, Junfan Zhu, Chenliang Xu
🎭 A latent shortcut model compresses multi-step diffusion sampling into a few steps, with spatial-temporal modeling of coordination across body parts and over time — keeping generation quality while greatly reducing inference cost for speech-driven gestures. - 🚗 IEDD: An Interactive Enhanced Driving Dataset for Autonomous Driving [Scientific Data 2026]
Haojie Feng, Xinrui Zhang, Mengjie Tian, Peizhi Zhang, Zhuoren Li, Junpeng Huang, Xiurong Wang, Junfan Zhu, Jianzhou Wang, Dongxiao Yin, Lu Xiong
🌲 As AutonomousDriving evolves toward VLA, sparse interactive scenarios and weak multimodal alignment remain critical bottlenecks. Existing datasets heavily bias toward straight-line cruising while severely under-representing long-tail interactive events (cut-in, merging, pedestrian crossing, head-on avoidance). IEDD introduces a physics-aware, interaction-dense dataset (plus IEDD-VQA multimodal extension) mined from 7.31M ego-centric scenes across Waymo, nuPlan, Lyft, INTERACTION, SIND — with 91% multi-agent interactions, dual Intensity–Efficiency metrics, pixel-level BEV-video alignment, rule-based hallucination-free language, and hierarchical L1–L4 VLM benchmarking. 🌍 It lays a scalable, causality-grounded foundation to evolve general-purpose VLMs into truly capable autonomous driving experts. 🤗 HuggingFace, LinkedIn, X. - 🛣️ SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models [arXiv 2026]
Haojie Feng, Peizhi Zhang, Xinrui Zhang, Zhuoren Li, Junpeng Huang, Xiurong Wang, Dongxiao Yin, Yuxiang Zhang, Junfan Zhu, Lu Xiong, et al.
🧪 A follow-on to IEDD. Instead of comparing driving VLA models on separately chosen datasets, SSP builds matched scenarios for the same safety-critical interaction event across synthetic, simulated and physical domains. Tested on cut-in and crossing scenarios with three VLA systems (scores 0.259–0.325), it challenges the common assumption that performance is necessarily better in the physical domain. - 📊 QuantEval: A Benchmark for Financial Quantitative Tasks in Large Language Models [arXiv 2026]
Zhaolu Kang, …, Junfan Zhu, …, Richeng Xuan (18 authors)
🧪 Evaluation and domain knowledge are the core bottlenecks of Quant + AI. Without expert-level, strong verifiers for evaluation, models cannot reliably assess performance in multi-step strategy generation, risk control, or real-world trading effectiveness. QuantEval is proposed in this context, providing a reproducible benchmark framework that goes beyond static question answering and shifts toward evaluation grounded in realistic trading details. It represents an initial exploration of evaluating financial “World Models.” 🌍
- Finalist & Track Winner, 🏆 Y Combinator Hackathon 2025
Inspired by Isaac Asimov’s Foundation, PsychoHistory is a probabilistic forecasting system that maps the branching futures of human events—combining history, data, and AI to model the flow of possibility. 🧠 Our approach blended SFT+RL, training the model not just what to predict but how to reason across alternative futures—like a psychohistorian trained on uncertainty itself. - Meritorious Winner, Mathematical Contest in Modeling.
- Finalist, Asia Supercomputer Challenge.
- Top 10 Algo Trader, Rotman International Trading Competition.
- Outstanding Thesis (1%).
- Program Committee Member & Session Chair, Physical World Models for Scaling Embodied AI Workshop (PWMS2026), IEEE/RSJ IROS 2026.
- Organizer, ACM SIGGRAPH 2026 Birds of a Feather — “World Models for Robotics: Bridging Graphics, Simulation, and Physical Intelligence”.
- Reviewer, NeurIPS 2026 Evaluations & Datasets Track (3 papers) · EMNLP 2026 Industry Track (2 papers) · AAAI 2027 · ACM CAIS 2026 AgentSkills Workshop (nominated by the program chairs).
- Book-Proposal Reviewer, Manning Publications — two proposals on world models (2026).
- Judge, Embodied Metal Hackathon 2026 (SF; robot learning, VLA, humanoid manipulation) · Silicon Valley Robotics Fair 2026 (AI Ideathon session).
- Invited Speaker & Moderator, AUTONOMOUS 2026 — “Rebuilding the Factory: Physical AI on the Production Line”.
- Moderator, Robotics & Data Summit, Robotics Center of Silicon Valley — “The Data Wars: Collect, Synthesize, or Scrape?”.
- Invited Podcast Guest, Innovator Coffee EP-38 — “World Models: The Missing Layer Between AI and the Physical World”.
- Member, IEEE · IEEE RAS Technical Committee on Robot Learning.
My portfolio boasts pioneering projects in MoE & Attention for scalable LLM, reflective multi-agent orchestrations, and full-stack GenAI applications.
- 1. Awesome-AI-Engineer-Review
In-depth review of industry trends in AI, LLMs, Machine Learning, Computer Science, and Quantitative Finance. - 2. Agentic RL: GRPO Reinforcement Learning for Agentic Search in LLMs
Search-R1 leverages Group Relative Policy Optimization (GRPO) to fine-tune LLMs at the token level, enabling stable reinforcement learning over multi-step search–reasoning trajectories. The model learns adaptive retrieval policies, deciding when to trigger searches and integrating results into its reasoning context for more precise answers. - 3. MiniGPT-and-DeepSeek-MLA-Multi-Head-Latent-Attention
Memory-efficient multi-head latent attention in PyTorch, that leverages low-rank approximation and decoupled rotary positional embeddings, to compress key–value representations, reducing inference memory while maintaining high performance in long-context language models. - 4. DeepSeek-MoE-Mixture-of-Experts-in-PyTorch
Implemented scalable 8-expert MoE model with top-k routing, expert load balancing, and capacity-aware gating; enabled parallel sparse activation and DeepSeek-R1-style distributed training scalability. - 5. MCP-MultiServer-Interoperable-Agent2Agent-LangGraph-AI-System
A decoupled real-time agent architecture connecting LangGraph agents to remote tools served by custom MCP servers via SSE and STDIO, enabling a scalable multi-agent system for LLM workflows. The design supports flexible multi-server connectivity and lays the groundwork for an Agent2Agent protocol, fostering seamless, cloud-deployable interoperability across diverse AI systems. - 6. LangGraph-Reflection-Researcher
Engineered LangGraph-based multi-agent system with self-reflection and retrieval-grounded alignment; integrated LangSmith trace for reasoning introspection, cutting hallucination 40% with iterative expert routing. - 7. Cognito-LangGraph-RAG-Chatbot
Advanced Retrieval Augmented Generation (RAG) chatbot that utilizes LangGraph to enhance answer accuracy and minimize hallucinations in LLM outputs. - 8. Cursor-FullStack-AI-App
Cursor Vibe Engineering: Full-stack micro SaaS AI application that processes GitHub URLs to generate insightful JSON reports powered by AI analytics. - 9. Cryptocurrency-Blockchain-FullStack
Comprehensive decentralized blockchain platform demonstrating practical applications of core blockchain concepts through a modular, full-stack approach.

