cv
Education, research experience, selected projects, and awards. Also available as a PDF.
General Information
| Name | 黄泓嘉 / Hongjia (Alex) Huang |
| Affiliation | Robotics Institute, Carnegie Mellon University |
| hh3043@nyu.edu | |
| Research Interests | Robotics, computer vision, and physics-grounded video generation |
| Profiles | |
Education
-
2026– Pittsburgh, US
M.S. in Robotics (MSR)
Carnegie Mellon University Robotics Institute -
2022–26 Shanghai, China
B.S. in Computer Science and Mathematics (double major)
New York University Shanghai Summa Cum Laude - Overall GPA: 3.93 / 4.0
- Selected coursework
- Graduate: Natural Language Processing with Representation Learning, Inference & Representation, Reinforcement Learning
- Computer Science: Machine Learning, Parallel Computing, Algorithms, Operating Systems, Data Structures
- Mathematics (Honors): Ordinary Differential Equations, Analysis, Numerical Analysis, Linear Algebra
- Mathematics: Partial Differential Equations, Linear and Nonlinear Optimization, Probability & Statistics
Research Experience
-
2024– Shanghai, China
Auxiliary Streams for Video-Diffusion World-Action Models
New York University Shanghai Independent research · Advised by Prof. Shengjie Wang (NYU Shanghai) and Prof. Tianyi Zhou (UMD) - Investigating whether training-time auxiliary structure can improve the policy of a frozen Wan2.2-5B video-diffusion world-action model, particularly its spatial out-of-distribution generalization.
- Added inference-free alignment teachers (VGGT, CoTracker, DINO/REPA) and co-denoised streams routed by a Mixture-of-LoRA-Experts, passing 3D structure from 2D video to the action head at test time.
- Earlier in the project, adapted pretrained video diffusion models with lightweight modifications so that generated video obeys physical laws, conditioning on physical constraints through a modified ControlNeXt and generating supervised training data with 3D physics engines (Genesis).
-
2025– Maryland, US
Physics-Informed Vision-Language-Action Models
University of Maryland, College Park Independent research · Advised by Prof. Furong Huang and Prof. Tianyi Zhou - Integrating trajectory planning into the action-prediction objective of VLA models such as GR00T N1.5.
- Using a 4D encoder that fuses multiple views across multiple timesteps to propagate spatial and temporal information jointly, aiding action prediction via future latent alignment.
- Designing an implicit 2D-to-3D mapping whose past-to-current 3D alignment, guided by the injected past 2D trace, aids action prediction.
- Implementing a cascaded pipeline in which a predicted future trace conditions the downstream action predictor.
-
2025 Shanghai, China
Atomic Environment Imaging for Efficient Machine-Learning Force Fields
New York University Shanghai Independent research · Advised by Prof. Shengjie Wang - Using 3D visualizations of molecular structure to improve force-field prediction models.
- Built a general pipeline that renders predefined multi-view images of the local environment centered on each atom.
- Added image encoders for models with direct force prediction, and coordinate encoders that reconstruct those images for gradient-based models that take only coordinates as input.
-
2025 Maryland, US
TraceGen: World Modeling in 3D Trace-Space
University of Maryland, College Park Summer research · Advised by Prof. Furong Huang - Addressed the small-data problem in manipulation with a compact 3D trace-space of scene-level trajectories, letting a world model predict future motion geometrically rather than in pixel space and learn from cross-embodiment, cross-environment, and cross-task video.
- Contributed to TraceForge, the data pipeline that converts heterogeneous human and robot video into consistent 3D traces, yielding a 123K-video, 1.8M observation–trace–language triplet pretraining corpus.
- Pretraining on that corpus reaches 80% success across four tasks from five target-robot videos, at 50–600× the inference speed of video-based world models, and 67.5% real-robot success from five uncalibrated handheld human videos.
-
2024–25 Shanghai, China
Efficient Training for Small-Molecule Force Fields
New York University Shanghai Independent research · Advised by Prof. Shengjie Wang (NYU Shanghai) and Prof. Tianyi Zhou (UMD) - Generated large pools of low-precision molecular dynamics samples, then used submodular selection to pick diverse conformations for recomputation at high precision.
- Showed that training GemNet on points chosen by similarity-based submodular functions outperforms both random and equal-timestep sampling.
- Implemented graph VAEs to learn geometry-aware features for submodular selection, and adapted EGNN to predict a coarse force field, yielding force-field-aware selection features.
Selected Projects
-
2024 Cost-Aware Finetuning of Language Models for Chemical Reaction Prediction
Final project, Natural Language Processing with Representation Learning (graduate) - Finetuned pretrained BART-based models for chemical reaction prediction under a compute budget.
- Constructed USPTO-50K_γ, a dataset pairing LLM-generated predictions with experimental data.
- Compared finetuning strategies, including a from-scratch LoRA implementation.
-
2023 Shanghai, China
Popularity Prediction of YouTube Videos
New York University Shanghai Summer research · Advised by Prof. Xianbin Gu - Applied the SlowFast action-recognition model to derive action labels for downstream prediction.
- Trained an MLP over visual features and the first seven days of view counts to predict total views at 30 days.
Honors and Awards
-
2026 - Summa Cum Laude, B.S., NYU Shanghai
-
2023–25 - Recognition Award, NYU Shanghai
-
2022–25 - Dean's Honor List, NYU Shanghai
-
2022–23 - Dean's Undergraduate Research Fund (DURF)
Technical Skills
-
Languages
- Python (proficient)
- C / C++ (familiar)
- LaTeX
-
Frameworks and tools
- PyTorch
- CUDA, OpenMP, MPI
-
Domains
- Robot learning and vision-language-action models
- Video diffusion and world models
- Machine-learned molecular force fields