cv

Education, research experience, selected projects, and awards. Also available as a PDF.

General Information

Name 黄泓嘉 / Hongjia (Alex) Huang
Affiliation Robotics Institute, Carnegie Mellon University
Email hh3043@nyu.edu
Research Interests Robotics, computer vision, and physics-grounded video generation
Profiles

Education

  • 2026–

    Pittsburgh, US

    M.S. in Robotics (MSR)
    Carnegie Mellon University
    Robotics Institute
  • 2022–26

    Shanghai, China

    B.S. in Computer Science and Mathematics (double major)
    New York University Shanghai
    Summa Cum Laude
    • Overall GPA: 3.93 / 4.0
    • Selected coursework
      • Graduate: Natural Language Processing with Representation Learning, Inference & Representation, Reinforcement Learning
      • Computer Science: Machine Learning, Parallel Computing, Algorithms, Operating Systems, Data Structures
      • Mathematics (Honors): Ordinary Differential Equations, Analysis, Numerical Analysis, Linear Algebra
      • Mathematics: Partial Differential Equations, Linear and Nonlinear Optimization, Probability & Statistics

Research Experience

  • 2024–

    Shanghai, China

    Auxiliary Streams for Video-Diffusion World-Action Models
    New York University Shanghai
    Independent research · Advised by Prof. Shengjie Wang (NYU Shanghai) and Prof. Tianyi Zhou (UMD)
    • Investigating whether training-time auxiliary structure can improve the policy of a frozen Wan2.2-5B video-diffusion world-action model, particularly its spatial out-of-distribution generalization.
    • Added inference-free alignment teachers (VGGT, CoTracker, DINO/REPA) and co-denoised streams routed by a Mixture-of-LoRA-Experts, passing 3D structure from 2D video to the action head at test time.
    • Earlier in the project, adapted pretrained video diffusion models with lightweight modifications so that generated video obeys physical laws, conditioning on physical constraints through a modified ControlNeXt and generating supervised training data with 3D physics engines (Genesis).
  • 2025–

    Maryland, US

    Physics-Informed Vision-Language-Action Models
    University of Maryland, College Park
    Independent research · Advised by Prof. Furong Huang and Prof. Tianyi Zhou
    • Integrating trajectory planning into the action-prediction objective of VLA models such as GR00T N1.5.
    • Using a 4D encoder that fuses multiple views across multiple timesteps to propagate spatial and temporal information jointly, aiding action prediction via future latent alignment.
    • Designing an implicit 2D-to-3D mapping whose past-to-current 3D alignment, guided by the injected past 2D trace, aids action prediction.
    • Implementing a cascaded pipeline in which a predicted future trace conditions the downstream action predictor.
  • 2025

    Shanghai, China

    Atomic Environment Imaging for Efficient Machine-Learning Force Fields
    New York University Shanghai
    Independent research · Advised by Prof. Shengjie Wang
    • Using 3D visualizations of molecular structure to improve force-field prediction models.
    • Built a general pipeline that renders predefined multi-view images of the local environment centered on each atom.
    • Added image encoders for models with direct force prediction, and coordinate encoders that reconstruct those images for gradient-based models that take only coordinates as input.
  • 2025

    Maryland, US

    TraceGen: World Modeling in 3D Trace-Space
    University of Maryland, College Park
    Summer research · Advised by Prof. Furong Huang
    • Addressed the small-data problem in manipulation with a compact 3D trace-space of scene-level trajectories, letting a world model predict future motion geometrically rather than in pixel space and learn from cross-embodiment, cross-environment, and cross-task video.
    • Contributed to TraceForge, the data pipeline that converts heterogeneous human and robot video into consistent 3D traces, yielding a 123K-video, 1.8M observation–trace–language triplet pretraining corpus.
    • Pretraining on that corpus reaches 80% success across four tasks from five target-robot videos, at 50–600× the inference speed of video-based world models, and 67.5% real-robot success from five uncalibrated handheld human videos.
  • 2024–25

    Shanghai, China

    Efficient Training for Small-Molecule Force Fields
    New York University Shanghai
    Independent research · Advised by Prof. Shengjie Wang (NYU Shanghai) and Prof. Tianyi Zhou (UMD)
    • Generated large pools of low-precision molecular dynamics samples, then used submodular selection to pick diverse conformations for recomputation at high precision.
    • Showed that training GemNet on points chosen by similarity-based submodular functions outperforms both random and equal-timestep sampling.
    • Implemented graph VAEs to learn geometry-aware features for submodular selection, and adapted EGNN to predict a coarse force field, yielding force-field-aware selection features.

Selected Projects

  • 2024
    Cost-Aware Finetuning of Language Models for Chemical Reaction Prediction
    Final project, Natural Language Processing with Representation Learning (graduate)
    • Finetuned pretrained BART-based models for chemical reaction prediction under a compute budget.
    • Constructed USPTO-50K_γ, a dataset pairing LLM-generated predictions with experimental data.
    • Compared finetuning strategies, including a from-scratch LoRA implementation.
  • 2023

    Shanghai, China

    Popularity Prediction of YouTube Videos
    New York University Shanghai
    Summer research · Advised by Prof. Xianbin Gu
    • Applied the SlowFast action-recognition model to derive action labels for downstream prediction.
    • Trained an MLP over visual features and the first seven days of view counts to predict total views at 30 days.

Honors and Awards

  • 2026
    • Summa Cum Laude, B.S., NYU Shanghai
  • 2023–25
    • Recognition Award, NYU Shanghai
  • 2022–25
    • Dean's Honor List, NYU Shanghai
  • 2022–23
    • Dean's Undergraduate Research Fund (DURF)

Technical Skills

  • Languages
    • Python (proficient)
    • C / C++ (familiar)
    • LaTeX
  • Frameworks and tools
    • PyTorch
    • CUDA, OpenMP, MPI
  • Domains
    • Robot learning and vision-language-action models
    • Video diffusion and world models
    • Machine-learned molecular force fields