
CoMA: Compositional Human Motion Generation with Multi-modal Agents
A multi-modal, compositional framework that refines complex human motion generation from textual descriptions, producing more realistic and diverse motion sequences.
About me
MSE Student in Computer Science at Johns Hopkins University
I work at the intersection of computer vision, generative models, and 3D understanding. I’m currently interested in vision-language-action models (VLA) and embodied AI—building intelligent agents that can perceive, reason, and act in the physical world.
01 / Selected work
Research on controllable generation and efficient visual models.

A multi-modal, compositional framework that refines complex human motion generation from textual descriptions, producing more realistic and diverse motion sequences.

An investigation of accelerated sampling for diffusion-based style transfer. UniPC achieved a fivefold speedup for high-quality stylized image generation in the study.
02 / Background
M.S.E. in Computer Science
Graduate study in computer science with interests in visual computing and generative artificial intelligence.
Berkeley Global Access Program
Coursework in computer vision, artificial intelligence, and computer networks. View CS 180 projects
UCInspire Research Program · GPA 4.0/4.0
Worked with Prof. Xiaohui Xie on human motion generation.
B.Eng. in Artificial Intelligence
GPA 92/100 (3.848/4.0) · Major rank 1/63 · Grade rank 3/253
03 / Experience
Applied work spanning recognition, graph learning, detection, and 3D vision.
Graduation project · 2024
A diffusion-based reconstruction pipeline using SMPL-X priors, multi-view synthesis, ViT encoders, and cross-view feature fusion to recover detailed geometry and texture.
Team lead · 2023–2024
Adapted IA-YOLO with YOLOv3 and a channel-attention mechanism for robust detection under challenging weather conditions.
Research project · 2023–2024
Combined spatial and temporal aggregation to capture evolving patterns in financial transaction graphs.
Independent study · 2023
Built an end-to-end facial authentication system with OpenCV, TensorFlow, and Flask, reaching 91% test accuracy.
04 / Recognition
Let’s connect