Zeyang Zhang 「张泽阳」

| CV | Email | Google Scholar |
| LinkedIn |

I am an AI Researcher in the Multimodal Speech Processing (MSP) Laboratory and the Accounting AI Research Lab.

I received my M.S. in Artificial Intelligence Engineering from Carnegie Mellon University, where I was advised by Prof. Carlos Busso at the Language Technologies Institute. My master's thesis is Cross-Modal Learning Through Hierarchical Quantized Embeddings.

Before that, I received my B.S. in Computer Science from the University of Illinois at Urbana-Champaign.

Goal: Build multimodal machine learning systems that jointly reason over text, video, audio, IMU, gaze, and other modalities, learning semantic and affective representations that deepen our understanding of human interaction and turning that understanding into the next generation of multimodal agents.

Research Interest: The intersection of representation learning, multimodal machine learning, signal processing, human interaction, AI agents, graph neural networks, and multimodal LLMs.

Email: zeyangz [AT] andrew.cmu.edu


  News
  • [08/2026] Our paper Hierarchical Quantized Cross-modal Masked Autoencoders for Audio-Visual Emotion Recognition was selected for an oral presentation at the 28th ACM International Conference on Multimodal Interaction (ICMI) 2026!
  • [07/2026] Our paper Hierarchical Quantized Cross-modal Masked Autoencoders for Audio-Visual Emotion Recognition has been accepted as a long paper at the 28th ACM International Conference on Multimodal Interaction (ICMI) 2026!
  • [07/2026] Our paper Speaker Diarization in Static and Egocentric Videos via Speech-Face Alignment has been accepted as a long paper at the 28th ACM International Conference on Multimodal Interaction (ICMI) 2026!
  • [05/2026] Graduated from Carnegie Mellon University with an M.S. in Artificial Intelligence Engineering (GPA: 3.94/4.00).
  • [04/2026] I successfully defended my master's thesis! Many thanks to my committee members, Prof. Carlos Busso and Prof. Max Simchowitz, for their support.
  • [01/2026] Our paper Contrastive Gated Fusion for Multilingual Speaker Verification, submitted to the ICASSP Grand Challenge GC-6: Face-Voice Association in Multilingual Environments (FAME), has been accepted for inclusion in the ICASSP Workshop Proceedings!
  • [05/2024] Graduated from the University of Illinois at Urbana-Champaign with a B.S. in Computer Science with Highest Honors (GPA: 3.94/4.00).

  Publications
sym

Hierarchical Quantized Cross-modal Masked Autoencoders for Audio-Visual Emotion Recognition
Zeyang Zhang and Carlos Busso. 2026.
INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI '26)
Oral Presentation

pdf | doi
sym

Speaker Diarization in Static and Egocentric Videos via Speech-Face Alignment
Zeyang Zhang, Xavier Yin, Zhaobo Zheng, Kumar Akash, Teruhisa Misu, and Carlos Busso. 2026.
INTERNATIONAL CONFERENCE ON MULTIMODAL INTERACTION (ICMI '26)
Poster Presentation

doi
sym

Contrastive Gated Fusion for Multilingual Speaker Verification
Zeyang Zhang, Katsuhiko Naito, and Hajer Dahmani. 2026.
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2026)

pdf | doi

  Projects

I am currently working on three projects:

  • Collective State Sensing with Multimodal Signal Processing and Fusion
    • We fuse face, speech, IMU, and gaze signals from each participant working in a group to infer their individual performance level.
    • We also plan to leverage these multimodal signals to predict engagement, leadership, and overall team performance during collaborative work.
  • Audio-Visual Distillation for General-Purpose IMU Encoder Pretraining
    • We unify IMU-only JEPA pretraining, audio-visual knowledge distillation, and Mantis-8M into a single framework for pretraining an IMU encoder that is both robust and generalizable.
    • Our model outperforms all baseline models on downstream human activity recognition tasks.
  • Predicting Future Earnings Changes Using Heterogeneous Graph Representations
    • We design a novel graph representation that captures the cash flow of each firm in every quarter.
    • We develop a temporal graph neural network that learns embeddings from these novel graph representations to predict future earnings changes.
    • Our temporal GNN model with unique graph structures outperforms all existing baselines on earnings change prediction.

Detailed descriptions, results, and papers are coming soon.


  Reviewer Service
International Conference on Multimodal Interaction (ICMI) 2026


  Life is Good
sym

My dog Kevin just turned 12, and he still greets me like a puppy every single day!

I play table tennis at CMU whenever I get the chance, come find me for a match! I also just picked up tennis two months ago, and I am already completely hooked!

My friends over the years, the people who have made all of this so much fun! I love traveling, hanging out with friends, and hunting down great restaurants!