Tuan-Anh Vu

Tuan-Anh Vu

Postdoctoral Researcher
· University of California, Los Angeles, working with Prof. M. Khalid Jawed
· Ph.D., HKUST 2024, advised by Prof. Sai-Kit Yeung

3D / 4D Scene Understanding Robot Perception Foundation Models Robust Visual AI Autonomous Systems
On the 2026–2027 academic job market

Research Overview

I am a Postdoctoral Researcher at the UCLA SCI Lab, working with Prof. M. Khalid Jawed. I completed my Ph.D. in Computer Science and Engineering at HKUST in 2024, advised by Prof. Sai-Kit Yeung. From 2023–2024 I was a researcher at CFAR, A*STAR, Singapore under the A*STAR Research Attachment Programme, mentored by Prof. Qing Guo and IEEE Fellow Prof. Ivor W. Tsang.

My research develops visual perception systems that are robust in the wild — handling transparency, camouflage, dynamic 3D geometry, and uncontrolled field conditions. I build on foundations in 3D/4D scene understanding, point cloud learning, physics-guided AI, and vision-language models to enable reliable perception for autonomous robots and intelligent systems, including UAVs, ground robots, and robot arms.

My longer-term goal is trustworthy embodied intelligence for real-world autonomy: robots that perceive and reason under distribution shift and visual ambiguity, maintain accurate 3D/4D representations of changing scenes, and act safely for navigation, monitoring and inspection, and targeted intervention.

I have mentored 2 Ph.D., 3 master's, and 11 undergraduate researchers at UCLA and HKUST, and served as a teaching assistant for more than 15 undergraduate and postgraduate course offerings. I co-organize the CV4DC and CV4Animals workshops at ICCV, CVPR, and ACCV, serve as a Guest Editor for IJCV, and sit on the program committees of CVPR, ECCV, ICCV, NeurIPS, and ICLR. I am also the founding CEO of Computer Vision for Developing Countries, a US-registered non-profit.

Research Themes Overview

I build toward trustworthy embodied intelligence in four steps: perceive what ordinary vision misses, model it in 3D and 4D as the scene changes, act on it safely on real platforms, and ground it in language domain experts can actually use.

01 Perceive

See what standard vision misses

Recover targets that defeat models trained on ordinary appearance — camouflaged, transparent, occluded, or visible only in thermal and near-infrared.

Robust Visual Perception5 papersIJCV, WACV
02 Model

Keep a 3D/4D world model correct while the scene moves

Joint reconstruction and motion, Gaussian splatting, neural fields, and registration that hold up when geometry deforms rather than only on static captures.

3D & 4D Scene Understanding9 papersECCV Oral, AAAI, 3DV, BMVC
03 Act

Turn perception into safe autonomy on real platforms

Onboard vision–language–action control, degeneracy-aware state estimation, and safety analysis for UAVs, ground robots, and manipulators in the field.

Robot Perception & Autonomy9 papersAAAI Oral, CoRL, T-RO
04 Ground

Make what a robot sees describable and queryable

Domain foundation models and referring segmentation that turn raw perception into something a marine biologist or agronomist can search, question, and trust.

Vision-Language & Foundation Models9 papersECCV Oral, MMM, WACV

Research Agenda

My vision is trustworthy AI from models to systems to applications: multimodal and spatiotemporal models that stay reliable under open-world shift; scalable, resource-aware systems that expose uncertainty rather than hide it; and decision-centered applications validated through repeatable experiments in the environments where they must operate. AI succeeds in a laboratory when a model predicts accurately on a fixed benchmark — it succeeds in the world only when the whole pipeline keeps working as sensors, environments, and tasks change.

The central question: how can an AI system preserve useful, calibrated behavior when its data, sensors, compute budget, and operational objective all change?

01
Models

Open-world multimodal and spatiotemporal learning

Create data-efficient models that integrate vision, geometry, language, and heterogeneous sensor signals while detecting and communicating uncertainty under distribution shift.

Rather than assuming every modality is always available and reliable, I study modality-aware representation learning: objectives that exploit complementary cues during training but stay stable when a sensor is missing, degraded, or replaced. Robustness is measured not by average accuracy alone, but by calibration under shift, worst-group performance, recovery after sensor degradation, and how much new data or computation adaptation costs.

  • Selective and calibrated prediction — uncertainty estimation, out-of-distribution detection, and abstention criteria tied to the cost of downstream errors
  • Data-efficient adaptation — self-supervised and test-time methods that learn from unlabeled streams without catastrophic drift
  • Spatial-temporal foundation models — representations that jointly encode appearance, geometry, motion, and semantic context across 2D–4D data
Grounded in
TransCues (WACV 2026) · Catch Me If You Can Describe Me (IJCV) · CamoVid60K (IJCV) · Physics-Regularized Latent Representations (preprint) · HiddenObject
02
Systems

Scalable dynamic world models, from cloud to edge

Turn learned representations into systems that update a 3D/4D world state online, operate within latency and energy budgets, and expose actionable reliability signals to users and decision modules.

The contribution here is not a faster inference engine but an observable AI pipeline, in which data provenance, model version, confidence, latency, and task outcome can be traced together. Streaming world models should maintain geometry, motion, semantics, and uncertainty as new observations arrive, across three scales usually studied in isolation.

  • Training scale — distributed experiments and controlled data engines for multimodal pretraining, adaptation, and rare-event generation
  • Representation scale — compact, online-updatable scene models with consistency checks, change detection, and memory management over long deployments
  • Deployment scale — profiling-guided compression, cloud–edge partitioning, runtime monitoring, and fallback behavior under compute or sensor failure
Grounded in
RFNet-4D (ECCV 2022 Oral) · Test-Time Augmentation (3DV 2024) · Cross4D-JEPA · DePT3R · NIR + metadata Gaussian splatting (AAAI 2026) · Degeneracy-aware VIO (preprint)
03
Applications

Decision-centered AI, validated where it has to work

Co-design AI with domain partners so that model confidence and system behavior improve measurable decisions — not benchmark accuracy alone.

I use demanding field applications as stress tests that reveal general AI and systems questions. Agricultural and marine platforms remain my development testbeds because they permit controlled study of occlusion, environmental shift, and rare failures before methods transfer elsewhere. The shared experiment across domains is to propagate a perturbation through the stack and measure how a sensing change alters confidence, latency, decisions, safety, and recovery.

  • Agriculture — crop intelligence under occlusion and shift, longitudinal digital twins and precision mapping, and reliable edge AI for targeted spraying, scouting, biological pest control, and robotic pollination
  • Manufacturing and inspection — multimodal detection of transparent, reflective, occluded, or novel parts, and dynamic reconstruction for robotic inspection and manipulation
  • Transportation and logistics — persistent 3D/4D mapping, asset monitoring, and risk-aware mobile autonomy where latency, coverage, and recovery matter as much as recognition
Grounded in
VLA-Flight (preprint) · AgriChrono · AgriDrone · Robotic pollination · MarineInst (ECCV 2024 Oral) · MarineVRS · Marine Video Kit · NSF CPS award (Senior Personnel) and 2.25M ACCESS credits as PI / Co-PI

Recent News

Aug 2026 Named Senior Personnel on a $1M NSF CPS award — “Physics-Guided Latent Space Models for Detecting Occluded Objects” (NSF #2551220).
Aug 2026 Awarded NSF ACCESS Accelerate (Co-PI, 1.5M credits) and a Thinking Machines Lab Tinker Research Grant (PI).
Aug 2026 Two papers accepted at BMVC 2026 — Off-Manifold Refinement and ForestMamba.
Jul 2026 Selected as an Outstanding Reviewer at ECCV 2026 — 782 of 12,280 reviewers (top 6.4%).
Jul 2026 Paper accepted at ECCV 2026 — MarineEVT: event-centric marine video understanding via visual tool reasoning.
Jun 2026 Co-organizing the CV4Animals Workshop at CVPR 2026 and serving as Guest Editor of the companion IJCV special collection.
May 2026 Awarded NSF ACCESS Discover (PI, 750K credits) and a NAIRR Pilot Start-Up allocation (PI, 2,000 GPU hours).
Mar 2026 Paper accepted at IJCV — CamoVid60K: large-scale video dataset for moving camouflaged animals.
Jan 2026 Paper accepted at IJCV — open-vocabulary camouflaged instance segmentation with diffusion.
Jan 2026 Invited talks at Sungkyunkwan University and Fulbright University Vietnam.
Jan 2026 Serving as Session Co-Chair at AAAI 2026.
Jan 2026 Appointed Executive Officer (Secretary) of IEEE Coastal Los Angeles Section, Region 6.
Dec 2025 Appointed Area Chair at ICASSP 2026.
Nov 2025 Two papers accepted at AAAI 2026 (one Oral, 4.3% rate); one at WACV 2026.
Oct 2025 Co-organized the CV4DC Workshop at ICCV 2025; the series continues at ACCV 2026.
Mar 2025 Founded and serve as CEO of Computer Vision for Developing Countries, a US-registered non-profit.
Feb 2025 Joined UCLA SCI Lab as a Postdoctoral Researcher, working with Prof. M. Khalid Jawed.
Aug 2024 Two papers at ECCV 2024 — MarineInst (Oral, 2.3%) and StyleCity3D. Participated in ECCV Doctoral Consortium, mentored by Prof. Leonidas Guibas (IEEE/ACM Fellow).

Publications * equal contribution  ·  † senior/lead  ·  # corresponding author

2026
MarineEVT thumbnail
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
Tuan-An To, Wong Yuk Kwan, Tuan-Anh Vu, Ziqiang Zheng, Sai-Kit Yeung
ECCV 2026 New
Recasts marine video understanding as event-centric reasoning, letting the model invoke visual tools to work out what happened rather than captioning frames in isolation.
2026
BMVC
2026
Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model
Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
BMVC 2026 New
Uses a frozen world model as an external critic to pull video generations back toward physically plausible motion, without retraining the generator.
2026
BMVC
2026
ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation
Trung Thanh Nguyen, Tuan-Anh Vu, Duc Viet Le, Yasutomo Kawanishi, Takahiro Komamizu, Ichiro Ide, Teja Kattenborn
BMVC 2026 New
Brings sparse state-space sequence modelling to forest-scale point clouds, using geometry-guided queries to segment structure where dense attention is impractical.
2026
Catch Me thumbnail
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
IJCV 2026 Journal New
Shows diffusion features retain the boundary evidence camouflage suppresses, enabling instance segmentation of camouflaged objects from open-vocabulary text instead of a fixed class list.
Earlier version at the ECCV 2024 CV4Ecology Workshop.
2026
CamoVid60K thumbnail
CamoVid60K: A Large-Scale Video Dataset for Moving Camouflaged Animals Understanding
Tuan-Anh Vu#, Ziqiang Zheng, Chengyang Song, Qing Guo, Ivor Tsang, Sai-Kit Yeung
IJCV 2026 Journal New
Supplies a large-scale video benchmark for camouflaged animals in motion, where the revealing signal is temporal rather than appearance-based.
Also at the CVPR 2025 CV4Animals Workshop — Oral, with Travel Award.
2026
TransCues thumbnail
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
WACV 2026 New
Segments transparent objects by explicitly modelling the two cues that survive transparency — boundary and reflection — inside a pyramid vision transformer.
Earlier version at the ECCV 2024 Transparent & Reflective Objects In the Wild Workshop.
2026
NIR Gaussian Splatting thumbnail
Reconstruction Using the Invisible: Intuition from NIR and Metadata for Enhanced 3D Gaussian Splatting
Gyusam Chang, Tuan-Anh Vu†, Vivek Alumootil, Harris Song, Deanna Pham, Sangpil Kim#, M. Khalid Jawed#
AAAI 2026 New
Shows near-infrared imagery and capture metadata recover geometry that RGB-only Gaussian splatting misses on low-texture and poorly lit surfaces.
Earlier version at the CVPR 2025 2nd Workshop on Neural Fields Beyond Conventional Cameras.
2026
Multi-Agent Behaviors thumbnail
Exploiting Geometric Structures for Modeling Multi-Agent Behaviors: A New Thinking
Bohao Qu, Xiaofeng Cao, Bing Li, Menglin Zhang, Tuan-Anh Vu, Di Lin, Qing Guo
AAAI 2026 Oral · 4.3% New
Recasts multi-agent behaviour modelling around the geometry of agent interactions rather than treating each agent's trajectory independently.
2026
Preprint
EgoPoseProbe: Diagnosing Scene Grounding and Feature Adaptation in Egocentric 3D Body Pose Estimation
Hoang M. Truong, Hai Nguyen-Truong, Tuan-Anh Vu#
Preprint
A diagnostic that separates how much egocentric 3D pose estimators depend on scene grounding versus feature adaptation, isolating where current methods actually fail.
2026
Preprint
Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning
Trung Thanh Nguyen, Hai Nguyen-Truong, Tu Vo, Hoang M. Truong, Tuan-Anh Vu#
Preprint
Learns 4D point-cloud representations by distilling dense cross-modal correspondence, removing the need for annotated motion supervision.
2026
Preprint
VLA-Flight: Lightweight Onboard Vision–Language–Action with Direct IMU Fusion for UAV Control
Tuan-Anh Vu*, Yashas Shashidhara*, M. Khalid Jawed
Preprint
Fuses IMU directly into a lightweight vision-language-action policy so a UAV can close the control loop onboard instead of offloading perception.
2026
Preprint
Do Simplicial, Hypergraph, and Spectral Biases Compose in Molecular GNNs? A Controlled Negative Result
Huyen-Thu Vu, Thanh-Hoang Nguyen-Vo, Tuan-Anh Vu#, Binh P. Nguyen#
Preprint
A controlled negative result: simplicial, hypergraph and spectral inductive biases in molecular GNNs do not compose, and stacking them buys no additive gain.
2026
Preprint
Information-Theoretic Point Selection and Observability-Constrained EKF for Degeneracy-Aware RGB-D Visual-Inertial Odometry
Tuan-Anh Vu, Xiaoyang Zhao, M. Khalid Jawed
Preprint
Selects points by information gain and constrains the EKF along unobservable directions, keeping RGB-D visual-inertial odometry stable where scene geometry degenerates.
2026
Hidden Object Thermal thumbnail
Physics-Regularized Latent Representations for Hidden Object Inference from Thermal Dynamics
Tuan-Anh Vu, M. Khalid Jawed
Preprint
Infers objects hidden from direct view by regularising the latent representation with the physics of how heat propagates through the occluding surface.
2026
Preprint
Consistency and Stability of LLM Reasoning Loops: A Latent Dynamical Systems Perspective
Tuan-Anh Vu*, Angela Yang*, Mohammad Sadoughi, Amit Kachroo, M. Khalid Jawed
Preprint
Treats iterative LLM reasoning as a dynamical system, so the stability of a reasoning loop can be characterised rather than judged only from its final answer.
2026
Preprint
EvoReg: Versatile and Robust Point Cloud Registration via Multi-Stage Alignment
Alan Nadelsticher Ruvalcaba, Joohyun Lee, Wenying Wu, Benjamin Ostrower, Junling Zhuang, Amir J. Bidhendi, Janie Huang, Luke DeVivo, Tuan-Anh Vu†, Supratik Mukhopadhyay
Preprint
Multi-stage alignment that keeps point-cloud registration robust across the range of overlap and initialisation regimes where single-stage methods break down.
2026
Preprint
Towards Human-Centered Safety in Vision-Language-Action Policies for Robotics: A Survey and Outlook
Nhat Chung*, Toan Nguyen*, Yun Xing, Tuan-Anh Vu, Bohao Qu, Yedi Zhang, Yue Cao, Jie Zhang, Yue Wang, Cheston Tan, Harold Soh, Qing Guo#, Jin Song Dong#, David Hsu#, Daniel Seita#, Ivor Tsang#
Preprint
Maps the safety landscape of vision-language-action policies around the human, setting out where today's robot policies leave people exposed.
2026
AquaVision thumbnail
AquaVision: An Edge-Deployed 360° Underwater Panoramic Monitoring System
Tan-Sang Ha, Tuan-Anh Vu#, Sai-Kit Yeung
Preprint
An edge-deployed 360° rig that makes continuous underwater panoramic monitoring practical without a surface tether or offboard compute.
2026
AgriDrone thumbnail
AgriDrone: A Multi-Sensor Dataset and Lightweight Vision Backbone for Perception in Greenhouse Drone Navigation
Tuan-Anh Vu, Evelyn Zhu, Xiaoyang Zhao, Angela Yang, Akshat Pandya, M. Khalid Jawed
Preprint
Pairs a multi-sensor greenhouse dataset with a vision backbone sized for the compute a drone can actually carry indoors.
2026
HiVLA thumbnail
Your Vision-Language-Action Model Already Has Attention Heads For Path Deviation Detection
Jaehwan Jeong, Evelyn Zhu, Jinying Lin, Emmanuel Jaimes, Tuan-Anh Vu, Jungseock Joo, Sangpil Kim#, M. Khalid Jawed#
Preprint
Finds that path-deviation signal is already present in a VLA policy's own attention heads, so drift can be detected by reading the policy rather than training a separate monitor.
2025
DePT3R thumbnail
DePT3R: Joint Dense Point Tracking and 3D Reconstruction of Dynamic Scenes in a Single Forward Pass
Vivek Alumootil, Tuan-Anh Vu#
Preprint
Recovers dense point tracks and 3D structure of a dynamic scene in a single forward pass, dropping the usual per-scene optimisation loop.
2025
HiddenObject thumbnail
HiddenObject: Modality Agnostic Fusion for Multimodal Hidden Object Detection
Harris Song*, Tuan-Anh Vu*†, Sanjith Menon, Sriram Narasimhan, M. Khalid Jawed
Preprint
Fuses arbitrary sensing modalities for hidden-object detection without assuming in advance which modality will be available at test time.
2025
Robotic Pollination thumbnail
Vision-Guided Targeted Grasping and Vibration for Robotic Pollination in Controlled Environments
Jaehwan Jeong*, Tuan-Anh Vu*, Radha Lahoti, Vivek Alumootil, Sangpil Kim#, M. Khalid Jawed#
Preprint
Closes the loop from visual flower detection to targeted grasp-and-vibrate, demonstrating robotic pollination in a controlled environment.
2025
AgriChrono thumbnail
AgriChrono: A Multi-modal Dataset Capturing Crop Growth and Lighting Variability with a Field Robot
Jaehwan Jeong, Tuan-Anh Vu†, Mohammad Jony, Shahab Ahmad, Md Mukhlesur Rahman, Sangpil Kim#, M. Khalid Jawed#
Preprint
Captures crop growth and lighting variation over time with a field robot, giving a dataset where the scene itself changes rather than only the viewpoint.
2025
Vision-Aware Text thumbnail
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
WACV 2025
Injects visual context into the text branch of referring segmentation, moving the model from object-level matching to context-level understanding.
2024
Time-Varying Point Clouds thumbnail
Reconstruction of Time-Varying Point Clouds via Bidirectional Weakly-Supervised Learning of Spatio-Temporal Cross Features
Preprint
Learns spatio-temporal cross features bidirectionally, so time-varying point clouds can be reconstructed under weak supervision.
2024
TTA 3D thumbnail
Test-Time Augmentation for 3D Point Cloud Classification and Segmentation
Tuan-Anh Vu*, Srinjay Sarkar*, Zhiyuan Zhang#, Binh-Son Hua, Sai-Kit Yeung
3DV 2024
Shows test-time augmentation carries over to 3D point clouds, improving classification and segmentation with no retraining.
2024
MarineInst thumbnail
MarineInst: A Foundation Model for Marine Image Analysis with Instance Visual Description
ECCV 2024 Oral · 2.3%
A marine foundation model that segments instances and describes them, so its output is usable by biologists rather than only by other models.
2024
StyleCity thumbnail
StyleCity: Large-Scale 3D Urban Scenes Stylization
ECCV 2024 US Patent (filed)
Stylises city-scale 3D scenes consistently across views and geometry, where per-image stylisation would flicker.
Also Oral at the ECCV 2024 CV4Metaverse Workshop.
2024
GPT-4V Marine thumbnail
Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study
Preprint
An early systematic probe of where GPT-4V succeeds and fails on marine imagery, establishing the gap that domain-specific models needed to close.
2024
Transferable Attacks thumbnail
Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography
Preprint
Shows typographic attacks transfer across vision-LLMs in driving scenes, exposing a practical attack surface for deployed autonomy.
2023
MarineVRS thumbnail
MarineVRS: Marine Video Retrieval System with Explainability via Semantic Understanding
OCEANS 2023 Oral
Retrieval over marine video with an explainable path from query to segment, so a result can be justified to a domain expert.
2023
Marine Video Kit thumbnail
Marine Video Kit: A New Marine Video Dataset for Content-based Analysis and Retrieval
MMM 2023 Oral
A marine video dataset built for content-based retrieval, addressing a domain shift that general-purpose video corpora do not cover.
2023
MarineGPT thumbnail
MarineGPT: Unlocking Secrets of Ocean to the Public
Ziqiang Zheng, Jipeng Zhang, Tuan-Anh Vu, Shizhe Diao, Yue Him Wong Tim, Sai-Kit Yeung
Preprint
A domain-tuned vision-language model that answers marine questions with the specificity general assistants lack.
2022
RFNet-4D thumbnail
RFNet-4D: Joint Object Reconstruction and Flow Estimation from 4D Point Clouds
ECCV 2022 Oral · 2.7%
Learns reconstruction and motion flow jointly from raw 4D point clouds, instead of reconstructing each frame and matching afterwards.
2022
Time-of-Day Style Transfer thumbnail
Time-of-Day Neural Style Transfer for Architectural Photographs
ICCP 2022 Oral
Transfers time-of-day appearance to architectural photographs while preserving structure, separating illumination from geometry.

Thesis & Dissertation

2024
PhD Dissertation thumbnail
Robust Scene Understanding in Challenging Scenarios
Tuan-Anh Vu
Ph.D. Dissertation · HKUST · 2024
Advised by Prof. Sai-Kit Yeung (HKUST).
Doctoral Consortiums:
2019
BSc Thesis thumbnail
Extend Traffic Signs Detection and Recognition Algorithm in Nighttime in Viet Nam
Tuan-Anh Vu
B.Sc. Thesis · HCMIU-VNU · 2019

Full list on Google Scholar and ORCID.

Grants & Awarded Resources

Named Senior Personnel on a $1M NSF CPS award, having generated the preliminary data and contributed to the proposal, and PI or Co-PI on more than $2M in high-performance computing and cloud allocations.

Aug 2026
Physics-Guided Latent Space Models for Detecting Occluded Objects
NSF CPS-CIR, CPS-FR  ·  Senior Personnel
PI: Prof. M. Khalid Jawed
$1,000,000total award
Aug 2026
NSF ACCESS Accelerate
High-performance computing and cloud services  ·  Co-PI
1,500,000credits
Aug 2026
Thinking Machines Lab — Tinker Research Grant
Cloud computing  ·  PI
$5,000cloud credits
May 2026
NSF ACCESS Discover
High-performance computing and cloud services  ·  PI
750,000credits
May 2026
NAIRR Pilot — Start-Up Allocation
National AI Research Resource  ·  PI
2,000GPU hours

Teaching & Student Supervision

Courses Taught / TA

Teaching assistant for more than 15 undergraduate and postgraduate course offerings across computer science, programming, multimedia computing, digital design, and artificial intelligence.

Introduction to Computer Science HKUST · COMP 1021
Programming with C++ HKUST · COMP 2011
Multimedia Computing HKUST · COMP 4431
Introduction to 3D Design HKUST · ISDN 2300
Physical Prototyping HKUST · ISDN 2400
Advanced Digital Design HKUST · ISDN 5300 / CSIT 6000L / CSIT 5940
Digital Image Processing HCMIU-VNU
Theoretical Models for Computing HCMIU-VNU
Introduction to Artificial Intelligence HCMIU-VNU
Principles of Programming Languages HCMIU-VNU

Current Mentees

Graduate Students
AI
Gyusam Chang
Visiting Ph.D. · Artificial Intelligence · Korea University
AI
Jaehwan Jeong
Visiting Ph.D. · Artificial Intelligence · Korea University
ME
Xiaoyang "Tony" Zhao
M.S. · Mechanical Engineering · UCLA
CS
Alan Nadelsticher Ruvalcaba
M.S. · Computer Science · Georgia Tech
Undergraduate Students
CS
Harris Song
CS
Vivek Alumootil
CS
Angela Yang
CS
Aaditya Raj
CS
Kevin Yao
M&P
Menghui "Evelyn" Zhu
DT
Russell Luo
ME
Deanna Pham
ME
Sanjith Menon
ME
Mehmet Arif Bacaksizlar
CE
Darren Chin
CS
Yashas Shashidhara

Professional Service

Leadership & Organization

CEO, Computer Vision for Developing Countries (non-profit, 2025–)
IEEE Executive Officer (Secretary), Coastal Los Angeles Section, Region 6 (2026–)
Guest EditorIJCV special collection on CV4Animals (2026–)
Session Co-Chair, AAAI 2026
Area Chair, ICASSP 2026
Workshop Co-organizerCV4Animals @ CVPR 2026
Workshop Co-organizerCV4DC @ ACCV 2024, ICCV 2025, ACCV 2026

Program Committee / Reviewer

Vision: CVPR '22–26, ICCV '23–25, ECCV '22–26, WACV '22–27, BMVC '26, ACCV '22–24, ISVC '25
Machine learning: NeurIPS '24–26, ICML '25, ICLR '25–26, AAAI '25–27, IJCAI '25–26, AISTATS '25–26, ECAI '25, IJCNN '27
Robotics & multimedia: ICRA '25–26, CoRL '26, ACM MM '25–26, ICMI '25–26, ICME '23
Journals: IJCV, ACM TOG, IEEE TIP, IEEE TMM, IEEE TBD, IEEE RA-L, IEEE TAI, Springer AI Review, Springer Visual Intelligence, Springer Scientific Reports, Elsevier ESWA, Neurocomputing, CAD, EAAI, IET Image Processing, PeerJ CS

Selected Honors & Awards

2026 Outstanding Reviewer, ECCV 2026 · 782 of 12,280 reviewers (top 6.4%)
2025 Doctoral Consortium with Travel Grant, WACV 2025 · Mentored by Prof. R. Souvenir & Prof. S. X. Huang
2024 Doctoral Consortium, ECCV 2024 · Mentored by Prof. Leonidas Guibas (IEEE/ACM Fellow)
2024 Doctoral Consortium, IEEE CAI 2024 · Mentored by Prof. Ivor W. Tsang (IEEE Fellow)
2024 2nd Place, USV-based Obstacle Segmentation Challenge · WACV 2024 Maritime CV Workshop
2023 2nd Place, USV-based Obstacle Segmentation Challenge · WACV 2023 Maritime CV Workshop
2022 UGC Research Travel Grant, HKUST (ECCV 2022 & ECCV 2024) · Travel Grant, EPFL CIS Edge AI Summer School
2020 Best Poster Award, Machine Learning Summer School Indonesia (MLSS-Indo)
2019 Postgraduate Scholarship, HKUST (2019–2024) · Scholarship for SENG Summer Camp for Elite Students, HKUST