About
Research Overview
I am a Postdoctoral Researcher at the UCLA SCI Lab, working with Prof. M. Khalid Jawed. I completed my Ph.D. in Computer Science and Engineering at HKUST in 2024, advised by Prof. Sai-Kit Yeung. From 2023–2024 I was a researcher at CFAR, A*STAR, Singapore under the A*STAR Research Attachment Programme, mentored by Prof. Qing Guo and IEEE Fellow Prof. Ivor W. Tsang.
My research develops visual perception systems that are robust in the wild — handling transparency, camouflage, dynamic 3D geometry, and uncontrolled field conditions. I build on foundations in 3D/4D scene understanding, point cloud learning, physics-guided AI, and vision-language models to enable reliable perception for autonomous robots and intelligent systems, including UAVs, ground robots, and robot arms.
My longer-term goal is trustworthy embodied intelligence for real-world autonomy: robots that perceive and reason under distribution shift and visual ambiguity, maintain accurate 3D/4D representations of changing scenes, and act safely for navigation, monitoring and inspection, and targeted intervention.
I have mentored 2 Ph.D., 3 master's, and 11 undergraduate researchers at UCLA and HKUST, and served as a teaching assistant for more than 15 undergraduate and postgraduate course offerings. I co-organize the CV4DC and CV4Animals workshops at ICCV, CVPR, and ACCV, serve as a Guest Editor for IJCV, and sit on the program committees of CVPR, ECCV, ICCV, NeurIPS, and ICLR. I am also the founding CEO of Computer Vision for Developing Countries, a US-registered non-profit.
I build toward trustworthy embodied intelligence in four steps: perceive what ordinary vision misses, model it in 3D and 4D as the scene changes, act on it safely on real platforms, and ground it in language domain experts can actually use.
See what standard vision misses
Recover targets that defeat models trained on ordinary appearance — camouflaged, transparent, occluded, or visible only in thermal and near-infrared.
Keep a 3D/4D world model correct while the scene moves
Joint reconstruction and motion, Gaussian splatting, neural fields, and registration that hold up when geometry deforms rather than only on static captures.
Turn perception into safe autonomy on real platforms
Onboard vision–language–action control, degeneracy-aware state estimation, and safety analysis for UAVs, ground robots, and manipulators in the field.
Make what a robot sees describable and queryable
Domain foundation models and referring segmentation that turn raw perception into something a marine biologist or agronomist can search, question, and trust.
Looking Ahead
Research Agenda
My vision is trustworthy AI from models to systems to applications: multimodal and spatiotemporal models that stay reliable under open-world shift; scalable, resource-aware systems that expose uncertainty rather than hide it; and decision-centered applications validated through repeatable experiments in the environments where they must operate. AI succeeds in a laboratory when a model predicts accurately on a fixed benchmark — it succeeds in the world only when the whole pipeline keeps working as sensors, environments, and tasks change.
The central question: how can an AI system preserve useful, calibrated behavior when its data, sensors, compute budget, and operational objective all change?
Open-world multimodal and spatiotemporal learning
Create data-efficient models that integrate vision, geometry, language, and heterogeneous sensor signals while detecting and communicating uncertainty under distribution shift.
Rather than assuming every modality is always available and reliable, I study modality-aware representation learning: objectives that exploit complementary cues during training but stay stable when a sensor is missing, degraded, or replaced. Robustness is measured not by average accuracy alone, but by calibration under shift, worst-group performance, recovery after sensor degradation, and how much new data or computation adaptation costs.
- Selective and calibrated prediction — uncertainty estimation, out-of-distribution detection, and abstention criteria tied to the cost of downstream errors
- Data-efficient adaptation — self-supervised and test-time methods that learn from unlabeled streams without catastrophic drift
- Spatial-temporal foundation models — representations that jointly encode appearance, geometry, motion, and semantic context across 2D–4D data
TransCues (WACV 2026) · Catch Me If You Can Describe Me (IJCV) · CamoVid60K (IJCV) · Physics-Regularized Latent Representations (preprint) · HiddenObject
Scalable dynamic world models, from cloud to edge
Turn learned representations into systems that update a 3D/4D world state online, operate within latency and energy budgets, and expose actionable reliability signals to users and decision modules.
The contribution here is not a faster inference engine but an observable AI pipeline, in which data provenance, model version, confidence, latency, and task outcome can be traced together. Streaming world models should maintain geometry, motion, semantics, and uncertainty as new observations arrive, across three scales usually studied in isolation.
- Training scale — distributed experiments and controlled data engines for multimodal pretraining, adaptation, and rare-event generation
- Representation scale — compact, online-updatable scene models with consistency checks, change detection, and memory management over long deployments
- Deployment scale — profiling-guided compression, cloud–edge partitioning, runtime monitoring, and fallback behavior under compute or sensor failure
RFNet-4D (ECCV 2022 Oral) · Test-Time Augmentation (3DV 2024) · Cross4D-JEPA · DePT3R · NIR + metadata Gaussian splatting (AAAI 2026) · Degeneracy-aware VIO (preprint)
Decision-centered AI, validated where it has to work
Co-design AI with domain partners so that model confidence and system behavior improve measurable decisions — not benchmark accuracy alone.
I use demanding field applications as stress tests that reveal general AI and systems questions. Agricultural and marine platforms remain my development testbeds because they permit controlled study of occlusion, environmental shift, and rare failures before methods transfer elsewhere. The shared experiment across domains is to propagate a perturbation through the stack and measure how a sensing change alters confidence, latency, decisions, safety, and recovery.
- Agriculture — crop intelligence under occlusion and shift, longitudinal digital twins and precision mapping, and reliable edge AI for targeted spraying, scouting, biological pest control, and robotic pollination
- Manufacturing and inspection — multimodal detection of transparent, reflective, occluded, or novel parts, and dynamic reconstruction for robotic inspection and manipulation
- Transportation and logistics — persistent 3D/4D mapping, asset monitoring, and risk-aware mobile autonomy where latency, coverage, and recovery matter as much as recognition
VLA-Flight (preprint) · AgriChrono · AgriDrone · Robotic pollination · MarineInst (ECCV 2024 Oral) · MarineVRS · Marine Video Kit · NSF CPS award (Senior Personnel) and 2.25M ACCESS credits as PI / Co-PI
Updates
Recent News
Research
Publications * equal contribution · † senior/lead · # corresponding author

2026
2026























Thesis & Dissertation

Doctoral Consortiums:
- WACV 2025 (Travel Award) — mentored by Prof. Richard Souvenir and Prof. Sharon X. Huang
- ECCV 2024 — mentored by Prof. Leonidas Guibas (IEEE & ACM Fellow, Stanford)
- IEEE CAI 2024 — mentored by Prof. Ivor W. Tsang (IEEE Fellow)
Full list on Google Scholar and ORCID.
Funding
Grants & Awarded Resources
Named Senior Personnel on a $1M NSF CPS award, having generated the preliminary data and contributed to the proposal, and PI or Co-PI on more than $2M in high-performance computing and cloud allocations.
Education & Mentoring
Teaching & Student Supervision
Courses Taught / TA
Teaching assistant for more than 15 undergraduate and postgraduate course offerings across computer science, programming, multimedia computing, digital design, and artificial intelligence.
Current Mentees
Community
Professional Service
Leadership & Organization
Program Committee / Reviewer
Selected Honors & Awards
