Research
We build world models—generative, predictive models of how things look, move, and evolve—and use them to solve inference problems in the physical and living world. One model serves two directions: run forward, it simulates and generates; run in reverse, it infers what produced a partial, noisy measurement.
Generative and World Models
We learn priors over images, video, 3D, and language: latent diffusion and flow models; unified multimodal conditioning; controllable and contextual generation; video and 3D synthesis.
Representative work
- UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation — arXiv 2026. Unified visual–text conditioning; removes text encoders from diffusion, with 59% faster inference at state-of-the-art quality.
- Pose-Aware Self-Supervised Learning with Viewpoint Trajectory Regularization — ECCV 2024 (Oral). Continuous, geometry-aware representation structure.
Ongoing
- Multimodal generation backbone: latent diffusion. A single latent-diffusion backbone handling text, image, and spatial conditioning in one conditioning path.
- Agentic physics-aware world models.
Agentic and Self-Improving Systems
We study evolutionary search over agent programs, self-play and adversarial training, experience memory, and inference-time optimization under cost budgets.
Representative work
- Memex (RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memory — arXiv 2026. Compresses agent working context without discarding evidence.
- Inference-Time Optimization with Confidence Dynamics — ICML 2026.
Ongoing
- Evolution via adversarial training and self-play. Co-evolving a generator against a learned validator or critic so neither side needs an external reward signal.
- Efficient recursive self-improvement for long-horizon tasks.
Science and Medicine
We develop physics-informed and multimodal AI for science and medicine, spanning inverse problems, learned simulators and clinical diagnosis and support.
Inverse problems
Recovering structure from partial, noisy measurements. Applications: compressed-sensing MRI, sparse-view CT, lung ultrasound, photoacoustic tomography, and functional ultrasound.
Representative work
- A Unified Model for Compressed Sensing MRI Across Undersampling Patterns — CVPR 2025. One model across sampling conditions, resolution-agnostic.
- Resolution-Agnostic Neural Operators for Multi-Rate Sparse-View CT — ECCV 2026. One model over a continuum of measurement rates.
- Ultrasound Lung Aeration Map via Physics-Aware Neural Operators.
Forward models
Learned simulators and surrogates. Applications: turbulence and chaotic systems, multiphysics PDE surrogates, wave and acoustic simulation, and satellite pose estimation.
Representative work
- Beyond Closure Models: Learning Chaotic Systems via Physics-Informed Neural Operators — an end-to-end learning approach using a physics-informed neural operator without a closure model or a coarse-grid solver.
Clinical and Healthcare Data
Multimodal diagnosis and decision support. Applications: ocular surface disease, surgical video and skill assessment, prognosis under class imbalance, and calibrated deferral to clinicians.
Representative work
- Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis — MICCAI 2024.
- Multi-Modal Self-Supervised Learning for Surgical Feedback Effectiveness Assessment — ML4H 2024 (Best Paper).
- Artificial Intelligence Models Utilize Lifestyle Factors to Predict Dry Eye Related Outcomes — Scientific Reports 2025.
Ongoing work
- Shared neural operator across modalities: one trained simulator serves both the inverse and the forward direction.
- Ranking-aware prompt evolution for multimodal clinical diagnosis.
- 4D brain imaging via functional ultrasound.