Preprints and Technical Reports
Reviewed Papers
EMNLP 2026 (Findings)
Hallucination as Exploit: Evidence-Carrying Multimodal Agents
Grounds multimodal-agent outputs in traceable evidence to expose and mitigate hallucination exploits.
Guijia Zhang, Hao Zheng, Harry Yang
ACM Multimedia 2026
AffordanceSAM: Segment Anything Once More in Affordance Grounding
Extends segment-anything models for precise, open-vocabulary affordance grounding.
Dengyang Jiang, Zanyi Wang, Hengzhuang Li, Liuzhuozheng Li, Sizhe Dang,
Guang Dai, Harry Yang, Mengmeng Wang
ECCV 2026
GKDT: General Keypoint Detection Transformer
Generalizes keypoint detection across categories with a unified transformer architecture.
Changsheng Lu, Yuxin Chen, Haokun GUI, Rong Wang, Jie Yang, Harry
Yang, Anton van den Hengel, Jiaya Jia
ECCV 2026
RotateAttention: RoPE-Aware Rotation and Range Rectification for INT4 Quantized
Attention in Video Generation
Enables accurate INT4 attention for faster, more memory-efficient video generation.
LIU Yaofu, Lan Wangli, lijinxi, Binhang Yuan, Harry
Yang
ECCV 2026
OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for
Subject-driven Image Generation and Manipulation
Uses video-derived priors for identity-consistent, diverse subject generation and editing.
Yexin Liu, Manyuan Zhang, Yueze Wang, Hongyu Li, Dian Zheng, Weiming
Zhang, Changsheng Lu, Xunliang Cai, Yan Feng, Peng Pei, Harry Yang
ECCV 2026
Distribution Matching Distillation Meets Reinforcement Learning
Combines distribution-matching distillation with reinforcement learning for efficient few-step generation.
Dengyang Jiang, Dongyang Liu, Zanyi Wang, Qilong Wu, Liuzhuozheng Li,
Hengzhuang Li, Xin Jin, David Liu, Changsheng Lu, Zhen Li, Bo Zhang, Mengmeng Wang, Steven Hoi, Peng
Gao, Harry Yang
arXiv 2026
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
Adapts Z-Image and Z-Image-Turbo from latent-space to pixel-space generation.
Dengyang Jiang, Ruoyi Du, Zhennan Chen, Dongyang Liu, Zanyi Wang,
Mingzhe Zheng, Xiangpeng Yang, Huanqia Cai, Aiming Hao, Yuming Jiang, Peng Gao,
Harry Yang (corresponding author), Steven Hoi
arXiv 2026
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus
Structure
Diagnoses when GUI agents rely on pixels versus structured state representations.
Guijia Zhang, Harry Yang
arXiv 2026
D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion
Models
Continuously tunes step-distilled diffusion models through on-policy self-distillation.
Dengyang Jiang, Xin Jin, Dongyang Liu, Zanyi Wang, Mingzhe Zheng, Ruoyi Du,
Xiangpeng Yang, Qilong Wu, Zhen Li, Peng Gao, Harry Yang (corresponding
author), Steven Hoi
arXiv 2026
AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers
Adaptively aligns hierarchical priors to improve diffusion-transformer generation.
Ruibin Min, Yexin Liu, Aimin Pan, Changsheng Lu, Jiafei Wu, Kelu Yao, Xiaogang Xu,
Harry Yang
ICML 2026
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided
Image-to-Video Generation
Improves semantic fidelity in text-guided image-to-video generation without additional training.
Yexin Liu, Wen-Jie Shu, Zile Huang, Haoze Zheng, Yueze Wang, Manyuan Zhang,
Ser-Nam Lim, Harry Yang
ICML 2026
Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
Scales confidence during autoregressive decoding to improve image-generation quality.
Harold Haodong Chen, Xianfeng Wu, Wen-Jie Shu, Rongjin Guo, Disen Lan, Harry Yang,
Ying-Cong Chen
CVPR 2026
Learning Latent Proxies for Controllable Single-Image Relighting
Learns latent proxies for controllable relighting from a single input image.
Haoze Zheng, Zihao Wang, Xianfeng Wu, Yajing Bai, Yexin Liu, Yun Li, Xiaogang Xu,
Harry Yang
CVPR 2026
Group Editing: Edit Multiple Images in One Go
Edits multiple images jointly for consistent, coordinated group transformations.
Yue Ma, Xinyu Wang, Qianli Ma, Qinghe Wang, Mingzhe Zheng, Xiangpeng Yang, Hao
Li, Chongbo Zhao, Jixuan Ying, Harry Yang, Hongyu Liu, Qifeng Chen
DenDiff: Density-Guided Diffusion for Quantity-Aware Image Synthesis
Uses density guidance to control object quantities in diffusion-based image synthesis.
Bo Gao, Haoyu Liang, Harry Yang, Ser-Nam Lim
CVPR 2026 (Findings)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
Introduces a large, diverse dataset for audio-driven talking-head synthesis.
Shunian Chen, Hejin Huang, Yexin Liu, Zihan Ye, Pengcheng Chen, Chenghao Zhu,
Michael Guan, Rongsheng Wang, Junying Chen, Jianye Hou, Bo Li, Guanbin Li, Ser-Nam Lim, Harry Yang,
Benyou Wang
ICLR 2026
AC-Foley: Reference-Audio-Guided Video-to-Audio Synthesis with Acoustic Transfer
Transfers acoustic characteristics from reference audio into video-to-audio generation.
Pengjun Fang, Yingqing He, Yazhou Xing, Qifeng Chen, Ser-Nam Lim, Harry Yang
ICLR 2026
EditAnyShape: Shape-Aware Image Editing via Trajectory-Guided Region Control
Uses shape-aware trajectories to control localized image edits.
Zeqian Long, Mingzhe Zheng, Kunyu Feng, Xinhua Zhang, Hongyu Liu, Harry Yang,
Linfeng Zhang, Qifeng Chen, Yue Ma
ICLR 2026
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
Formulates referring video segmentation as flow matching from videos to masks.
Zanyi Wang, Dengyang Jiang, Liuzhuozheng Li, Sizhe Dang, Chengzu Li, Harry Yang,
Guang Dai, Mengmeng Wang, Jingdong Wang
Technical Report
INT4 Quantization for FlashAttention
Quantizes FlashAttention to INT4 for more efficient video-generation inference.
Full paper accepted to ECCV 2026.
Yaofu Liu, Harry Yang
arXiv 2026
Thinking in Loops: Scaling Visual ARC with Looped Transformers
Scales Visual ARC reasoning by recurrently looping transformer computation.
Wen-Jie Shu, Xuerui Qiu, Rui-Jie Zhu, Harold Haodong Chen, Yexin Liu, Harry Yang
arXiv 2026
SAGE-GRPO: Manifold-Aware Exploration for Reinforcement Learning in Video
Generation
Uses manifold-aware exploration to stabilize reinforcement learning for video generation.
Mingzhe Zheng, Weijie Kong, Yue Wu, Dengyang Jiang, Yue Ma, Xuanhua He, Bin Lin,
Kaixiong Gong, Zhao Zhong, Liefeng Bo, Qifeng Chen, Harry Yang
Journal of Technology in Behavioral Science
Reducing depressive symptoms through AI-guided narrative self-films: Results from a
randomized controlled trial
Tests whether AI-guided narrative self-films can reduce depressive symptoms.
Elvin Yao, Harry Yang
Health Communication
Message Framing and Causal Attributions Shape Public Reactions to Parkinson's
Disease: Evidence from China and the United States
Studies how message framing and causal attribution shape perceptions of Parkinson's disease.
Conditionally accepted; final version submitted.
Elvin Yao, Harry Yang
Publication forthcoming
AAAI 2026
Next Patch Prediction for AutoRegressive Visual Generation
Predicts the next visual patch to improve autoregressive image generation.
Yatian Pang, Peng Jin, Shuo Yang, Bin Lin, Bin Zhu, Zhenyu Tang, Liuhan Chen,
Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Li Yuan
COLM 2025
Meta-Learning for Speeding Up Large Model Inference in Decentralized Environments
Uses meta-learning to accelerate large-model inference across decentralized environments.
Yuzhe Yang, Yipeng Du, Ahmad Farhan, Claudio Angione, Yue Zhao, Harry Yang,
Fielding Johnston, James Buban, Patrick Colangelo
NeurIPS 2025
Hierarchical Fine-Grained Preference Optimization for Physically Plausible Video
Generation
Optimizes hierarchical fine-grained preferences for more physically plausible videos.
Harold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang
(corresponding author), Ser-Nam Lim
NeurIPS 2025
When Semantics Mislead Vision: Mitigating Hallucinations in MLLMs
Reduces multimodal-model hallucinations caused by misleading semantic priors.
Yan Shu, Hangui Lin, Yexin Liu, Yan Zhang, Gangyan Zeng, Yan Li, Yu Zhou, Ser-Nam
Lim, Harry Yang, Nicu Sebe
NeurIPS 2025 NextVid Workshop (Oral)
VideoGen-of-Thought: Step-by-Step Generation of Multi-Shot Videos
Generates coherent multi-shot videos through an explicit step-by-step planning process.
Mingzhe Zheng, Yongqi Xu, Haojian Huang, Xuran Ma, Yexin Liu, Wenjie Shu, Yatian
Pang, Feilong Tang, Qifeng Chen, Harry Yang, Ser-Nam Lim
ICCV 2025
DreamDance: Animating Human Images by Enriching 3D Geometry Cues
Enriches 3D geometry cues to animate a person from a single image.
Yatian Pang, Bin Zhu, Bin Lin, Mingzhe Zheng, Francis E. H. Tay, Ser-Nam Lim,
Harry Yang, Li Yuan
ICCV 2025
Model Reveals What to Cache: Profiling-Based Feature Reuse
Profiles diffusion models to identify and reuse cacheable intermediate features.
Xuran Ma, Yexin Liu, Yaofu Liu, Xianfeng Wu, Mingzhe Zheng, Zihao Wang, Ser-Nam
Lim, Harry Yang
CVPR 2025
Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly
Benchmarks cases where multimodal models perceive correctly but answer incorrectly.
Yexin Liu, Zhengyang Liang, Yueze Wang, Xianfeng Wu, Feilong Tang, Muyang He,
Jian Li, Zheng Liu, Harry Yang, Ser-Nam Lim, Bo Zhao
IEEE TCSVT 2025
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a
Single View
Reconstructs 3D scenes from one view using normal-integrated affine geometry.
Xianzu Wu, Zhenxin Ai, Harry Yang, Ser-Nam Lim, Jun Liu, Huan Wang
ICLR 2025
Intervening Anchor Token: Decoding Strategy in Alleviating Hallucinations
Intervenes on anchor tokens during decoding to reduce language-model hallucinations.
Feilong Tang, Zile Huang, Chengzhi Liu, Qiang Sun, Harry Yang, Ser-Nam Lim
ICLR 2023
Make-A-Video: Text-to-Video Generation without Text-Video Data
Generates videos from text without requiring paired text-video training data.
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan
Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, Yaniv Taigman
ECCV 2022
Long Video Generation with Time-Agnostic VQGAN and Time-Sensitive Transformer
Combines a time-agnostic VQGAN with a time-sensitive transformer for long-video generation.
Songwei Ge, Thomas Hayes, Harry Yang, Xi Yin, Guan Pang, David Jacobs, Jia-Bin
Huang, Devi Parikh
CVPR 2017
High-Resolution Image Inpainting using Multi-Scale Neural Patch Synthesis
Synthesizes multi-scale neural patches for high-resolution image inpainting.
Chao Yang, Xin Lu, Zhe Lin, Eli Shechtman, Oliver Wang, Hao Li
August 2026
Hallucination as Exploit: Evidence-Carrying Multimodal Agents
accepted to EMNLP 2026 Findings.
August 2026
Serving as Area Chair for ICLR 2027.
August 2026
Served as a judge and roundtable guest at the
2026 FIREBIRD Hackathon, organized by the HKUST Mainland Alumni
Association.
July 2026
Invited to give an online seminar at the Graduate School of Artificial
Intelligence, UNIST.
July 2026
Awarded the 2026-2027 GHMUA Exchange Events on Greater Bay Area
Fund for the GBA Workshop on Generative Video and Human-AI
Collaboration.
July 2026
Presenting at ICML 2026 in Seoul.
July 2026
Awarded the AIS Support Fund for Interdisciplinary Research
Collaboration for VitaClaw.
July 2026
AffordanceSAM accepted to ACM Multimedia 2026.
July 2026
Awarded the AUA Scholars Award Program 2026 to support an
academic visit to Tsinghua University.
June 2026
Serving as Area Chair for AAAI 2027.
June 2026
Four papers accepted to ECCV 2026.
June 2026
DenDiff received the Compute Transparency Champion
award at CVPR 2026.
May 2026
Serving on the Program Committee for Pacific Graphics (PG)
2026.
May 2026
Recognized as Gold Reviewer for ICML 2026.
May 2026
Two papers (AlignVid, ScalingAR) accepted to
ICML 2026.
Apr 2026
Serving as Area Chair for ICIP 2026.
Apr 2026
Giving a talk on World Models at the Von Neumann
Institute.
Apr 2026
Speaker at
The Scaling Summit: AI Agents & Autonomous
Systems, hosted by
499 & City University of Hong Kong, co-hosted
by
0G. Panel:
Silicon Meets Carbon: From Embodied AI to Biological
Longevity, with Prof. Yu Wang (CityU), Dr. Roy Rong (Longevity Pioneer), Winnie Qiu
(Avinasi Labs), moderated by June (Facet.ai & BayAI Venture Association).
Mar 2026
Serving as Area Chair for NeurIPS 2026.
Feb 2026
Five papers accepted to CVPR 2026 (2 in main conference and 3 in
findings).
Feb 2026
Visit XbotPark in Songshanhu and meet with Prof. Zexiang
Li.
Feb 2026
Panel speaker at Desci Hong Kong 2026.
Jan 2026
Three papers accepted to ICLR 2026.
Dec 2025
Awarded 2026 HKUST-UStA Global Knowledge Network Awards/Joint Seed
Funding.
Dec 2025
Exhibiting Cosmos Mapping: Unlimited Exploration (with Yuyang
Jiang) at Touching the Void: Art Without an Object. Dec 19-21. Blanc Gallery, 15 E 40th
St, New York.
Dec 2025
Selected by
Hong Kong Monetary Authority (HKMA) for GenAI Sandbox
testing (Phase 2).
SCMP
Dec 2025
Giving talk and serving as panel at SIGGRAPH Asia Birds of a
Feather: "Working in an Interdisciplinary Department: When Art and Technology Intertwine".
Dec 2025
Organizer for the exhibition 【AMC ×
CMA】共息信号:缠联的感知与算法|港科大
× 港科广 第二届跨校区艺术展览 (Entangled Signals: Perceptions and Algorithms Entwined).
Dec 2025
Invited Mr. Xiangchen Kong (Zhejiang TV) and Mr. Yu
Chen (Renowned Designer) for job talks.
Nov 2025
Giving a keynote speech at From Vibe to Viable, Build Real Apps with AI
Coding + No-Code Workshop - Hong Kong.
Oct 2025
Serving as Area Chair for ICLR 2026 and
AAAI 2026.
Sep 2025
Giving a
talk
at
HKUST-GZ CMA Seminar.
July 2025
Internship placements: Congratulations to my first-year PhD students for securing
research internships at Kuaishou (Kling),
Tencent, and Bytedance.
June 2025
Awarded GRF/ECS 2025-26 funding.
June 2025
Serving as Associate Editor for APSIPA Transactions on Signal and
Information Processing.
June 2025
Visiting Meta AI, New York.
June 2025
- Interviewed by RTHK on GenAI and video generation,
discussing applications in Hong Kong local community. (Cantonese)
- Interviewed by CNN on Embodied AI, discussing the current status,
challenges, and future directions.
Feb 2025
- Panelist at the Hong Kong Web3 & AI Builder
Workshop, invited by Prof. Xiaofan Liu, speaking on “When AI Meets
Web3: Redefining the Future for Developers and Builders.”
- Panel speaker at DeSci HK 2025 during the 2025 HKG Consensus Web3
Conference (hosted by CityU), invited by Prof. Yu Wang.
Jan 2025
Awarded HKUST-POSTECH Joint Research Seed Grant.
Dec 2024
Video project approved for HKSTP Incubation Program (3 years).
Aug 2024
Hosted a roundtable at Foresight 2024 in Hong Kong.
Sep 2024
Visited Abu Dhabi (Royal Family meeting) and Token2049 Singapore.