For a complete and up-to-date list of my publications, please see my Google Scholar.

Efficient Model Adaptation

Alfa paper thumbnail
AAAI 2026

Alfa: Attentive Low-Rank Filter Adaptation for Structure-Aware Cross-Domain Personalized Gaze Estimation

He-Yen Hsieh, Wei-Te Mark Ting, H. T. Kung

Alfa adapts gaze models to new users and domains with only a few unlabeled samples by reweighting useful spatial patterns already learned in pre-trained filters. The same low-rank adaptation idea also extends to diffusion-based language models for reasoning tasks.

Efficient adaptation 路 Cross-domain personalization 路 Gaze estimation 路 PEFT/LoRA 路 Model compression
DFT Gaze paper thumbnail
ICIP 2025 Spotlight

DFT Gaze: Distilled and Fine-Tuned Gaze Estimation for Personalization on Tiny Devices

He-Yen Hsieh, Ziyun Li, Sai Qian Zhang, Wei-Te Mark Ting, Kao-Den Chang, Barbara De Salvo, Chiao Liu, H. T. Kung

DFT Gaze combines distillation and fine-tuning into one solution to train a tiny 281K-parameter model for accurate personalized gaze estimation on resource-constrained devices.

AR/VR 路 Eye tracking 路 Edge AI 路 5-shot personalization 路 Model compression

Efficient AI Deployment

EDIT paper thumbnail
NeurIPS OPT 2025

EDIT: Early Diffusion Inference Termination for dLLMs Based on Dynamics of Training Gradients

He-Yen Hsieh, Hong Wang, H. T. Kung

EDIT speeds up diffusion-based language model reasoning by using training-gradient dynamics to stop inference once generation has stabilized.

Reasoning LLMs 路 Diffusion language models 路 Efficient inference 路 Test-time compute

Generative AI

Gen4Gen paper thumbnail
BMVC 2025

Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition

Chun-Hsiao Yeh*, Ta-Ying Cheng*, He-Yen Hsieh*, Chuan-En Lin, Yi Ma, Andrew Markham, Niki Trigoni, H. T. Kung, Yubei Chen

Gen4Gen builds a generative data pipeline for creating high-quality multi-concept image-text pairs and improving multi-concept personalized generation.

Generative AI 路 Synthetic data 路 Multi-concept personalization 路 Text-to-image generation 路 Benchmark dataset
GazeGen paper thumbnail
arXiv 2024

GazeGen: Gaze-Driven User Interaction for Visual Content Generation

He-Yen Hsieh, Ziyun Li, Sai Qian Zhang, Wei-Te Mark Ting, Kao-Den Chang, Barbara De Salvo, Chiao Liu, H. T. Kung

GazeGen turns eye gaze into a natural control signal for image and video generation, enabling hands-free visual content manipulation and interaction.

Interactive AI 路 Gaze control 路 Visual content generation 路 AR/VR 路 Human-computer interaction

Computer Vision

S3R paper thumbnail
ECCV 2022

Self-Supervised Sparse Representation for Video Anomaly Detection

Jhih-Ciang Wu*, He-Yen Hsieh*, Ding-Jie Chen, Chiou-Shann Fuh, Tyng-Luh Liu

S3R detects abnormal events in long videos by learning self-supervised sparse representations, reducing the need for dense manual anomaly annotations.

Video anomaly detection 路 Self-supervised learning 路 Weak supervision 路 Safety AI 路 Video understanding
Adaptive Image Transformer paper thumbnail
CVPR 2021

Adaptive Image Transformer for One-Shot Object Detection

Ding-Jie Chen, He-Yen Hsieh, Tyng-Luh Liu

Adaptive Image Transformer detects new object categories from only one query example by adapting proposal features to better match the given object.

Object detection 路 One-shot learning 路 Data-efficient vision 路 Visual detection 路 Attention
One-Shot Action Detection via Attention Zooming In paper thumbnail
ICASSP 2023

One-Shot Action Detection via Attention Zooming In

He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu

This work localizes unseen actions in long videos from only one support image by using attention zooming to progressively refine action proposals.

Action localization 路 One-shot learning 路 Video understanding 路 Attention zooming 路 Temporal action detection
Aggregating Bilateral Attention paper thumbnail
WACV 2023

Aggregating Bilateral Attention for Few-Shot Instance Localization

He-Yen Hsieh, Ding-Jie Chen, Cheng-Wei Chang, Tyng-Luh Liu

ABA improves few-shot instance localization by aggregating query-support attention, helping models find unseen actions and objects with only a few examples.

Few-shot learning 路 Instance localization 路 Action localization 路 One-shot object detection 路 Attention
Contextual Proposal Network paper thumbnail
WACV 2022

Contextual Proposal Network for Action Localization

He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu

CPN generates high-quality temporal action proposals by modeling multi-scale temporal context and boundary cues in long untrimmed videos.

Temporal action localization 路 Action proposal generation 路 Video understanding 路 Boundary prediction 路 Context modeling
Temporal Action Proposal Generation via Deep Feature Enhancement paper thumbnail
ICIP 2020

Temporal Action Proposal Generation via Deep Feature Enhancement

He-Yen Hsieh, Ding-Jie Chen, Tyng-Luh Liu

This work improves temporal action proposal generation by enhancing video segment features and expanding multi-granularity proposal representations.

Temporal action proposal generation 路 Feature enhancement 路 Video understanding 路 Action localization 路 Proposal representation

Notes

* Equal contribution.

More publications are available on my Google Scholar.