I am currently working on edge-cloud VLMs, quantizing the vision encoder for on-device deployment on a Qualcomm NPU, with a manuscript in preparation on where evaluation breaks between development and deployment. In parallel, I am extending my PhD research line into synthetic data generation for cross-environment ML transfer and explainable AI; a Google TPU Research Cloud proposal is under review. To validate the orchestration pattern for that work, I built a cost-efficient multi-agent LLM pipeline as a PoC.
In 2025, I completed my Ph.D. at Ghent University-imec, and was previously a RA at Academia Sinica. Both my PhD and MSc research were conducted on real-world data, covering 5000+ hours of uncurated audio streams and extended video across multiple deployment sites.
My research focuses on building solutions at the mechanism-level while under real-world constraints. Recognition includes 13 peer-reviewed articles (7 journals, ICIP oral), 150+ citations, APICTA award, and Top 3% master’s graduate honor.
PhD in Computer Science Engineering
Ghent University, Belgium
MSc in Computer and Communication Engineering
National Cheng Kung University, Taiwan
BSc in Electrical Engineering
National Cheng Kung University, Taiwan
ResearchEdge-Cloud VLM: vision encoder quantization for on-device deployment
Quantizing a VLM vision encoder for a Qualcomm QCS8550 NPU. Eleven configurations evaluated on GPU, four actually executable on target. The paper is about where evaluation breaks between development and deployment. Manuscript to be submitted.
ProposalGoogle TPU Research Cloud proposal, under review
Reality-Bounded Synthetic Data via Decoupling and Recombination: extending the PhD line to cross-environment event and behaviour transfer, and person-level stress-testing for deepfake detection.
BuildMulti-agent LLM orchestration PoC
Cost-efficient routing, context isolation, and model gating, validated on a CPU/API-only setup before committing to GPU-heavy compute.
MilestonePh.D. conferred by Ghent University-imec
From Lab to Street: transferable and privacy-friendly deep learning for urban surveillance, built on 5000+ hours of uncurated real-world audio.
PaperSource-free model transferability assessment
Published in MDPI Sensors: ranking model transferability for smart surveillance without access to source data.
PaperEmbedding-based pair generation for contrastive learning
Published in Frontiers in Robotics and AI: audio-visual representation learning on real surveillance streams.