👋 My name is Xu Zhang. I am currently a PhD student at City University of Hong Kong, advised by Prof. Zhichao Lu. My current research interests lie in MLLM safety, LLM agent security, and vision-language learning.

🔥 News

  • [2026-04] One paper is accepted to ACL 2026 main conference, thanks for all of my collaborators.
  • [2024-summer] 🎓 Graduated from UESTC and joined CityU as a PhD student! 😄
  • [2023-12] One paper is accepted to AAAI 2024, thanks for all of my collaborators.
  • [2023-03] One paper is accepted to ICME 2023, thanks for all of my collaborators.

📝 Publications

ACL 2026 Main
Paper teaser

CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks

Xu Zhang, Hao Li, Zhichao Lu✉

  • Recent studies reveal implicit attacks, in which benign text and image inputs jointly express unsafe intent. Such joint-modal threats are difficult to detect and remain underexplored, largely due to the scarcity of high-quality implicit data.
  • We propose ImpForge, an automated red-teaming pipeline that leverages reinforcement learning with tailored reward modules to generate diverse implicit samples across 14 domains. Building on this dataset, we further develop CrossGuard, an intent-aware safeguard providing robust and comprehensive defense against both explicit and implicit threats.
  • AAAI 2024
    Paper teaser

    Negative Pre-aware for Noisy Cross-Modal Matching

    Xu Zhang, Hao Li, Mang Ye✉

    In this paper, we present a novel Pre-aware Cross-modal (NPC) matching solution for visual-language model fine-tuning on noisy downstream tasks. It is featured in two aspects:

  • We propose to estimate the negative impact of each sample. It does not need additional correction mechanisms that may predict unreliable correction results, leading to self-reinforcing error. This adaptively adjusts the contribution of each sample to avoid noisy accumulation.
  • For maintaining stable performance with increasing noise, we establish a memory bank to estimate the negative impact of each sample. The memory bank can still maintain high quality at a high noise ratio.
  • ICME 2023
    Paper teaser

    Image-text retrieval via preserving main semantics of vision

    Xu Zhang, Xinzheng Niu✉, Philippe Fournier-Viger, Xudong Dai

    Several approaches can map images and texts into a common space to create correspondences between the two modalities. However, due to the content (semantics) richness of an image, redundant secondary information in an image may cause false matches. We presents a semantic optimization approach, implemented as a Visual Semantic Loss (VSL), to assist the model in focusing on an image’s main content. We leverage the annotated texts corresponding to an image to assist the model in capturing the main content of the image, reducing the negative impact of secondary content.

    🎖 Honors and Awards

    • 2024 Academic Youngster Scholarship by Graduate School of UESTC
    • 2023.04 First Prize Scholarship for Graduate Students
    • 2021.06 Provincial Outstanding Graduates
    • 2019.10 Chinese National Scholarship

    🎓 Educations

    • 2024.09 - Present Ph.D, City University of Hong Kong, Kowloon Tong, Hong Kong SAR, China.
    • 2021.09 - 2024.06 Master's Degree, University of Electronic Science and Technology of China (UESTC), Chengdu, Sichuan, China.
    • 2017.09 - 2021.06 Bachelor's Degree, Xihua University, Chengdu, Sichuan, China.