Publications

Selected Publications and Prior Work

This page currently highlights selected publications by the PI. Work produced after the launch of LUMI Lab will be marked separately as LUMI Lab research.

2026

  1. How Instruction Hierarchy Breaks Under Pairwise Multi-Conflict Pressure

    Yerin HwangJiwon MoonChangho LeeKwangmin KiTaegwan KangJinsik LeeHonglak LeeKyomin Jung

    EMNLP 2026

    A benchmark study of how instruction-hierarchy compliance deteriorates as multiple conflicts accumulate.

  2. Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

    Jiwon MoonYerin HwangKyomin Jung

    EMNLP 2026

    A study of how the language a prompt is written in affects whether multilingual models respect instruction priority.

  3. Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

    Jiwon Moon*Yerin Hwang*Dongryeol LeeTaegwan KangYongil KimKyomin Jung

    Findings of EACL 2026

    A systematic study of whether superficial code variations bias LLM-based code evaluation.

  4. A Benchmark for the Generation and Evaluation of Scientific Architecture Diagrams in HTML/SVG

    Kyeongman ParkKang-il LeeJanghoon HanYerin HwangTaegwan KangMinwoo LeeChangho LeeKyomin Jung

    Findings of EMNLP 2026

    A benchmark for generating and evaluating scientific architecture diagrams in HTML and SVG.

2025

  1. Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

    Yerin Hwang*Dongryeol Lee*Kyungmin MinTaegwan KangYongil KimKyomin Jung

    EMNLP 2025

    An analysis of visual biases that distort large vision–language model judgments.

  2. Can You Trick the Grader? Adversarial Persuasion of LLM Judges

    Yerin HwangDongryeol LeeTaegwan KangYongil KimKyomin Jung

    Findings of EMNLP 2025

    A study of how persuasive language can shift LLM-judge decisions on correctness-based evaluation tasks.

  3. LLMs can be easily Confused by Instructional Distractions

    Yerin HwangYongil KimJahyun KooTaegwan KangHyunkyung BaeKyomin Jung

    ACL 2025

    A study of inputs that resemble instructions and divert models from the user's intended task.

  4. Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the Effect of Epistemic Markers on LLM-based Evaluation

    Dongryeol Lee*Yerin Hwang*Yongil KimJoonsuk ParkKyomin Jung

    NAACL 2025 · Oral

    An evaluation of whether uncertainty expressions systematically influence LLM judges.

An asterisk after an author name indicates equal contribution.