사진: Yerin Hwang

Yerin Hwang황예린

조교수

동국대학교 컴퓨터AI학부

언어·멀티모달 모델을 평가자나 에이전트로 쓸 때, 또 모델이 주어진 지시를 따라야 할 때 생기는 신뢰성 문제를 연구합니다. 이런 실패를 어떻게 측정하고 줄일 수 있을지가 주된 관심사입니다. 서울대학교에서 인공지능 전공으로 박사학위를, 전기정보공학부에서 학사학위를 받았습니다.

동국대학교에 오기 전에는 SNU MILAB과 LG AI Research, 그리고 독일 Max Planck Institute for Security and Privacy에서 연구했습니다.

연구 관심 분야

  • Natural Language Processing
  • Large Language Models
  • LLM Agents
  • LLM Evaluation
  • Trustworthy AI
  • AI Safety
  • Multimodal AI
  • Korean NLP

함께할 분을 찾고 있습니다

신뢰할 수 있는 AI와 꼼꼼한 실증 연구에 관심 있는 대학원 진학 희망자와 학부 연구인턴의 문의를 환영합니다!

Join LUMI

Publications

Show publicationsHide publications

LUMI Lab 설립 이후의 연구는 LUMI Lab 연구로 따로 표시됩니다.

2026

  1. How Instruction Hierarchy Breaks Under Pairwise Multi-Conflict Pressure

    Yerin HwangJiwon MoonChangho LeeKwangmin KiTaegwan KangJinsik LeeHonglak LeeKyomin Jung

    EMNLP 2026

    A benchmark study of how instruction-hierarchy compliance deteriorates as multiple conflicts accumulate.

  2. Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

    Jiwon MoonYerin HwangKyomin Jung

    EMNLP 2026

    A study of how the language a prompt is written in affects whether multilingual models respect instruction priority.

  3. Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

    Jiwon Moon*Yerin Hwang*Dongryeol LeeTaegwan KangYongil KimKyomin Jung

    Findings of EACL 2026

    A systematic study of whether superficial code variations bias LLM-based code evaluation.

  4. A Benchmark for the Generation and Evaluation of Scientific Architecture Diagrams in HTML/SVG

    Kyeongman ParkKang-il LeeJanghoon HanYerin HwangTaegwan KangMinwoo LeeChangho LeeKyomin Jung

    Findings of EMNLP 2026

    A benchmark for generating and evaluating scientific architecture diagrams in HTML and SVG.

2025

  1. Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

    Yerin Hwang*Dongryeol Lee*Kyungmin MinTaegwan KangYongil KimKyomin Jung

    EMNLP 2025

    An analysis of visual biases that distort large vision–language model judgments.

  2. Can You Trick the Grader? Adversarial Persuasion of LLM Judges

    Yerin HwangDongryeol LeeTaegwan KangYongil KimKyomin Jung

    Findings of EMNLP 2025

    A study of how persuasive language can shift LLM-judge decisions on correctness-based evaluation tasks.

  3. LLMs can be easily Confused by Instructional Distractions

    Yerin HwangYongil KimJahyun KooTaegwan KangHyunkyung BaeKyomin Jung

    ACL 2025

    A study of inputs that resemble instructions and divert models from the user's intended task.

  4. Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the Effect of Epistemic Markers on LLM-based Evaluation

    Dongryeol Lee*Yerin Hwang*Yongil KimJoonsuk ParkKyomin Jung

    NAACL 2025 · Oral

    An evaluation of whether uncertainty expressions systematically influence LLM judges.

* 공동 제1저자.