LUMI Lab · Dongguk University

LUMI Lab

Language Understanding and Machine Intelligence

We research natural language processing, large language models, and trustworthy AI—including LLM evaluation, reliable agents, multimodal systems, and AI safety—under Prof. Yerin Hwang at Dongguk University.

We are growing. Graduate students, undergraduate research interns, and research collaborators are welcome to get in touch.

Researching reliability where it matters most

Our work sits at the intersection of natural language processing, trustworthy AI, agentic systems, and multimodal evaluation. Strong benchmark scores say little about how a model behaves once instructions conflict, evaluations carry bias, or a first attempt needs fixing. We investigate where AI systems fail under that kind of realistic pressure, and build the benchmarks, analyses, and methods that make their behavior more reliable.

Ongoing Projects

What we are working on

Capable systems still fall short in real use. Right now we focus on three of those gaps: weighing cost against capability, keeping priorities straight under conflict, and recovering from a result that missed.

  • Ongoing

    Cost-Aware Tool Use in LLM Agents

    Agents that call external tools have to weigh what a task needs against what each option costs. We study how dependable that trade-off is, and what pulls it off course.

    • LLM Agents
    • Tool Use
    • Cost-Aware Reasoning
  • Ongoing

    Instruction Hierarchy Under Accumulated Conflict

    A model is expected to respect which source of instructions outranks another. We study how well that ordering holds up once conflicting instructions start to pile up.

    • AI Safety
    • Instruction Following
    • Model Behavior
  • Ongoing

    Feedback and Correction in Image Editing

    When an edit misses what the user actually wanted, can a model tell that something went wrong and put it right? We work on evaluating multimodal systems on recovery, not only on first attempts.

    • Multimodal Evaluation
    • Image Editing
    • Feedback

Research Areas

From evaluation to reliable action

We study intelligence not only through what models know, but through how they judge, follow priorities, interact with tools, and respond when their first attempt is wrong.

  • Trustworthy Evaluation

    We study bias, robustness, uncertainty, and faithfulness when language and vision–language models are used as evaluators.

    • LLM-as-a-Judge
    • Evaluation
    • Robustness
  • Reliable Agents

    We examine instruction hierarchy, tool use, cost-aware decisions, and safety when agents face conflicting or misleading signals.

    • Agent Reliability
    • Tool Use
    • AI Safety
  • Feedback & Multimodal Systems

    We evaluate whether models can diagnose failures, express actionable feedback, and carry out precise corrections without causing new errors.

    • Multimodal AI
    • Feedback
    • Image Editing
  • Data-Centric & Korean NLP

    We develop evaluation and training data, metrics, and analyses that reflect language-specific and real-world model behavior.

    • Data Generation
    • Evaluation Data
    • Korean NLP

Publications

Selected Work by the PI

Selected research that informs the questions we are pursuing at LUMI.

  1. How Instruction Hierarchy Breaks Under Pairwise Multi-Conflict Pressure

    Yerin HwangJiwon MoonChangho LeeKwangmin KiTaegwan KangJinsik LeeHonglak LeeKyomin Jung

    EMNLP 2026

  2. Fooling the LVLM Judges: Visual Biases in LVLM-Based Evaluation

    Yerin Hwang*Dongryeol Lee*Kyungmin MinTaegwan KangYongil KimKyomin Jung

    EMNLP 2025

  3. LLMs can be easily Confused by Instructional Distractions

    Yerin HwangYongil KimJahyun KooTaegwan KangHyunkyung BaeKyomin Jung

    ACL 2025

News

Latest from the lab

  1. Three papers accepted to EMNLP 2026

    Two to the main conference and one to Findings.

Join LUMI

Come work on these questions with us.

We welcome students and collaborators who want to study trustworthy AI through careful experiments and evaluations they can stand behind.