Ongoing
Cost-Aware Tool Use in LLM Agents
Agents that call external tools have to weigh what a task needs against what each option costs. We study how dependable that trade-off is, and what pulls it off course.
LUMI Lab · Dongguk University
Language Understanding and Machine Intelligence
We research natural language processing, large language models, and trustworthy AI—including LLM evaluation, reliable agents, multimodal systems, and AI safety—under Prof. Yerin Hwang at Dongguk University.
We are growing. Graduate students, undergraduate research interns, and research collaborators are welcome to get in touch.
Our work sits at the intersection of natural language processing, trustworthy AI, agentic systems, and multimodal evaluation. Strong benchmark scores say little about how a model behaves once instructions conflict, evaluations carry bias, or a first attempt needs fixing. We investigate where AI systems fail under that kind of realistic pressure, and build the benchmarks, analyses, and methods that make their behavior more reliable.
Ongoing Projects
Capable systems still fall short in real use. Right now we focus on three of those gaps: weighing cost against capability, keeping priorities straight under conflict, and recovering from a result that missed.
Ongoing
Agents that call external tools have to weigh what a task needs against what each option costs. We study how dependable that trade-off is, and what pulls it off course.
Ongoing
A model is expected to respect which source of instructions outranks another. We study how well that ordering holds up once conflicting instructions start to pile up.
Ongoing
When an edit misses what the user actually wanted, can a model tell that something went wrong and put it right? We work on evaluating multimodal systems on recovery, not only on first attempts.
Research Areas
We study intelligence not only through what models know, but through how they judge, follow priorities, interact with tools, and respond when their first attempt is wrong.
We study bias, robustness, uncertainty, and faithfulness when language and vision–language models are used as evaluators.
We examine instruction hierarchy, tool use, cost-aware decisions, and safety when agents face conflicting or misleading signals.
We evaluate whether models can diagnose failures, express actionable feedback, and carry out precise corrections without causing new errors.
We develop evaluation and training data, metrics, and analyses that reflect language-specific and real-world model behavior.
Publications
Selected research that informs the questions we are pursuing at LUMI.
News
Two to the main conference and one to Findings.
Join LUMI
We welcome students and collaborators who want to study trustworthy AI through careful experiments and evaluations they can stand behind.