Yekyung Kim

Yekyung Kim by the coast
Photo by Yapei Chang

I am a fourth-year Ph.D. student at the University of Maryland, CLIP Lab, advised by Mohit Iyyer. I began my Ph.D. at UMass Amherst and moved to UMD with my advisor.

My goal is to build self-improving models that people can rely on as partners in learning and problem-solving. My research focuses on evaluating and improving language models for tasks where output quality is not easily verifiable, like long-form writing.

Self-improving models

Evaluation

What ideas get lost when models write, and why?

Different LLMs converge on the same arguments, often favored by LLM judges and reward models (Argument Collapse).

Can we trust long-form outputs?

Factuality and faithfulness in long-form generation, and long-context understanding (FABLES, VeriScore, OneRuler).

Alignment

Can models learn through simulated interaction?

Building simulated environments where models learn from feedback and revision (ongoing).

Post-training with synthetic datasets

Using synthetic data to improve instruction following (BLEUBERI) and retrieval over complex, real-world documents (ongoing).

News

Aug 2026 Argument Collapse was accepted to EMNLP 2026 (main)!
Jun 2026 Started interning at the Document Intelligence Lab, Adobe (primary mentor: Joe Barrow)
Sep 2025 BLEUBERI was accepted to NeurIPS 2025!
Jul 2025 OneRuler was accepted to COLM 2025!
Sep 2024 VeriScore was accepted to EMNLP Findings 2024.
Jun 2024 Is It Safe to Cross? was named a Best Paper Finalist at UR 2024!
May 2024 FABLES was accepted to COLM 2024.
  1. Argument Collapse: people offer varied arguments while LLMs converge on the same argument

    Argument Collapse: LLMs Flatten Long-Form Public Debate

    Yekyung Kim*, Yapei Chang*, Chau Minh Pham, and Mohit Iyyer
    (* equal contribution)
    EMNLP 2026
  2. BLEUBERI: comparing an answer with a reference to provide a BLEU reward

    BLEUBERI: BLEU is a Surprisingly Effective Reward for Instruction Following

    Yapei Chang, Yekyung Kim, Michael Krumdick, Amir Zadeh, Chuan Li, Chris Tanner, and Mohit Iyyer
    NeurIPS 2025
  3. OneRuler: measuring long-context understanding across languages

    One Ruler to Measure Them All: Benchmarking Multilingual Long-Context Language Models

    Yekyung Kim, Jenna Russell, Marzena Karpinska, and Mohit Iyyer
    COLM 2025
  4. VeriScore: extracting verifiable claims and checking web evidence

    VeriScore: Evaluating the Factuality of Verifiable Claims in Long-Form Text Generation

    Yixiao Song, Yekyung Kim, and Mohit Iyyer
    EMNLP Findings 2024
  5. FABLES: checking a book summary for faithful claims and missing content

    FABLES: Evaluating Faithfulness and Content Selection in Book-Length Summarization

    Yekyung Kim, Yapei Chang, Marzena Karpinska, Aparna Garimella, Varun Manjunatha, Kyle Lo, Tanya Goyal, and Mohit Iyyer
    COLM 2024
  6. Is It Safe to Cross? GPT-4V assesses street-crossing risk from an approaching car

    Is It Safe to Cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing

    Hochul Hwang, Sunjae Kwon, Yekyung Kim, and Donghyun Kim
    Ubiquitous Robots (UR) 2024 (Best Paper Finalist)
  7. LINDA: a model interpolates between two sentences with an adjustable mixing ratio

    LINDA: Unsupervised Learning to Interpolate in Natural Language Processing

    Yekyung Kim, Seohyeong Jeong, and Kyunghyun Cho
    arXiv 2021
  8. InfoVerse: combining model signals to map data and select informative, diverse examples

    InfoVerse: A Universal Framework for Dataset Characterization with Multidimensional Meta-information

    Jaehyung Kim, Yekyung Kim, Karin de Langis, Jinwoo Shin, and Dongyeop Kang
    ACL 2023
  9. Style and content shifts detected using model metadata

    Meta-Crafting: Improved Detection of Out-of-Distributed Texts via Crafting Metadata Space

    Ryan Koo, Yekyung Kim, Dongyeop Kang, and Jaehyung Kim
    AAAI 2024  (Student Abstract & Poster)
  10. Select uncertain diverse sequences for human labeling

    Deep Active Learning for Sequence Labeling Based on Diversity and Uncertainty in Gradient

    Yekyung Kim
    Life-long Learning for Spoken Language Systems Workshop @ AACL 2021
  11. Jamo and character context for Korean named entity recognition

    Learning Sub-Character Level Representation for Korean Named Entity Recognition

    Yejin Kim and Yekyung Kim (equal contribution)
    FLAIRS 2020
  12. Music listening tweets predict future chart rankings

    #Nowplaying the Future Billboard: Mining Music Listening Behaviors of Twitter Users for Hit Song Prediction

    Yekyung Kim, Bongwon Suh, and Kyogu Lee
    SoMeRA Workshop @ SIGIR 2014
  13. Explore topic trends and linked source tweets

    A Visual Analytics Approach to Summarizing Tweets

    Ramik Sadana, Yekyung Kim, Bongwon Suh, and Eunyee Koh
    Industry Day @ SIGIR 2014

Industry projects

Before my Ph.D., I worked at Hyundai Motor Group and LG Electronics as a research engineer. I was selected as a specialist in AI and conducted research at CMU LTI as a visiting scientist mentored by Jaime Carbonell.

  1. Airstar airport robot

    Airstar, Incheon Airport Robot

    LG Electronics
  2. Hyundai in-car AI assistant

    AI Assistant for Cars

    Hyundai Motor Group
  3. LG ThinQ chatbot

    Chatbot for Home Appliances

    LG Electronics

Outside research

Pixel avatar of Yekyung at her laptop

I enjoy video games, especially Dark Souls, Darkest Dungeon, Hollow Knight, and pixel-art games like Stardew Valley. I’m also a fan of escape rooms, but clueless lol. I love finding great boba spots.