BEIJING INSTITUTE OF TECHNOLOGY

Human Language Technology Group

HLT 语言技术研究组

Institute of Language Intelligence and Social Computing · School of Computer Science and Technology · Beijing Institute of Technology

Understanding language.
Connecting context and action.

We study how language connects perception and action: from machine translation and multimodal interaction to agents that understand tasks, learn from experience and use tools reliably.

RESEARCH DIRECTIONS

What we study

Starting with language, grounded in real-world understanding and interaction.

01 / LANGUAGE

Language understanding & translation

How can we recover meaning across languages and forms of expression? We explore information extraction, machine translation, in-image translation and translation evaluation.

RATE · PRIM
02 / INTERACTION

Multimodal interaction & agents

How can vision, audio and context turn task understanding into action? We study smart-home interaction, embodied planning, GUI agents and agent memory.

PEAP · HomeBench · LifeMem
03 / RELIABILITY

Model reliability & safety

Can models follow requirements, recognize risks and use tools safely? We study reliability from generated outputs to execution through capability improvement and evaluation.

ReFF · SafeToolBench

SELECTED WORK & OPEN SOURCE

Publications & open source

All repositories
2026EMNLP 2026 · Main
LifeMem Perception & agents

LifeMem: Enabling Lifelong Experience Reuse for LLM Agents

Extracting, organizing and retrieving skills from task trajectories so agents can reuse experience across environments.

2026ACL 2026 · Main
RATE Language & translation

Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation

Evaluating non-literal translation through a human-annotated benchmark and an agentic framework for reflection, retrieval and comparison.

2026ACL 2026 · Main
PEAP Perception & agents

PEAP: Proactive Embodied Action Sequence Planning with Joint Understanding of Vision and Audio Perception

Jointly understanding visual and audio context for proactive embodied action-sequence planning.

2026ACL 2026 · Findings
PEC-Home Perception & agents

PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes

Interpreting progressively elliptical instructions in smart-home interactions using conversational context.

2026EMNLP 2026 · Findings
VGEBench Perception & agents

Towards Generalizable Visually Grounded Exploration of Household Devices

Benchmarking visually grounded exploration of household devices and generalization to new devices and contexts.

2026arXiv · 2026
LivingScreen Perception & agents

Benchmarking Living-Screen-Native GUI Agents on Short-Video Platforms

Evaluating GUI agents on the continuously changing interfaces of short-video platforms.

2025EMNLP 2025 · Main
PRIM Language & translation

PRIM: Towards Practical In-Image Multilingual Machine Translation

Exploring practical multilingual translation within images, connecting textual meaning with visual presentation.

2025ACL 2025 · Main
HomeBench Perception & agents

HomeBench: Evaluating LLMs in Smart Homes with Valid and Invalid Instructions Across Single and Multiple Devices

Evaluating smart-home instruction understanding across single and multiple devices, including valid and invalid instructions.

2025EMNLP 2025 · Findings
SafeToolBench Reliability & safety

SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs

Evaluating the safety of tool use by large language models through a dedicated benchmark.

2025AAAI 2025
ReFF Reliability & safety

ReFF: Reinforcing Format Faithfulness in Language Models across Varied Tasks

Reinforcing output-format faithfulness across tasks to make language-model outputs more reliable for downstream applications.

Find more publications on the full publication list and ACL Anthology.

TEACHING

Teaching & learning

Build understanding from core ideas and interactive examples.

DS2026

DATA STRUCTURES

Data Structures

Course information, chapter summaries and interactive materials released after class.

PEOPLE & COLLABORATION

People & collaboration

Institute of Language Intelligence and Social Computing · School of Computer Science and Technology · Beijing Institute of Technology

Group lead

Yuhang Guo 郭宇航

We welcome discussions on language intelligence, multimodal interaction and reliable agents.

Doctoral students

  • 姚嘉树
  • 田炎智
  • 曾理
  • 刘艳平
  • 钱博傲

Master’s students

  • 邱昱力
  • 单赢宇
  • 郑林昊
  • 文浩宇
  • 苏炜
  • 张红博
  • 艾华喜
  • 陈旺科