Research

Evaluation is part of the system, not an afterthought.

My interests sit at the boundary between AI behavior and system design: how an agent's actions are constrained, recorded, evaluated, and communicated to people who need to trust the result.

01

AI systems

How agentic systems can remain observable, bounded, and reproducible when models interact with tools, simulators, and external state.

02

Trustworthy AI

Evaluation methods that make safety assumptions, failure modes, evidence provenance, and decision boundaries explicit.

03

Embodied AI

Trace-level analysis of agent behavior in interactive environments, with particular interest in safety-critical action sequences.

04

Agent evaluation

Auditable benchmarks and reporting systems that separate deterministic checks, model judgments, and empirical outcomes.

Current system research

In Progress

SafeTrace-Agent

An auditable safety evaluation system for embodied LLM agents, designed to preserve the path from environment observation to action, safety event, evaluator decision, and report evidence.

Implemented artifacts

  • Structured trace schema and safety taxonomy
  • Rule-based baseline and evaluator pipeline
  • Report generation with evidence-linked findings
  • AI2-THOR adapter and technical report draft

Current boundary

Real simulator experiments are ongoing. This portfolio does not report success rates, comparative results, or completed empirical conclusions before those runs are finished and reviewed.

First-author manuscript

In Progress

Spatial-scale driving mechanisms of land-use conflict in the Yellow River Basin

A first-author study examining how the drivers of land-use conflict vary across spatial scales in the Yellow River Basin. The manuscript is being prepared for SCI submission and is not described as submitted or accepted.

Research role

Project lead within a National Undergraduate Innovation Program, responsible for coordinating the research process and developing the first-author manuscript.

Reporting policy

A paper link will be added after a public preprint or formal submission artifact is available. No publication or acceptance claim is made at this stage.