Undergraduate researcher · Software engineer

Building AI systems that can be examined, not merely observed.

I am Chenyu Zhang, a Computer Science undergraduate working across trustworthy AI, embodied agents, and verifiable software systems. I am preparing for 2027 Fall master's study in Computer Science and AI.

Selected work

Systems with explicit evidence, boundaries, and status.

01

RWA Credit Sentinel

Finalist

A bounded autonomous underwriting and capital-action system for invoice-backed financing. Eight specialized agents turn public evidence into a risk credential, enforce deterministic vault policy, cap executable principal, and anchor both the credential and execution intent on Casper Testnet.

Read
02

SafeTrace-Agent

In Progress

An evaluation system for auditing the safety behavior of embodied LLM agents. The current artifact set includes a trace schema, safety taxonomy, rule baseline, evaluators, report generation, an AI2-THOR adapter, and a technical report draft.

Read
03

LocalVault AI

A local-first confidential contract intelligence workspace for consumer Windows hardware. It combines QVAC local inference with deterministic policy checks, citation-bound findings, amendment drafting, prompt-injection defenses, and auditable report export without remote AI APIs.

Read

Research direction

Questions I want to pursue in graduate study.

01

AI systems

How agentic systems can remain observable, bounded, and reproducible when models interact with tools, simulators, and external state.

02

Trustworthy AI

Evaluation methods that make safety assumptions, failure modes, evidence provenance, and decision boundaries explicit.

03

Embodied AI

Trace-level analysis of agent behavior in interactive environments, with particular interest in safety-critical action sequences.

04

Agent evaluation

Auditable benchmarks and reporting systems that separate deterministic checks, model judgments, and empirical outcomes.

Verification over assertion

Claims should resolve to something a reader can inspect.

The portfolio distinguishes finished artifacts from active research. Public code, deployed systems, Testnet records, and documented evaluation protocols are linked where available. Simulator experiments and the first-author manuscript remain explicitly in progress.

Open GitHub profile