Projects

Selected systems, documented at their actual stage.

These projects explore bounded autonomy, auditable evaluation, and privacy-preserving AI. Each entry separates implemented work from current limitations and links to public proof where available.

01Finalist

Casper Agentic Buildathon 2026

RWA Credit Sentinel

A bounded autonomous underwriting and capital-action system for invoice-backed financing. Eight specialized agents turn public evidence into a risk credential, enforce deterministic vault policy, cap executable principal, and anchor both the credential and execution intent on Casper Testnet.

RWA Credit Sentinel visual showing the underwriting, credential, and execution flow
Public finals system: evidence intake, verifiable credential, and bounded capital intent.

My work

Designed and built the end-to-end system: agent workflow, content-hashed evidence, server-locked policy, Casper contract integration, live state verification, deterministic safety benchmark, and public judge experience.

Evidence and boundary

The finals record exposes two independently verifiable application state transitions. The public demo reads both contract dictionaries and compares 13 fields against the published evidence without requiring a wallet.

02In Progress

Auditable safety evaluation for embodied LLM agents

SafeTrace-Agent

An evaluation system for auditing the safety behavior of embodied LLM agents. The current artifact set includes a trace schema, safety taxonomy, rule baseline, evaluators, report generation, an AI2-THOR adapter, and a technical report draft.

My work

Developing the evaluation protocol and implementation with an emphasis on traceability: observations, actions, safety events, evaluator decisions, and report evidence remain linked throughout the pipeline.

Evidence and boundary

Real simulator experiments are still running. No simulator performance numbers or completed empirical claims are reported here.

03

QVAC Hackathon · preliminary screening

LocalVault AI

A local-first confidential contract intelligence workspace for consumer Windows hardware. It combines QVAC local inference with deterministic policy checks, citation-bound findings, amendment drafting, prompt-injection defenses, and auditable report export without remote AI APIs.

My work

Built the document review flow, policy matrix, source-grounded evidence model, local-inference integration, reproducible validation set, and zero-cloud audit evidence.

Evidence and boundary

The public repository documents the validation procedure, hardware context, policy pack, robustness cases, and exported evidence logs.