When passing tests becomes the reward
How I built an execution-based RL evaluation pipeline, repaired a verifier blind spot, and measured whether the repair improved training.
ENGINEERING & RESEARCH
I build AI systems and study
what makes them reliable.
Founding Engineer at Miravoice, working on production voice agents. My research interests sit around model evaluation, reinforcement learning, and the gap between passing a test and getting it right.
Verifier-RL
An experimental study of how incomplete code graders affect reinforcement learning—and whether repairing a verifier improves what a model actually learns.
I ran 12 GRPO post-training experiments on Qwen2.5-Coder-1.5B, comparing three verifier conditions across four paired seeds on a booking-capacity task.
A repaired verifier rejected all 38 observed false acceptances while retaining all 440 audit-passing program draws, without increasing the test budget. Replication did not establish a training benefit.
The engineering work included sandboxed parallel code execution, durable results, and full-state training recovery.
Evaluation, conversation recovery, and reliable deployment at Miravoice.
APPLIED AI · SYSTEMSAs a Founding Engineer, I evaluate instruction following, tool-call correctness, latency, and interruption recovery. I use failed-conversation replays to guide prompt and runtime changes.
My work also spans speech recognition and dialogue-state failures, asynchronous deployment orchestration, and campaign delivery.
Evidence-grounded patent research with retrieval and agent workflows.
RETRIEVAL · AGENTSA LangGraph system that routes patent questions, expands technical queries, and synthesizes reports from retrieved patent and litigation evidence.
It combines Qdrant vector search with PostgreSQL metadata, concurrent enrichment, programmatic citation-ID checks, and structured critic feedback with capped retries.
Studying spurious correlations and learned masks in graph models.
GRAPH LEARNING · ROBUSTNESSI co-developed a graph-learning pipeline with learned node masks, counterfactual feature perturbations, and consistency and sparsity objectives.
The work included diagnosing attention-mask collapse, adjusting initialization, and evaluating classification performance and learned masks on synthetic motif graphs.
How I built an execution-based RL evaluation pipeline, repaired a verifier blind spot, and measured whether the repair improved training.
From four scored answers to a policy update: my notes on group-relative advantages, clipping, and a loss that can be zero while learning continues.
I’m an engineer based in the Bay Area, working on AI systems at Miravoice and exploring research questions through independent projects.
I like working across the full path from an experiment to a system that has to work in practice. My work spans model evaluation, reinforcement learning, and the engineering that makes AI systems reliable.
I’m especially interested in what our tests capture, what they miss, and how those choices shape model behavior.
Have a question or a shared interest?
Say hello.