Lead the applied research vertical of a team building an agentic AI platform, and build the production retrieval and evaluation systems it runs on.
Responsibilities
- Lead the team’s applied research vertical, with a research agenda centered on agentic orchestration and execution-grounded evaluation
- Prototyping agentic execution, in which models draft and execute plans, use structured errors to revise them, and recover failed runs by selecting alternative plans
- Redesigned and shipped the platform’s document retrieval system, combining filename-aware routing with content retrieval across document families and versions
- Built the evaluation layer for the platform’s LLM-generated SQL, measuring the correctness, stability, and groundedness of both queries and their executed results
- Re-engineered the group’s evaluation suite for concurrent execution, cutting a full run from about 3 hours to about 8 minutes
- Built an internal tool that provisions a vLLM inference endpoint for any model in about 30 seconds, now used across the group
Technologies & Skills
- Agentic systems, retrieval-augmented generation, and LLM evaluation
- Python, LangGraph, and vLLM