Applied AI Research Engineer
Own the evaluation harness that decides whether an AI approach is fit for a client system: retrieval quality, failure modes, cost per task and regression tracking over time.
What you own
- Build reproducible evaluation pipelines for retrieval, extraction and agentic workflows.
- Design failure taxonomies from real client documents and processes, not public benchmarks.
- Produce reference implementations the delivery practice can reuse directly.
- Write up each finding, including negative results, for publication on the Labs updates page.
What we need
- Production experience shipping an LLM or ML feature that real users depended on.
- Strong Python or TypeScript, and comfort reading model and API documentation critically.
- Ability to state clearly why a result is or is not trustworthy.
Helpful: Experience with structured document extraction · Prior work on evaluation or observability tooling