What you’ll do
- Own the tutor engine: system prompt design, model routing, cache strategy.
- Design the streaming architecture end-to-end.
- Instrument and optimize inference cost across every route.
- Build the evaluation harness that runs before every model upgrade.
What you bring
- 6+ years backend engineering, at least 2 with LLM APIs in production.
- Fluent in Node/TypeScript and Postgres.
- Have shipped a product where inference cost was a real line item.
- Comfortable making cost/quality tradeoffs with real data.
Bonus, not required
- Familiarity with Anthropic's prompt caching.
- Background in eval design.
- Experience running LLMs behind an SLA.