We are building the next generation of production AI agent systems — systems that reason, plan, call tools, delegate across specialized sub-agents, and deliver real business outcomes end to end. We are looking for a senior engineer who has gone deep on modern machine learning, transformers and agents, and who wants to build at the frontier of what language models and multimodal systems can actually do in production today.
This is not a "wrap an API around an LLM" role. We expect candidates who understand why things work — the mechanics of attention, the failure modes of autoregressive decoding, the cost/latency trade-offs of caching and tool routing, the pitfalls of context engineering at scale. And who can turn that understanding into systems that ship, hold up under load, and keep improving.
Key Responsibilities:
- Architect agentic systems — plan, memory, tool use, multi-agent delegation, evaluation loops, guardrails. Pick the right abstraction for the problem, not the one on the hype curve.
- Push model capability into production. Design prompt and context strategies, tool interfaces, retrieval and reranking, structured output, streaming, and evaluation — across text and multimodal inputs (vision, documents, audio).
- Own the evaluation story. Build offline eval sets, online LLM-as-judge loops, regression harnesses. Know the difference between a metric that moves your users and a metric that moves only your dashboard.
- Squeeze the system. Prompt caching, batching, speculative decoding, model routing, token budget management, latency targets. Know your P50/P99 and why they look the way they do.
- Contribute upstream. Read SDK source when docs are thin, open PRs against open-source agent frameworks, write crisp bug reports when a vendor's orchestration service returns a weird 500.
- Mentor and set the bar. Your design reviews, code reviews and technical writing shape how the rest of the team thinks about agents.