Responsibilities
Build and evolve production AI agent pipelines for CVE research, patch generation, validation, container image remediation, and end-to-end merge requests.
Design the agent capabilities behind those workflows: prompts, tools, model selection, context, guardrails, and integrations.
Build the feedback loops that make our systems better: tracing production runs, creating datasets, defining evals and scorers, and using results to guide iteration.
Analyze real-world performance to answer the questions that matter: Did this change improve patch success, quality, latency, or cost Can we prove it?
Operate what you build in production across Kubernetes, Argo Workflows, AWS, and GitOps.
Partner closely with Product and engineering to focus on the customer problems with the highest security impact-and ship quickly.
Hands-on experience building with LLMs, agentic architectures, and AI workflows-such as LangGraph, Claude Agent SDK, OpenAI Agents SDK, or equivalent.
Strong experience using coding agents such as **Claude Code, Codex, or Cursor** as part of your day-to-day engineering workflow.
A solid grasp of the mechanics of data science and applied AI: evaluation design, experimental rigor, noisy metrics, statistical reasoning, feedback loops, and evidence-based decision-making.
The judgment to distinguish an impressive demo from a system that reliably works in production.
Strong software engineering skills and an experiment-driven mindset: form a hypothesis, build, measure, learn, and iterate.
Comfort working independently in a fast-moving environment with high ownership and little unnecessary process.
Clear communication and good product instincts.












