we are looking for a ML Engineer.
Own the model plane end to end – vLLM inference, fine-tuning (GRPO/RL and LoRA), and the evaluation-and-promote pipeline that takes an owned model from checkpoint to production serving. We train and serve our own models on GPUs in our own cluster.
Own the model plane end to end – vLLM inference, fine-tuning (GRPO/RL and LoRA), and the evaluation-and-promote pipeline that takes an owned model from checkpoint to production serving. We train and serve our own models on GPUs in our own cluster.
Requirements:
2+ years of ML engineering or applied ML research
Deep familiarity with HuggingFace transformers
Experience running inference servers (vLLM, TGI, or Triton)
Python and CUDA fundamentals
Understanding of quantization (AWQ, GPTQ, GGUF)
Bonus: GRPO/RLHF/DPO training, GPU workloads on Kubernetes
2+ years of ML engineering or applied ML research
Deep familiarity with HuggingFace transformers
Experience running inference servers (vLLM, TGI, or Triton)
Python and CUDA fundamentals
Understanding of quantization (AWQ, GPTQ, GGUF)
Bonus: GRPO/RLHF/DPO training, GPU workloads on Kubernetes
This position is open to all candidates.











