@ 1001ai
ML Infrastructure Engineer — 1001 (1001.ai). P1. $30M Series A AI company; enterprise deployments including hybrid/on-prem/sovereign cloud. London/Europe, hybrid, GCC travel.
Hands-on ML infra hire who has deployed, served, and monitored models in production. Owns the platform underneath every model shipped: ML serving infrastructure across multiple enterprise deployments; model deployment pipelines (registry, artifacts, datasets, deployment security, multi-tenant serving); monitoring/logging/observability for model performance; inference latency/throughput/cost optimization across CPU and GPU; shared ML platform components for self-serve product teams.
REQUIREMENTS: 4+ years infrastructure or platform engineering with a meaningful chunk in ML/MLOps. Has ACTUALLY run ML in production. Hands-on with model serving (Triton, TorchServe, Ray Serve, vLLM, BentoML, KServe, or equivalent). Strong Kubernetes, containers, CI/CD, Terraform. Production cloud (AWS/GCP/Azure). Observability stack (Prometheus, Grafana, OpenTelemetry). Python and TypeScript. Bonus: GPU optimization, distributed training, large-model serving.
CALIBRATION (ruthless — the previous candidate was withdrawn because "we can do better"; client team is from top labs and top schools and expects that bar): target ML platform/infra engineers from frontier labs (DeepMind, Anthropic, OpenAI, Mistral, Cohere), ML-heavy scale-ups (Hugging Face, Weights & Biases, Modal, Together AI, Baseten, Replicate, CoreWeave), or elite tech companies' ML platform teams (Google, Meta, Uber Michelangelo, Spotify, Netflix). Evidence of serving LLMs/large models in production is a major plus. On-prem/air-gapped/sovereign deployment experience is a differentiator. London/Europe strongly preferred. NOT generic DevOps/SRE with no ML, NOT data engineers who only built ETL, NOT notebooks-only ML people.