1001 AI — ML Infrastructure Engineer (P1)
About 1001 AI: AI-native OS for MENA heavy industries (aviation, construction, ports, energy). Lux Capital / General Catalyst / CIV backed. Engineering team mostly ex-Scale AI / Palantir. Locations: London + Doha + Dubai + Abu Dhabi.
One-line pitch: Hands-on ML infra hire who has deployed, served, and monitored models in production. Owns the platform underneath every model we ship.
Responsibilities:
- Build and operate ML serving infrastructure across multiple enterprise deployments
- Own model deployment pipelines — model registry, model artifacts, datasets, model deployment security, multi-tenant serving
- Stand up monitoring, logging, and observability for model performance
- Optimize inference latency, throughput, and cost across CPU and GPU workloads
- Maintain shared ML platform components so product teams can self-serve
Requirements:
- 4+ years in infrastructure or platform engineering, with a meaningful chunk in ML / MLOps
- Has actually run ML in production
- Hands-on with model serving (Triton, TorchServe, Ray Serve, vLLM, BentoML, KServe, or equivalent)
- Strong with Kubernetes, containers, CI/CD, and infrastructure-as-code (Terraform)
- Production cloud experience (AWS / GCP / Azure)
- Monitoring and observability stack (Prometheus, Grafana, OpenTelemetry, or equivalent)
- Python and TypeScript
- Comfortable with hybrid / on-prem / sovereign cloud enterprise deployments (not just clean SaaS)
Bonus:
- GPU optimization, distributed training, large-model serving
- Prior enterprise deployment experience
- Mentorship of platform users / consumer teams
Pedigree priority: Anyscale, Modal Labs, Replicate, Together AI, Fireworks AI, Baseten, OctoAI, BentoML, RunPod, Scale AI infrastructure, Anthropic Compute, OpenAI Compute, Google DeepMind infra, AWS SageMaker, GCP Vertex, Databricks engineering, Snowflake engineering, Meta PyTorch infra, NVIDIA NeMo, Hugging Face inference infra.
Locations: London, Doha, Dubai, or Abu Dhabi.