@ 1001ai
ML Infrastructure Engineer at 1001 AI (P1). Hands-on ML infra hire who has deployed, served, and monitored models in production. Owns the platform underneath every model shipped.
Responsibilities: Build and operate ML serving infrastructure across multiple enterprise deployments; own model deployment pipelines (model registry, model artifacts, datasets, model deployment security, multi-tenant serving); stand up monitoring, logging, and observability for model performance; optimize inference latency, throughput, and cost across CPU and GPU workloads; maintain shared ML platform components so product teams can self-serve.
Requirements: 4+ years in infrastructure or platform engineering, with a meaningful chunk in ML / MLOps; has actually run ML in production; hands-on with model serving (Triton, TorchServe, Ray Serve, vLLM, BentoML, KServe, or equivalent); strong with Kubernetes, containers, CI/CD, infrastructure-as-code (Terraform); production cloud experience (AWS / GCP / Azure); monitoring stack (Prometheus, Grafana, OpenTelemetry, or equivalent); Python and TypeScript.
Bonus: GPU optimization, distributed training, large-model serving.
Company: 1001 AI building AI-native OS for critical MENA industries. Lux/GC/CIV. Mostly ex-Scale AI / Palantir.