@ Wafer
Member of Technical Staff — Wafer (San Francisco, on-site)
Wafer's mission is to maximize intelligence per watt: using AI to optimize AI infrastructure for orders-of-magnitude better energy and cost efficiency per token. Wafer commercializes this by serving optimized LLMs as an inference cloud. Six people, recently raised a large round, doubling headcount. Visa sponsorship available. SF office, in person.
Role: core engineer on the inference stack. Deep work on LLM inference optimization — GPU kernels, compilers, serving systems, quantization, batching/scheduling, KV-cache and memory management, energy/cost-per-token efficiency.
CALIBRATION (from Theo): the bar is incredibly, incredibly high — inference people who are demonstrably top-tier ONLY. Think: contributors/committers to vLLM, SGLang, TensorRT-LLM, FlashAttention, Triton (OpenAI), llama.cpp, TVM, XLA; kernel and performance engineers from NVIDIA (TensorRT/cuDNN), OpenAI, Anthropic, DeepMind, Together AI, Fireworks, Groq, Cerebras, Modal, Baseten; PhDs/researchers in ML systems from top schools (Stanford, Berkeley, MIT, CMU, UW) with real systems artifacts. Hard skills: CUDA/Triton kernels, torch.compile/inductor, distributed inference, speculative decoding, quantization (AWQ/GPTQ/FP8). Anti-patterns to reject: generic MLOps/platform engineers, model-training-only researchers with no systems depth, API-integration "AI engineers", short tenures, no artifacts. US relocation to SF required (visa sponsorship offered, so international top-tier is acceptable if clearly exceptional).