Inference Engineer (Member of Technical Staff)
Own performance and reliability across the whole serving path, from the GPUs up to the API agents call
San Francisco Bay Area · Full-time
You'll work on performance and reliability across the whole serving path — from the GPUs up through the API layer that agents actually call. This is a founding-team role open to all experience levels, including new grads: you'll ship code that runs in production against real agent workloads, watch your changes show up directly in customer cost and latency, and learn the deep end of the stack on the job.
What you'll work on
- GPU side: push throughput and bring down time-to-first-token and tokens/sec on workloads that are heavy on prompt, long-running, and high-fanout — batching, scheduling, parallelism, quantization.
- Cross-platform serving: we're building for both NVIDIA and AMD GPUs — bring models up on each, close the performance gap between them, and keep the stack portable rather than vendor-locked.
- CPU side: the serving control plane that keeps large agent workloads fast and cheap — you'll own pieces of it end to end.
- Bring up and tune new open-weight model families as they release.
- Build the profiling and load-testing harness that tells us — honestly — what the system is doing under real agent load, not synthetic benchmarks.
- Decide what we fork, what we contribute back upstream, and what we build ourselves.
You probably have
- Strong software engineering fundamentals — systems, performance, or distributed computing, whether from work, internships, research, or personal projects. New grads are welcome.
- A genuine interest in LLM inference. Maybe you've run models locally, poked at vLLM, SGLang, or llama.cpp, or just read the serving papers because you couldn't help it.
- Curiosity about what's actually happening under the hood — you'd rather profile and find the real bottleneck than guess.
- A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
- A bias toward shipping. Most wins here come from a careful change that lands this week, not a six-month rewrite.
Bonus
- Open-source contributions to vLLM, SGLang, TensorRT-LLM, llama.cpp, or similar.
- Experience with GPU programming on either vendor — CUDA, Triton, or ROCm / HIP — profiling tools (Nsight, rocprof, PyTorch Profiler), quantization, or MoE serving.
- Cloud infrastructure experience — AWS, container orchestration, databases, load balancers, networking, or storage.
Compensation
Competitive base + meaningful early-stage equity. Founding engineers shape the platform, the technical direction, and the team we build around it.
How to apply
Send a note, a CV or GitHub, and a paragraph on the most interesting inference-related thing you've built, shipped, or dug into to hello@coralbricks.ai — subject line MTS Inference. A class project, a home-lab experiment, or a blog post you wrote all count. If you've contributed to one of the projects above, link the PR.