Inference Research Engineer (Member of Technical Staff)
Research and build new ways to make LLM inference faster and cheaper for real agent workloads
San Francisco or remote · Full-time
About Coral Bricks
Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.
We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.
The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.
The role
You'll research and build new ways to make LLM inference faster and cheaper, then prove them against real agent workloads. The work sits between research and systems engineering: form a hypothesis about where time or memory is going, design the experiment, implement the change, and measure whether it survives contact with production-shaped traffic.
This role is distinct from our GPU infrastructure role. You won't own the day-to-day operation of the fleet or deployment platform. You'll own the performance ideas that change what the serving system can do: new scheduling policies, cache strategies, parallelism approaches, quantization methods, and model-specific optimizations.
This is a founding-team role open to all experience levels, including new grads. You'll work close to production, see your changes show up directly in customer cost and latency, and learn the deep end of the stack on the job.
What you'll work on
- Research ways to push throughput and bring down time-to-first-token and inter-token latency on workloads that are heavy on prompts, long-running, and high-fanout — batching, scheduling, parallelism, caching, and quantization.
- Cross-platform serving: we're building for both NVIDIA and AMD GPUs — bring models up on each, close the performance gap between them, and keep the stack portable rather than vendor-locked.
- Prototype changes across the GPU and CPU sides of the serving stack when the research requires it, and carry successful ideas far enough to demonstrate them under production-shaped load.
- Bring up and tune new open-weight model families as they release.
- Build the profiling and load-testing harness that tells us — honestly — what the system is doing under real agent load, not synthetic benchmarks. The developers in our Discord are where those workloads come from, and the first to tell you whether a change actually helped.
- Decide what we fork, what we contribute back upstream, and what we build ourselves.
You probably have
- Strong software engineering fundamentals — systems, performance, or distributed computing, whether from work, internships, research, or personal projects. New grads are welcome.
- A genuine interest in LLM inference. Maybe you've run models locally, poked at vLLM, SGLang, or llama.cpp, or just read the serving papers because you couldn't help it.
- Curiosity about what's actually happening under the hood — you'd rather profile and find the real bottleneck than guess.
- A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
- A bias toward shipping. Most wins here come from a careful change that lands this week, not a six-month rewrite.
Bonus
- Open-source contributions to vLLM, SGLang, TensorRT-LLM, llama.cpp, or similar.
- Experience with GPU programming on either vendor — CUDA, Triton, or ROCm / HIP — profiling tools (Nsight, rocprof, PyTorch Profiler), quantization, or MoE serving.
- Experience designing and evaluating systems research: careful baselines, useful instrumentation, reproducible experiments, and skepticism about surprising results.
Compensation
$120,000–$200,000 base salary, plus 0.25%–2.0% equity. Both bands are wide on purpose: this role is open from new grad through senior, and we'd rather post the real span than a number that only fits one end of it. Where you land depends on experience, and cash and equity move together — take less of one and we'll weight the other.
Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.
Founding engineers shape the platform, the technical direction, and the team we build around it.
How to apply
Send a note, a CV or GitHub, and a paragraph on the most interesting inference-related thing you've built, shipped, or dug into to hiring@coralbricks.ai — subject line MTS Inference. A class project, a home-lab experiment, or a blog post you wrote all count. If you've contributed to one of the projects above, link the PR.