Distributed Systems Engineer, GPU Infrastructure
Own the clusters, GPU fleet, and production systems that turn inference research into a reliable service
San Francisco or remote · Full-time
About Coral Bricks
Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.
We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.
The team is small, technical, and shipping.
The role
You'll own the operational systems behind our inference platform: the clusters, GPU fleet, deployment machinery, and control plane that keep models available and traffic moving. When research produces a faster serving technique or a new model drops, you'll turn it into a repeatable, observable, production launch.
This role is distinct from our inference research role. You won't be measured on inventing a new attention kernel. You'll be measured on whether we can provision capacity, place workloads, ship changes, recover from failures, and operate a growing fleet without heroics.
This is a founding-team role with broad ownership. You'll work across cloud infrastructure, distributed systems, networking, storage, deployment, and the serving layer where they meet.
What you'll work on
- Own our GPU clusters and fleet across cloud providers: capacity, provisioning, machine images, drivers, networking, storage, health, and cost.
- Build the control-plane systems that place workloads, manage capacity, drain and replace unhealthy nodes, and recover cleanly from failures.
- Turn model launches into a reliable process: bring up new weights, validate serving configurations, roll out safely, watch production behavior, and roll back when needed.
- Build deployment and release systems for inference servers and the services around them, with fast feedback and clear failure modes.
- Create the observability we need to operate the fleet: metrics, logs, traces, dashboards, alerts, and tools that make incidents diagnosable instead of mysterious.
- Improve reliability at every layer — autoscaling, load balancing, failover, backpressure, graceful degradation, and capacity planning.
- Automate recurring operational work so the fleet can grow faster than the team operating it.
You probably have
- Strong backend or distributed-systems fundamentals and experience owning production services end to end.
- Experience with Linux, containers, networking, and at least one major cloud platform. You can debug across application, host, and infrastructure boundaries.
- Good instincts around reliability: staged rollouts, observability, failure isolation, incident response, and simple systems that are easy to operate.
- Comfort working from symptoms to root cause. A failed launch, an unhealthy node, or a latency spike is a systems problem to investigate, not a ticket to hand off.
- A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
- A bias toward shipping and automation. You fix the immediate problem, then build the mechanism that keeps it from becoming routine work.
Bonus
- Experience operating GPU or accelerator fleets, including NVIDIA or AMD drivers, topology, health checks, and failure modes.
- Experience with Kubernetes, Nomad, Slurm, ECS, or another cluster scheduler — especially if you've had to work below its happy path.
- Familiarity with vLLM, SGLang, TensorRT-LLM, PyTorch distributed, NCCL, or other model-serving and collective-communication systems.
- Experience with multi-cloud capacity, bare-metal provisioning, or scarce-resource scheduling.
- You've built an internal platform, scheduler, deployment system, or piece of infrastructure that other engineers trusted in production.
Compensation
$120,000–$200,000 base salary, plus 0.25%–2.0% equity. Where you land depends on experience, and cash and equity move together — take less of one and we'll weight the other.
Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.
Founding engineers shape the platform, the technical direction, and the team we build around it.
How to apply
Send a note, a CV or GitHub, and a paragraph about the most demanding production system you've owned to hiring@coralbricks.ai — subject line Distributed Systems Engineer. Tell us what failed, how you diagnosed it, and what you changed afterward.