AI Infrastructure Engineer
Automate the cloud infrastructure and operational workflows behind a fast-moving AI platform, and grow into the part of the stack that suits you
San Francisco or remote · Full-time
About Coral Bricks
Our mission is to make frontier intelligence affordable and accessible to everyone. Frontier models are finally here — but almost nobody can afford to use them freely. People have token anxiety: they meter every call, ration every context window, and settle for weaker models because the best ones are priced out of everyday use.
We're building the inference platform that ends that, starting with the workloads that feel the squeeze hardest: research and coding agents that swarm across multiple models, plan, call tools for hours, and reason over big context. Classic LLM serving was never built for them — rate limits that throttle real workloads, queues that stretch a 20-minute job into a 4-hour one, costs that grow with every agent turn. Same models, same prompts — many times the tokens per second at a fraction of the cost.
The team is small, technical, and shipping. We also build in the open: a lot of the day-to-day happens in our Discord, where the developers building on Coral Bricks tell us what broke, compare numbers with us, and push on what we work on next.
The role
You'll help build and operate the cloud infrastructure around our AI platform. One day that might mean automating a deployment that still has manual steps; the next, adding the dashboard that makes a production issue obvious or writing a tool that turns a recurring operational task into a button or a command.
This is an early-career role designed for someone with around one to two years of professional experience. We don't expect you to arrive as an expert in GPU clusters, distributed systems, or LLM serving internals. We do expect you to be a solid programmer, comfortable in a terminal, eager to understand how production systems behave, and ready to take ownership of concrete projects while learning the deeper parts of the stack.
You'll work closely with experienced systems engineers and the founders. As you grow, so will the scope you own — and which direction it grows in is open. Some of this work leads toward the GPU fleet and the systems that operate it; some of it leads toward the inference performance work that makes the fleet fast.
What you'll work on
- Automate recurring infrastructure, operational, and internal workflow tasks with scripts, services, and internal tools.
- Improve how we build, test, deploy, and roll back services across development and production environments.
- Help operate our AWS infrastructure, including compute, containers, networking, storage, permissions, and secrets.
- Build useful observability: metrics, logs, dashboards, alerts, and runbooks that help the team find problems quickly.
- Benchmark and load-test serving configurations so we argue about real numbers instead of guesses.
- Support model and service launches by validating configurations, monitoring rollouts, and improving the launch process after each release.
- Investigate production issues across application and infrastructure boundaries — often starting from a developer's report in our Discord rather than an alert — then turn what you learn into a lasting fix or better automation.
- Use AI coding agents to move quickly, while reviewing their work carefully and understanding the systems you change.
- Keep infrastructure understandable: clear code, small changes, useful documentation, and fewer one-off manual procedures.
You probably have
- Around one to two years of professional software engineering, infrastructure, DevOps, site reliability, or ML systems experience. Strong internships, open-source work, or substantial personal projects can count too.
- Solid programming fundamentals in Python, TypeScript, Go, or a similar language. You can write maintainable code, not just one-off shell commands.
- Hands-on experience with Linux, Git, containers, and at least one cloud platform, whether from work or projects.
- Experience with some part of the software delivery loop: CI/CD, deployment automation, infrastructure as code, monitoring, or production support.
- A methodical approach to debugging. You follow the evidence, ask good questions, and keep going when the first explanation is wrong.
- A high work ethic and excitement about early-stage startups. The pace is fast, the problems are open-ended, and everyone does a bit of everything.
- A bias toward automation and shipping. When you do something twice, you start thinking about how the system should do it for you.
Bonus
- Experience with AWS services such as ECS, EC2, ECR, IAM, CloudWatch, S3, or Amplify.
- Familiarity with Terraform, Pulumi, CloudFormation, or another infrastructure-as-code tool.
- Experience operating a production service, participating in incident response, or building tools used by other engineers.
- Curiosity about LLM serving, GPUs, vLLM, SGLang, Kubernetes, or distributed systems. Prior production experience with them is not required.
- Any experience profiling or benchmarking something and making it measurably faster.
- A project where you used an AI coding agent to build something ambitious, automate a workflow, or explore an unfamiliar system.
Compensation
$100,000–$150,000 base salary, plus 0.1%–0.75% equity. Where you land depends on experience, and cash and equity move together — take less of one and we'll weight the other.
Equity vests over four years with a one-year cliff. Health, dental, and vision coverage, and flexible time off.
Early engineers shape the platform, the technical direction, and the team we build around it.
How to apply
Send a note, a CV or GitHub, and one example of something you've automated, deployed, operated, or made faster to hiring@coralbricks.ai — subject line AI Infrastructure Engineer. Tell us what the problem was, what you built, and what you learned when it met the real world. Work projects, internships, open-source contributions, home labs, and personal projects all count.