Inference Engineer
About Gradium
Gradium is a frontier voice AI company on a mission to redefine how humans interact with machines.
Voice remains the most human, highest-stakes channel, but also the most broken. Long wait times, rigid IVRs, low automation, and poor handoffs create frustration on both sides of the line. Gradium is rebuilding voice from the ground up with proprietary models: real-time understanding, autonomous resolution, and seamless escalation when humans matter most.
We're a team of world-class talent, with initial traction and a clear belief: voice will be the next major frontier of applied AI. We recently raised a $100m seed round and are backed by top-tier investors. Our goal is not incremental improvement: it's to make AI-powered voice interactions feel reliable, scalable, and economically transformative.
Expect early-stage reality: high autonomy, fast decisions, unreasonable ambition, few handoffs, and very little process unless it earns its place.
The Role
We're looking for an inference engineer to make Gradium's voice AI the fastest and most reliable on the market. You'll work across the full inference stack, from optimizing serving pipelines to writing CUDA when it's needed, and directly shape the latency our customers feel on every single request.
You'll work at the core of the company, partnering closely with the founders and the research team to take models from prototype to production through a wide variety of deployment stacks. This role exists because our edge is not just model quality but how fast we serve it at scale, and that is an engineering problem that needs an owner.
What You'll Do
Serve voice models at ultra-low latency: Optimize end-to-end inference pipelines so our models respond faster than anything else on the market. You'll directly impact the latency customers feel on every single request.
Go deep on the GPU stack: Write and tune CUDA kernels, implement quantization (FP8/INT8), and apply parallelism strategies to get the most out of every GPU. Profile and debug performance bottlenecks across the full pipeline.
Run high-throughput serving in production: Deploy and maintain model serving infrastructure that stays fast and stable under real load.
Bring new models to prod: Work closely with research to take new models from prototype to production, fast.
Who You Are
Founder mindset: You act with urgency, take full ownership, and don't wait for permission or perfect information. You are comfortable making high-stakes decisions in ambiguous environments and see the founding team as partners, not hierarchy.
Production-obsessed engineer: You care about how systems behave under real load, not just in a demo. You measure everything, chase latency and reliability relentlessly, and take pride in things that don't break.
Deep GPU and serving expertise: You have strong C++ and CUDA programming skills and proficiency with PyTorch or JAX. You've served generative AI models in production (vLLM or equivalent) and have a track record of measurable latency or throughput improvements.
AI-fluent operator: You use AI-powered tools in your daily work and can always articulate why. You're relentless about pushing what's possible and reinventing how you and your team operate.
Bonus points for Triton, TensorRT, or custom kernel development; pipeline or tensor parallelism on multi-GPU clusters; real-time audio or streaming inference constraints; Kubernetes and distributed systems.
Still interested?
We'd love to hear about you!