I build software and design infrastructure systems that help founders, ML researchers, and customer engineering teams run GPU training and inference workloads on production infrastructure.
After eight years as a Meteorologist in the U.S. Marine Corps, I moved into systems engineering across bare metal, GPU platforms, cloud infrastructure, and research environments. For the last several years, I’ve worked in fast-moving startups turning ambiguous infrastructure and product requirements into shipped, reliable systems.
My work spans customer workload onboarding, distributed training, batch inference, ML pipeline observability, autoscaling infrastructure, utilization-based billing, and workload resiliency. I've written and deployed custom Kubernetes controllers, host agents, and control-plane integrations, and help adapt manifests, containers, serving stacks, model caches, telemetry, and deployment workflows so customer workloads run reliably across virtual machines, Kubernetes, and bare-metal GPU fleets.
I have hands-on production experience with NVIDIA Hopper/Blackwell and AMD Instinct environments, vLLM, SGLang, TensorRT-LLM, Triton, PyTorch, Candle, tch-rs, NCCL debugging, kernel and BIOS tuning, high-bandwidth data movement, and large-scale GPU performance engineering.
My infrastructure work spans Kubernetes, Slurm on Kubernetes (SUNK), Software Defined Networking, AWS, GCP, DigitalOcean, co-located bare metal, Terraform/OpenTofu, Terragrunt, Argo CD, Cluster API, Kamaji, QEMU/KVM, Prometheus/Grafana, Weka, VAST, Ceph, Postgres, Google Cloud SQL, Rust, Go, and Nix/NixOS.
Recent work includes multi-cluster observability for 12,000+ GPUs across 10+ providers and regions, customer workload onboarding for training and inference, hosted metrics, automated Slack alerting, Stripe-backed utilization billing, Cluster API-compatible infrastructure APIs, custom Kubernetes controllers, host-side agents, control-plane automation, and reproducible test/dev environments using Nix, Tilt, Kubernetes, and QEMU/KVM.
I’m also an advocate for local inference, with hands-on experience with tools for M1/M2 Apple Silicon using Asahi Linux diagnostics, and exploring Apple Container orchestration with MLX and JACCL.
Links:
- Resume PDF: https://kennethdsheridan.github.io/resume/resume.pdf (web)
- Forgejo: https://forgejo.kennysheridan.io
- Codeberg: https://codeberg.org/kennysheridan
- GitHub: https://github.com/kennethdsheridan
- LinkedIn: https://www.linkedin.com/in/kennethdashensheridan/




